- Team
- Models
- Location
- Remote (IST ±4)
- Type
- Full-time
- Package
- Competitive · equity
About the role
You'll sit between research output and production reality — taking models that work in a notebook and making them serve at 12ms p50 with a contract behind the number.
Expect to spend as much time on evaluation harnesses and failure modes as on throughput.
What you'll do
- Build and tune the model-serving path: batching, caching, routing, and fallback
- Own the evaluation harness — offline benchmarks and online quality signals
- Design guardrails, and the observability to prove they're working
- Advise clients on build-vs-buy and model selection with evidence, not vibes
What we're looking for
- 4+ years shipping ML systems to production, not just training them
- Strong Python, and comfort dropping into a lower-level language when latency demands it
- Hands-on with inference servers and the throughput/latency trade-offs they force
- You can explain a p99 regression to a non-specialist stakeholder
Nice to have
- Experience with LLM serving, agent orchestration, or retrieval systems
- Published evaluation or benchmarking work
Not quite your role? See the other three openings.