- Team
- Infrastructure
- Location
- Remote (IST ±4)
- Type
- Full-time
- Package
- Competitive · equity
About the role
You'll work on the substrate: multi-region compute scheduling, autoscaling under bursty inference load, and the storage and networking seams that everything above depends on.
This is a small team with real ownership. You will not be handed a ticket queue — you'll be handed a constraint and trusted to design your way out of it.
What you'll do
- Design and operate multi-region GPU compute with predictable cost and 99.99% availability
- Own infrastructure-as-code, CI, and the deployment path from merge to production
- Set the reliability bar: SLOs, error budgets, on-call practice, and post-incident review
- Work directly with client engineering teams during embedded engagements
What we're looking for
- 6+ years building and operating production distributed systems
- Deep Kubernetes and cloud-provider experience, and the scars to go with it
- Fluency in a systems language — Go, Rust, or equivalent
- You've been on-call for something that mattered, and improved it
Nice to have
- GPU scheduling or inference-serving experience
- Experience in a consulting or embedded-engineering model
Not quite your role? See the other three openings.