DevOps Engineer
Kenpath Labs is a voice AI lab in Bengaluru. We build speech models that work in the languages most systems skip.
Our open model, Svara TTS v1, has passed a million downloads. Our current model, Svara TTS Turbo, speaks 80+ languages in hundreds of voices, switches language mid-sentence, and returns first audio in about 80 ms. It runs today behind a self-serve API at platform.kenpathlabs.com.
We are a small team. You would work on a model that people already use in production, and you would see your work ship.
Location & work mode
This role is based in Bengaluru, India, and works from our office. We care about focused time and short feedback loops, which is easier in a room together while the team is this size.
What you'll do
- Run the GPU inference fleet the API serves from: capacity, autoscaling, rollout, and what happens when a node dies mid-request.
- Keep time to first audio low in production, not just in a benchmark — that means real work on cold starts, model loading, and request routing.
- Own deployments end to end: containers, CI/CD, and releases that do not need a person watching them.
- Build the observability that makes an incident short — metrics, logs, traces, and alerts that fire on things that matter.
- Handle multi-region: we serve from the EU, the US and India, and Enterprise customers ask about residency.
- Keep the platform secure and the secrets where they belong.
What we're looking for
- Experience running production infrastructure that people depend on, including the on-call side of it.
- Real Kubernetes and container experience, and a clear sense of when not to reach for them.
- Comfort with GPU workloads: scheduling, drivers, memory, and why the second request is faster than the first.
- Infrastructure as code, and CI/CD you have built rather than inherited.
- Enough Python to read and debug the services you are deploying.
- Clear writing. We work from written arguments more than meetings.
Bonus
- You have served ML models at scale — vLLM, Triton, or something you had to write yourself.
- You have worked with streaming or low-latency systems, WebSockets included.
- You have run infrastructure across more than one cloud or GPU provider.
- You have been the person who made a slow deploy fast.
What we offer
- Competitive salary, discussed early in the process.
- Direct ownership of a system in production, not a slice of a backlog.
- A small team and a short path from idea to shipped.
- Hardware and compute you need to do the work.
- Based in our Bengaluru office.
Apply
This goes straight to a person, not a queue. A short note about what you have built beats a long one about what you are looking for.