Speech AI Engineer
Kenpath Labs is a voice AI lab in Bengaluru. We build speech models that work in the languages most systems skip.
Our open model, Svara TTS v1, has passed a million downloads. Our current model, Svara TTS Turbo, speaks 80+ languages in hundreds of voices, switches language mid-sentence, and returns first audio in about 80 ms. It runs today behind a self-serve API at platform.kenpathlabs.com.
We are a small team. You would work on a model that people already use in production, and you would see your work ship.
Location & work mode
This role is based in Bengaluru, India, and works from our office. We care about focused time and short feedback loops, which is easier in a room together while the team is this size.
What you'll do
- Train and fine-tune text-to-speech and speech-to-speech models, from data through to a checkpoint that ships.
- Own the quality bar for a language or a voice family: what good sounds like, how it is measured, and what has to change to get there.
- Work on the parts that make speech feel live — streaming inference, time to first audio, and the tradeoffs between latency and quality.
- Build the evaluation that tells us we improved: objective metrics, listening tests, and the tooling around both.
- Take on the hard cases directly — code-switching, low-resource languages, expressive delivery, and voice cloning.
What we're looking for
- Experience training speech or audio models end to end, in PyTorch, and getting them to a state someone else can run.
- A working understanding of modern TTS: neural codecs, autoregressive and diffusion approaches, vocoders, and where each breaks.
- Comfort with the unglamorous half — data collection, cleaning, alignment, and normalization, which is usually where the quality actually comes from.
- The judgement to tell a real improvement from a metric that moved.
- Clear writing. We work from written arguments more than meetings.
Bonus
- You have worked on Indic or African languages, or any low-resource setting.
- You have shipped speech-to-speech or real-time voice systems.
- You have published, or contributed to open-source speech work.
- You speak more than one of the languages we support, and can hear when a voice is wrong.
What we offer
- Competitive salary, discussed early in the process.
- Direct ownership of a system in production, not a slice of a backlog.
- A small team and a short path from idea to shipped.
- Hardware and compute you need to do the work.
- Based in our Bengaluru office.
Apply
This goes straight to a person, not a queue. A short note about what you have built beats a long one about what you are looking for.