Smallest.ai raises $13M to build ultra-fast voice AI that sounds genuinely human
TechCrunch Marina Temkin ● Covered by 3 sources
Smallest.ai just banked $13M to make talking to AI agents feel less robotic. Their trick: small, fast voice models instead of bulking up big LLMs.
Based on reporting by TechCrunch, Marina Temkin — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
There's a specific moment in every AI customer support call where you just know. Maybe it's a beat too long before the response, or a phrasing that's technically correct but weirdly stilted. Smallest.ai, founded in late 2024, thinks it has found the fix, and it's not what most of the industry is chasing.
Instead of racing to make giant language models respond faster, the startup builds small, specialized voice models that mimic how humans actually converse: listening, thinking, and talking almost at the same time, with room to interrupt and no dead air. Founder and CEO Sudarshan Kamath put it plainly to TechCrunch — an LLM waits for your whole prompt before it starts processing, which is fine in a text box but excruciating out loud. Real conversation doesn't work in clipped audio chunks.
The company just closed a $13 million Series A led by Seligman Ventures, with Sierra Ventures and 3one4 Capital also writing checks, pushing total funding past $21 million. The model works as a real-time layer for narrow, well-defined conversations — handling accents, background noise, and dozens of languages without lag. When a customer asks something outside its wheelhouse, the system does something almost charmingly human: it puts the caller on brief hold and quietly kicks the question to a bigger foundational model to sort out, the same way a call center rep might say 'let me check on that.'
Kamath's bet is that this two-model setup — a lean, fast voice layer paired with an LLM held in reserve for the hard stuff — becomes the standard architecture for voice agents generally, not just something Smallest.ai does. Customers already include RingCentral and Truecaller, and Kamath is pitching newer customer-support platforms like Sierra and Decagon too, arguing that building deep voice expertise in-house would just distract them from their actual product.
The competitive field isn't empty. ElevenLabs remains the name most people know in voice AI, Cartesia is chasing similar real-time ambitions, and Sarvam has carved out a niche in regional languages. Smallest.ai's angle is narrower than most: no dubbing, no podcast tools, just enterprise voice agents built for one goal. Kamath says it outright — he wants a model that beats the Turing test, where you genuinely can't tell if you're talking to code or a person.
My take — AI-written commentary, not fact-checked reporting
Splitting voice handling from reasoning is the obviously correct architecture, and I'm mildly annoyed it's taken this long for someone to make it their whole company instead of a feature. The real test isn't whether it fools you on a good call — it's whether the handoff to the big model stays invisible when things get messy, which is exactly the moment most voice AI still faceplants.
Read more about this at: TechCrunch