While AI agents are increasingly capable of solving customer support problems, most people can still tell instantly when they’re speaking to a machine rather than a human. Smallest.ai, a startup founded in late 2024, is betting that the next leap in voice agents will come not from making large language models (LLMs) faster, but from using smaller, specialized models designed for natural human conversation. Simply put, the company aims to make talking to an AI agent indistinguishable from talking to a person.
To achieve this, Smallest.ai is developing a compact voice model that mimics how humans process information—listening, thinking, and speaking simultaneously. “While I’m speaking to you, you’re already thinking, and you might interrupt me if I talk for too long,” Sudarshan Kamath, founder and CEO of Smallest.ai, told TechCrunch. He added that this is exactly how the startup’s model is designed to operate.
To fuel this mission, Smallest.ai has raised $13 million in a Series A round, led by Seligman Ventures with participation from Sierra Ventures and 3one4 Capital. The fresh capital brings the startup’s total funding to over $21 million—a strategic boost as the demand for hyper-realistic voice AI accelerates into 2026, with enterprises prioritizing customer experience and reduced latency.
The Latency Problem in Voice AI
“The way an LLM works is you give it an entire prompt, and then it starts thinking,” Kamath explained. While that latency is acceptable in text chat, even a short pause feels unnatural in a voice conversation. “If you think about how we are talking, I’m not giving you a large clipping of my audio, and then you start thinking.”
The startup’s model serves as a real-time intelligence layer that enables natural customer conversations on specific topics, with virtually zero response lag. However, if the model encounters a subject outside its limited knowledge base, Smallest.ai hands off the query to a large foundational model, briefly placing the customer on hold to “research” the issue—just as a human would.
A Hybrid Approach for Future AI Agents
Kamath believes that all AI agents will soon rely on two models: a small voice model for real-time interaction and an “offline” LLM that is called upon as needed to solve complex problems. Unlike large foundational models, Smallest.ai focuses strictly on voice-specific nuances, such as handling diverse accents, supporting dozens of languages, and operating in noisy environments—critical features for global enterprise deployment.
The startup’s existing customers include voice-centric companies like RingCentral and Truecaller. Kamath noted that any customer support firm, including newer players like Sierra and Decagon, is a potential customer for the startup.
Why Outsourcing Voice AI Makes Sense
When asked why a well-funded AI customer support company wouldn’t build its own voice model, Kamath said that for these startups, becoming “extremely good at doing voice is a distraction from their core business.” This perspective positions Smallest.ai as a key enabler in the AI ecosystem, allowing customer support providers to focus on their primary offerings while leveraging specialized voice technology.
Smallest.ai competes with voice AI leader ElevenLabs, as well as Cartesia and regional players like Sarvam that focus on local languages. While some competitors apply voice AI to use cases like audio dubbing and podcasting, Smallest.ai concentrates exclusively on real-time conversational voice agents for enterprise clients.
“We want our models to break the Turing test,” Kamath said. “You should speak to our model and not know it’s AI or human. That’s the sole focus of the company.” As the industry moves toward more immersive AI interactions in 2026, Smallest.ai’s dedication to conversational authenticity could set a new standard for what users expect from voice-enabled systems.
