Founder and CEO Sudarshan Kamath argues that existing voice AI systems remain limited because they rely on separate components for speech recognition, language model processing, orchestration, memory, text-to-speech, and safety controls. That layered design introduces delays and contributes to conversations that still sound artificial.
“Humans don’t wait for someone to finish speaking before they begin thinking. We listen, think, and respond simultaneously,” Kamath said. “Voice AI needs to work the same way. By rethinking the stack instead of simply scaling models, we’re reducing latency to the point where voice interactions feel genuinely human.”
Smallest.ai describes Voice 4.0 as the next stage in the evolution of conversational AI. According to the company, earlier voice systems relied on rigid interactive voice response trees before progressing to machine learning-powered voice bots and later generative AI voice agents. Hydra is intended to move beyond those approaches by allowing multiple processes to occur simultaneously instead of one after another.
The model can also be combined with the company's speech-to-text technology to support rapid transcription with latency measured in milliseconds.
Hydra expands Smallest.ai's broader voice AI platform, which includes speech recognition models such as Pulse STT Pro and Lightning V3.1. The company says those models rank among the top voice AI systems on the Artificial Analysis benchmark.
Since introducing Lightning, Smallest.ai has expanded the platform to support 38 languages while adding features including emotion detection, speaker diarization, data redaction, and noise reduction. The technology has been adopted by customers including RingCentral, Truecaller, Kogta Financial, and Readymode. According to the company, some deployments have reduced customer support costs by as much as 80%.
Kamath said the company is focused specifically on real-time conversational voice agents rather than broader AI assistants. Instead of relying entirely on a large language model, Smallest.ai uses a smaller voice model to manage live conversations and can hand more complex requests to a larger foundation model when additional reasoning is required.
The company believes that specialized voice models are better suited to handling conversational details such as interruptions, diverse accents, multiple languages, and noisy environments while maintaining low response times.
Smallest.ai enters an increasingly competitive voice AI market that includes companies such as ElevenLabs, Cartesia, Fish Audio, and Sarvam. While some competitors focus on applications like audio generation or dubbing, Smallest.ai says its primary focus is enterprise voice agents for customer interactions.
Seligman Ventures partner Ashish Kakran said the company is taking a different approach by redesigning the underlying voice AI architecture rather than requiring developers to assemble multiple independent models. “Developers now increasingly talk to their machines instead of typing code,” he said. “Smallest.ai is taking a fundamentally different approach to the category by rethinking architecture itself. Customers get an efficient vertically integrated stack and don’t need to waste time stitching models together.”
Kamath said the company's long-term goal is to make AI voice agents indistinguishable from human speakers during live conversations. “We want our models to break the Turing test,” he said. “You should speak to our model and not know it’s AI or human. That’s the sole focus of the company.”
This analysis is based on reporting from SiliconAngle.
Image courtesy of Smallest.ai.
This article was generated with AI assistance and reviewed for accuracy and quality.