“Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can't be solved from a transcript,” CEO and co-founder Carter Huffman said.
Modulate’s main platform, Velma, analyzes characteristics of speech including emotion, tone, intent, emphasis, conversational behavior and whether a voice is synthetic. Those signals can then be combined to identify events such as suspected fraud, harassment, customer frustration, policy violations or problems with an AI voice agent.
The system can operate while a conversation is taking place, allowing applications to respond to events before a call or interaction has ended.
Rather than relying on a single large foundation model, Velma is powered by Modulate’s Ensemble Listening Model architecture. The system coordinates more than 100 specialized audio models and selects combinations suited to individual tasks.
Modulate says that architecture can be up to 1,000 times more efficient than using one large model. The company also says Velma produces twice the true-positive accuracy of traditional large language models while generating seven times fewer false positives.
Its technology is already processing more than 10 million hours of audio each month, according to the company, with total analyzed audio recently passing 600 million hours.
Modulate has also posted results on public benchmarks. Its transcription technology reached the top position on Hugging Face’s Open ASR Leaderboard, while its deepfake speech detection system has also ranked first on a Hugging Face benchmark. The company says its deepfake detector achieves 98.9% accuracy on public benchmark data, while batch transcription is priced at $0.03 per hour.
Beyond transcription, Modulate says organizations are using its models to detect synthetic voices and suspicious behavior, monitor the performance of voice AI agents, identify harassment and child grooming in online conversations, and mask voices to protect people working in higher-risk situations.
The company initially built a presence in gaming through ToxMod, its voice moderation system. It is now expanding further into security, communications, customer experience and AI agent supervision as voice-based AI applications become more widely used.
Steve Jurvetson, co-founder of Future Ventures, said Modulate has “gained a significant technical lead in audio-native AI,” adding that demand for the technology is expanding into areas including AI agents, security and customer experience.
The new capital will support additional research in AI and machine learning as well as product development, engineering and partnerships. Modulate also plans to expand its SDKs, APIs and deployment options as it tries to make its audio models easier for outside developers to incorporate into their own products.
“Developers shouldn't have to rebuild the audio intelligence layer every time they create a new voice experience,” Huffman said. “Our mission is to build the models and infrastructure that let them focus on the application they want to create.”
With the funding, Modulate is betting that understanding how something is said — not simply turning speech into text — will become an increasingly important part of building and supervising voice-based AI systems.
This analysis is based on reporting from Yahoo Finance.
Image courtesy of Modulate.
This article was generated with AI assistance and reviewed for accuracy and quality.