“A lot of people, especially young people, turn to these systems for support, and usually they aren’t actually getting the help they need. But in many cases, they’re actively being harmed, and people unfortunately have taken their lives already,” Arul said. “Those sorts of safety vulnerabilities, where people aren’t necessarily actively trying to break the system — they’re engaging in a natural way — and the system has context pollution or it doesn’t understand the nuance, and then takes really dangerous action, we’re trying to prevent that.”
Circuit Breaker Labs describes its testing agents as AI “crash-test dummies.” Instead of relying only on standard prompts, the system is built to reflect how people actually communicate, including slang, coded language, typos and differences in how people express themselves across age groups and cultural backgrounds.
“The way a six-year-old girl versus a 45-year-old man, or someone who speaks English as a first language versus a second language, or … gamer slang versus someone else who uses a different kind of slang, all of those can really trip up a model,” Shirali said. “Models are really good at handling standard speech patterns, but nobody actually talks like that and so if the model misunderstands nuance or slang, it can go really badly.”
The startup works with human domain experts to design those simulations, then uses them in red-team testing intended to expose weaknesses before real users encounter them. Circuit Breaker Labs says it can run tens of thousands to hundreds of thousands of simulated conversations per day.
Its tests are meant to examine both obvious and gradual failures. That includes whether a model recognizes signs of suicidal ideation, understands slang correctly and continues responding safely as a conversation develops over multiple turns.
Circuit Breaker Labs then applies its own scoring system to the results, producing assessments designed to be auditable and explainable. The company is currently focused on higher-risk AI products, including coaching, journaling and mental health support applications.
Arul declined to identify the company’s major customers. Circuit Breaker Labs remains small, with five employees including the two founders, and the product is still in an early stage.
The company sees a broader use case beyond explicitly mental-health-focused products. Its testing approach could also be applied to conversational systems where repeated interaction may encourage users to form strong emotional or parasocial attachments.
That concern has become more prominent as families have taken legal action against chatbot companies. Character.AI settled several wrongful death lawsuits earlier this year involving underage users, while multiple families have sued OpenAI over allegations involving ChatGPT, suicide and delusional behavior.
Circuit Breaker Labs is positioning its product as a way for developers to identify those kinds of risks before deployment rather than after harmful interactions occur.
“People are becoming more skeptical of AI or more resistant to adopt it across the board,” Arul said, while arguing that blocking useful systems over safety concerns would be “regressive.”
The company’s goal, he said, is to make those systems safer. “We want to help build that trust for people.”
This analysis is based on reporting from TechCrunch & Crypto Briefing.
Image courtesy of Circuit Breaker Labs.
This article was generated with AI assistance and reviewed for accuracy and quality.