Tavus has introduced Griffin, a full-duplex video-to-video AI model designed for real-time face-to-face conversation. The company calls Griffin its first Human Interaction Model, or HIM, and says the system continuously processes video and audio while generating speech, facial expressions, gestures and other responses as an interaction unfolds.
Unlike conversational systems that wait for one person to finish speaking before responding, Griffin repeatedly reassesses an exchange at sub-second intervals. That allows it to react while either participant is talking, including through interruptions, backchannels, pauses, changes in expression and shifts in tone.
The model combines capabilities Tavus previously developed across separate systems for perception, conversational timing and video generation. Griffin brings those functions into a single architecture where visual and audio perception, response decisions, speech generation and video output operate continuously.
Visual information is central to that approach. Griffin can use facial expressions, gaze and a person’s surroundings as part of its understanding of a conversation, rather than relying only on spoken words. Tavus says that can allow the model to respond differently when someone pauses to think, looks away or shows an object to the camera.
Griffin’s generation system also controls more than a virtual person’s face. Tavus says it generates the full scene in real time from a reference image, including body movements and elements of the surrounding environment.
The company divided the system into two main components. A continuous conversational modeling engine processes incoming audio and video and determines what the model should say and how it should behave. A separate audiovisual generation system turns those instructions into streamed speech and video.
For speech, Griffin-Lite can clone a voice using about 10 seconds of reference audio, according to Tavus. Its video system produces 720p output in 320-millisecond segments and can adjust elements such as gaze, gestures and emotional expression as new instructions arrive.
Tavus is releasing Griffin-Lite to a limited group of trusted testers as a research preview rather than making it generally available. The company said it is working on disclosure and other safety mechanisms before a broader release because increasingly realistic conversational AI can be difficult for users to distinguish from another person.
That concern is reflected in Tavus' own testing. In a live study involving 54 participants, each person had a one-minute video conversation without initially being told that the other participant was an AI. Twenty-six participants, or 48%, later said they believed they had spoken with a real person.
Tavus compared that result with its previous system, which combined Phoenix-4.5, Sparrow-2 and Raven-1. In a similar study involving 41 participants, one person, or 2.4%, thought the AI was human.
The company also tested Griffin-Lite on Nvidia's VideoFDB benchmark for full-duplex audiovisual conversation. On the generation portion, Griffin scored 3.83 out of 5, compared with a human reference score of 3.92 and 2.80 for the next-highest published system.
On VideoFDB's perception track, Griffin scored 3.73. The human reference was 4.20, while the strongest reported baseline scored 3.44. Tavus said Nvidia independently conducted the evaluation using the benchmark's published methodology.
“For decades, we’ve imagined computers as partners, not just tools,” Tavus CEO Hassaan Raza said. “AI has become incredibly intelligent, but we still have to meet the machine on its terms. Griffin is a step toward changing that — toward machines that understand how we naturally communicate and meet us where we are. That’s what we mean by human computing.”
Griffin builds on Tavus models including Phoenix-4.5 for real-time rendering, Raven-1 for multimodal perception and Sparrow-2 for conversational dynamics. Tavus said its existing technologies are used by more than 150,000 developers and businesses.
The company sees potential uses for Griffin in areas such as tutoring, conversation practice and visual customer support. In those scenarios, the model could combine what someone says with their behavior and what appears on camera rather than requiring the user to translate the problem into a text command.
For now, those applications remain part of a controlled research preview. Tavus said Griffin-Lite will not be generally available to customers while it continues evaluating the safety implications of AI systems capable of producing increasingly human-like face-to-face interactions.
About this article: This article was generated with AI assistance and reviewed by our editorial team to ensure it follows our editorial standards for accuracy and independence. We maintain strict fact-checking protocols and cite all sources.
Word count: 714Reading time: 0 minutes
Explore More AI Resources
Continue with high-value guides related to this topic.
Join thousands of weekly readers — the latest AI news, free in your inbox.
🤖 3 times a week📊 Industry analysis💡 Breaking news
Enjoying this article?
Get a free month of ChatAI Plus or Pro
🎁 Limited time offer
🔒 Once per user
ChatAI Plus & Pro unlock multi-model chat (GPT, Claude, Gemini & more), our productivity tools, and ad-free reading. Subscribe to our newsletter and take a 2-minute survey to claim one month free — no strings attached.