Tavus Griffin Passes Video Turing Test: 48% Mistook AI for Human, How to Choose Real-Time Video AI
Tavus released Griffin, the first Human Interaction Model (HIM) to pass the real-time video Turing test. In a 54-person blind study, 48% believed they were talking to a real human. NVIDIA VideoFDB score: 3.83 vs human reference 3.92. Griffin uses a full-duplex video-to-video architecture unifying perception, conversation modeling, and audiovisual generation.
Public-source compilation
Synthesized from public posts/docs. Prefer the original source for primary claims.
TL;DR
On October 1, 2026, San Francisco AI company Tavus released Griffin—the world's first Human Interaction Model (HIM) and the first AI to pass the real-time video Turing test. In a 54-person blind study, 48% of participants believed they had spoken with a real human after a one-minute video call, compared to a 2.4% pass rate for the previous industry-leading system. Griffin-Lite is available as a research preview to select trusted testers, with a more powerful model to follow after safety work is completed.
The Video Turing Test: What 48% Means
The study recruited 54 participants from the US and Europe through an independent research platform. They were told they would have a one-minute video call with another participant about "what they were looking forward to this year"—but their partner was actually a PAL (Personal AI Link) powered by Griffin-Lite.
After the call, participants rated their partner on a 7-point scale across five dimensions:
- Seeming natural: 5.4
- Seeming trustworthy: 5.6
- Would enjoy talking with again: 5.8 (held at 5.4 among those who said AI)
- Partner was really listening: 5.5
- Conversation flowed naturally: 4.9 (lowest of five axes)
Over half of participants said the possibility their partner wasn't real had not crossed their mind during the call. 26 people (48%) ultimately believed their partner was real. The previous industry-leading system (Phoenix 4.5 + Sparrow-2 + Raven-1) managed only a 2.4% pass rate (1 of 41 participants).
NVIDIA VideoFDB Benchmark: #1 on Both Tracks
NVIDIA independently built and scored VideoFDB, the leading industry benchmark for full-duplex audio-visual conversation.
Generation Track (tests whether the model produces the right behavior—fluency, affect matching, nonverbal behavior):
| System | Score (out of 5) |
|---|---|
| Human ground truth | 3.92 |
| Griffin-Lite | 3.83 |
| Gemini 2.5 + Anam (next-best) | 2.80 |
Griffin leads the next-best system by 1.03 points and is only 0.09 below human ground truth—more than 12x closer to human performance than any other system.
Perception Track (tests whether the model understands the moment—fluency, conversational flow, visual grounding):
| System | Score (out of 5) |
|---|---|
| Human ground truth | 4.20 |
| Griffin-Lite | 3.73 |
| MiniCPM-o 4.5 (strongest baseline, audio-only) | 3.44 |
| Gemini 2.5 Flash Native | 3.17 |
| OpenAI gpt-realtime (audio-only) | 2.97 |
Griffin is the only model evaluated on both tracks. Takeover-rate alignment was best in class on both tracks (62.8% generation, 73.8% perception).
Technical Architecture: Full-Duplex Video-to-Video
Griffin is not a traditional cascade system (ASR → LLM → TTS → face rendering). It is a unified video-to-video architecture with two engines:
1. Continuous Conversational Modeling Engine
- Perceives audio and video simultaneously (not audio alone)
- Makes conversational decisions at sub-second intervals (mini-turns), not once per turn
- Decides when and how to respond—speak, hold, yield, backchannel, nod, or stay silent
- Emits expressive control signals (emotional tone, stance, facial expression, gesture)
- Can be interrupted, can interrupt, adjusts mid-response without losing context
2. Audio-Visual Generation Engine
Streaming Speech Generation:
- Fast autoregressive diffusion transformer (VDiT) generates latent speech progressively
- Tavec Codec: convolutional autoencoder mapping 48kHz audio to compact continuous latent—40 values per frame × 100 frames/second, no codebooks
- Fully causal decoder emits audio packets as small as 10ms—speech begins before the sentence is finished
- Clones a speaker's voice from ~10 seconds of audio
- 1 minute of speech = 6,000 latent frames (vs 2.88 million raw audio samples)
Streaming Video Generation:
- Few-step autoregressive diffusion generator: real-time 720p video in 320ms chunks
- VAE compresses time 8x—one latent = 8 frames at 25fps = 320ms of video
- Inputs: reference image + streaming audio + streaming controls (gesture, gaze, emotion)
Three-Stage Distillation:
- Distribution Matching Distillation: teacher distilled into few-step student (fast, not streaming)
- Teacher Forcing → Autoregressive: converted to generate one latent at a time
- Self-Forcing: trained on own generated history for long rollouts without drift
Cascade Systems vs. Griffin
| Feature | Traditional Cascade | Griffin |
|---|---|---|
| Architecture | ASR → LLM → TTS → face (sequential relay) | Unified video-to-video, all concurrent |
| Turn-taking | Waits for user to finish, then thinks, then talks | Never stops thinking; decides every sub-second |
| Information loss | Every handoff discards info (tone, visual cues) | Perception is continuous—audiovisual throughout |
| Interruption | Poor—cannot be interrupted mid-response | Can interrupt and be interrupted instantly |
| Backchanneling | Cannot nod or say "mm-hm" while you talk | Can backchannel, nod, and react during your speech |
Safety Challenges and Release Strategy
Tavus explicitly acknowledges the responsibility: the same properties that make Griffin powerful also allow it to deceive humans. CEO Hassaan Raza and Head of Research Ioannis Patras stated:
- Working on safe disclosure features
- Collaborating with organizations tackling AI safety
- Griffin-Lite will not be available for customers until safety concerns are addressed
- Inviting external participation in evaluations and testing
- A more powerful model will follow "very soon" after safety work is complete
The existing Tavus platform serving 150,000 developers and businesses (Phoenix, Raven, Sparrow) is unaffected; Griffin is not yet on the platform.
Industry Impact
Griffin's breakthrough has far-reaching implications:
- Education: AI tutors could notice confusion before a student says anything
- Customer support: AI can see a faulty product via camera and guide repairs in real time
- Interview practice: Simulate real interviewer nonverbal reactions
- Language learning: True conversational immersion practice
But 48% misidentification also means: a familiar face on screen can no longer be treated as proof the person is real. AI impersonation risk has become more tangible after this milestone.
Conclusion
Tavus Griffin marks the crossing point where AI video interaction transitions from "mechanical response" to "human-level conversation." The full-duplex video-to-video architecture, dual-track #1 on NVIDIA VideoFDB, and 48% video Turing test pass rate collectively demonstrate that AI is approaching human face-to-face communication capability faster than expected. But safety teams are still working to ensure this technology won't be misused. For developers, this is a significant signal that AI agents are moving from text to multimodal real-time interaction.
Topic hub
YouTube AI Content Policy Hub
Answer-style evergreen hub for AI labels, auto detection, and disclosure—not just breaking news.
Explore YouTube AI Content Policy Hub →Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
ChatGPT ads + content distribution
Sell compliance checklists and automated distribution
OpenClaw Agent short-video growth
Lean into hybrid workflows as labels get stricter
Related tutorials
- How to build an AI content automated distribution system with n8n + ChatGPT: a complete 30-minute tutorial
- Claude Code automated writing practice: build an AI content production pipeline in 30 minutes
- AI Agent drives automated website operations: Build a fully automatic content pipeline in 30 minutes
- AI Agent-Driven Content Automation: n8n MCP Building Guide from Scratch
Related news
- Hikvision Guanlan AI Applications: From Vision Recognition to Industry Agents
- Cloudflare Open-Sources Clef Decision Models: Returns Probabilities Not Text, How to Cut AI Agent Costs
- DeepSeek Harness v0.2 Desktop Launch: Zero-Config AI Coding Agent, How to Choose?
- ElevenLabs Doubles to $22B Valuation with Eleven v4 Turbo: How to Choose Voice AI Agents?