WayToClawEarn
High impactTavus

Tavus Griffin Passes Video Turing Test: 48% Mistook AI for Human, How to Choose Real-Time Video AI

Tavus released Griffin, the first Human Interaction Model (HIM) to pass the real-time video Turing test. In a 54-person blind study, 48% believed they were talking to a real human. NVIDIA VideoFDB score: 3.83 vs human reference 3.92. Griffin uses a full-duplex video-to-video architecture unifying perception, conversation modeling, and audiovisual generation.

Edisen Lu · WayToClawEarnVia TavusPublished Oct 2, 2026

Reviewed from public sources · AI-assisted drafting under editorial oversight. How we work · Original source

Public-source compilation

Synthesized from public posts/docs. Prefer the original source for primary claims.

How we review content · Primary source · Tavus

TL;DR

On October 1, 2026, San Francisco AI company Tavus released Griffin—the world's first Human Interaction Model (HIM) and the first AI to pass the real-time video Turing test. In a 54-person blind study, 48% of participants believed they had spoken with a real human after a one-minute video call, compared to a 2.4% pass rate for the previous industry-leading system. Griffin-Lite is available as a research preview to select trusted testers, with a more powerful model to follow after safety work is completed.

The Video Turing Test: What 48% Means

The study recruited 54 participants from the US and Europe through an independent research platform. They were told they would have a one-minute video call with another participant about "what they were looking forward to this year"—but their partner was actually a PAL (Personal AI Link) powered by Griffin-Lite.

After the call, participants rated their partner on a 7-point scale across five dimensions:

  • Seeming natural: 5.4
  • Seeming trustworthy: 5.6
  • Would enjoy talking with again: 5.8 (held at 5.4 among those who said AI)
  • Partner was really listening: 5.5
  • Conversation flowed naturally: 4.9 (lowest of five axes)

Over half of participants said the possibility their partner wasn't real had not crossed their mind during the call. 26 people (48%) ultimately believed their partner was real. The previous industry-leading system (Phoenix 4.5 + Sparrow-2 + Raven-1) managed only a 2.4% pass rate (1 of 41 participants).

NVIDIA VideoFDB Benchmark: #1 on Both Tracks

NVIDIA independently built and scored VideoFDB, the leading industry benchmark for full-duplex audio-visual conversation.

Generation Track (tests whether the model produces the right behavior—fluency, affect matching, nonverbal behavior):

SystemScore (out of 5)
Human ground truth3.92
Griffin-Lite3.83
Gemini 2.5 + Anam (next-best)2.80

Griffin leads the next-best system by 1.03 points and is only 0.09 below human ground truth—more than 12x closer to human performance than any other system.

Perception Track (tests whether the model understands the moment—fluency, conversational flow, visual grounding):

SystemScore (out of 5)
Human ground truth4.20
Griffin-Lite3.73
MiniCPM-o 4.5 (strongest baseline, audio-only)3.44
Gemini 2.5 Flash Native3.17
OpenAI gpt-realtime (audio-only)2.97

Griffin is the only model evaluated on both tracks. Takeover-rate alignment was best in class on both tracks (62.8% generation, 73.8% perception).

Technical Architecture: Full-Duplex Video-to-Video

Griffin is not a traditional cascade system (ASR → LLM → TTS → face rendering). It is a unified video-to-video architecture with two engines:

1. Continuous Conversational Modeling Engine

  • Perceives audio and video simultaneously (not audio alone)
  • Makes conversational decisions at sub-second intervals (mini-turns), not once per turn
  • Decides when and how to respond—speak, hold, yield, backchannel, nod, or stay silent
  • Emits expressive control signals (emotional tone, stance, facial expression, gesture)
  • Can be interrupted, can interrupt, adjusts mid-response without losing context

2. Audio-Visual Generation Engine

Streaming Speech Generation:

  • Fast autoregressive diffusion transformer (VDiT) generates latent speech progressively
  • Tavec Codec: convolutional autoencoder mapping 48kHz audio to compact continuous latent—40 values per frame × 100 frames/second, no codebooks
  • Fully causal decoder emits audio packets as small as 10ms—speech begins before the sentence is finished
  • Clones a speaker's voice from ~10 seconds of audio
  • 1 minute of speech = 6,000 latent frames (vs 2.88 million raw audio samples)

Streaming Video Generation:

  • Few-step autoregressive diffusion generator: real-time 720p video in 320ms chunks
  • VAE compresses time 8x—one latent = 8 frames at 25fps = 320ms of video
  • Inputs: reference image + streaming audio + streaming controls (gesture, gaze, emotion)

Three-Stage Distillation:

  1. Distribution Matching Distillation: teacher distilled into few-step student (fast, not streaming)
  2. Teacher Forcing → Autoregressive: converted to generate one latent at a time
  3. Self-Forcing: trained on own generated history for long rollouts without drift

Cascade Systems vs. Griffin

FeatureTraditional CascadeGriffin
ArchitectureASR → LLM → TTS → face (sequential relay)Unified video-to-video, all concurrent
Turn-takingWaits for user to finish, then thinks, then talksNever stops thinking; decides every sub-second
Information lossEvery handoff discards info (tone, visual cues)Perception is continuous—audiovisual throughout
InterruptionPoor—cannot be interrupted mid-responseCan interrupt and be interrupted instantly
BackchannelingCannot nod or say "mm-hm" while you talkCan backchannel, nod, and react during your speech

Safety Challenges and Release Strategy

Tavus explicitly acknowledges the responsibility: the same properties that make Griffin powerful also allow it to deceive humans. CEO Hassaan Raza and Head of Research Ioannis Patras stated:

  • Working on safe disclosure features
  • Collaborating with organizations tackling AI safety
  • Griffin-Lite will not be available for customers until safety concerns are addressed
  • Inviting external participation in evaluations and testing
  • A more powerful model will follow "very soon" after safety work is complete

The existing Tavus platform serving 150,000 developers and businesses (Phoenix, Raven, Sparrow) is unaffected; Griffin is not yet on the platform.

Industry Impact

Griffin's breakthrough has far-reaching implications:

  • Education: AI tutors could notice confusion before a student says anything
  • Customer support: AI can see a faulty product via camera and guide repairs in real time
  • Interview practice: Simulate real interviewer nonverbal reactions
  • Language learning: True conversational immersion practice

But 48% misidentification also means: a familiar face on screen can no longer be treated as proof the person is real. AI impersonation risk has become more tangible after this milestone.

Conclusion

Tavus Griffin marks the crossing point where AI video interaction transitions from "mechanical response" to "human-level conversation." The full-duplex video-to-video architecture, dual-track #1 on NVIDIA VideoFDB, and 48% video Turing test pass rate collectively demonstrate that AI is approaching human face-to-face communication capability faster than expected. But safety teams are still working to ensure this technology won't be misused. For developers, this is a significant signal that AI agents are moving from text to multimodal real-time interaction.

TavusGriffinHuman Interaction Modelvideo Turing testfull-duplex AINVIDIA VideoFDBreal-time video AIAI safety

View source →

Educational reference only: cases summarize public sources and may use AI-assisted drafting under editorial review. Not financial advice; outcomes are not guaranteed.
Tavus Griffin Video Turing Test Passed: 48% Mistaken for Human, HIM Model Explained · WayToClawEarn