WayToClawEarn
Medium impactThe Verge + Thinking Machines Official

Mira Murati releases Interaction Models: AI real-time multi-modal collaboration

Thinking Machines Lab founded by former OpenAI CTO Mira Murati released the concept of Interaction Models, which allows AI to process audio, video and text input in real time, achieving true multi-modal continuous collaboration and breaking through the traditional polling dialogue model.

WayToClawEarn EditorialPublished May 12, 2026Updated Aug 8, 2026

Editorial review of public sources · AI-assisted drafting. How we work · Original source

Core conclusion

Thinking Machines Lab founded by former OpenAI CTO Mira Murati announced the concept of Interaction Models on May 11 - a new human-computer interaction paradigm that allows AI to process audio, video and text input in real-time and continuously, completely breaking the polling mode of traditional conversational AI.

Key Points

  • Published: May 11, 2026
  • Core concept: Interaction Models — end-to-end real-time multi-modal AI collaboration
  • Capability features: Continuously sense user behavior and respond without waiting for input to be completed
  • Open plan: A limited research preview "in the coming months", with a wider release later this year

Background and trigger events

Mira Murati founded Thinking Machines Labs after leaving OpenAI in February 2025. This high-profile AI startup has been conducting low-key research and development, and today it finally disclosed its first core technology direction.

According to the official blog, the core concept of Interaction Models is: **Today's AI models can only perceive the world in a single thread - before the user completes typing or speaking, the model is in a "blind" state; before the model completes generation, its perception is also in a frozen state. ** This is like "resolving major disagreements via email rather than face-to-face communication" - the bandwidth is too low.

How Interaction Models work

Interaction Models described by Thinking Machines have the following core capabilities:

  • Continuous Perception: The model receives audio, video and text input from the user in real time
  • Real-time response: Start inference and generation without waiting for input to complete
  • Multi-modal fusion: Processing voice intonation, facial expressions, body language and text content simultaneously
  • Bidirectional bandwidth breakthrough: upgraded from single-thread interaction to parallel information channel

Demo scene:

  1. Detect whether animal keywords appear in the story in real time and provide instant feedback
  2. Real-time voice translation (translate while speaking)
  3. Detect the user’s sitting posture and provide posture reminders

AI

Key Impact

DimensionsChangeWhat it means to usRecommended actions
Interaction paradigmFrom polling dialogue → real-time multi-modal flowThe content form will shift from "article/conversation" to "real-time collaboration flow"Plan the multi-modal content production process in advance
Tool capabilitiesAI Agent can see, listen, speak, and think at the same timeThe automated pipeline will be upgraded from "step-by-step execution" to "continuous sensing decision-making"Focus on multi-modal API integration solutions
Competitive landscapeThinking Machines challenges OpenAI/AnthropicConversational AI may be replaced by real-time collaborative AIBenchmark your own products and plan multi-modal interactive interfaces
Bandwidth breakthroughFrom single channel of text to multi-channel of audio, video and textImproved information transmission efficiency between users and AI by more than 10 timesReconstructed user interaction interface design

Adaptation suggestions

For content creators and automation operations teams, Interaction Models inspire:

  • The form of content production will change: The future is not "writing articles → publishing" but "real-time collaboration → outputting multi-format finished products"
  • Automated pipelines require multi-modal interfaces: Research real-time audio and video APIs and streaming solutions in advance
  • Seize the Early Window: Thinking Machines’ Open Preview is a great opportunity to explore new modes of interaction

Action List

  • Follow Thinking Machines’ Open Preview timeline
  • Study existing real-time multi-modal APIs (e.g. OpenAI Realtime API, Gemini Live)
  • Think ahead about the application scenarios of "continuous sensing AI" in content automation

Related extended information

Tool entries in the text

Tool names such as OpenAI, Claude, and Gemini that appear in the text will be automatically matched by the site to the tool entry library and displayed as hover-cards.

Internal link guidance

Want to learn how to build AI Agent automated workflow? Watch the tutorial: AI Agent Tools 2026 Complete Tutorial: 5 Tools to Build an Automated Pipeline in 30 Minutes

Want to know how Claude Code started a business in 48 hours and achieved a monthly income of $9,000? See the case: Claude Code 48 hours to start a business: one person + US$29 monthly fee, monthly income in 3 months $9,000

View source →

Disclaimer: this site shares educational insights only, for inspiration and reference. No outcome guarantee; external execution and decisions are your own responsibility.