AI Briefing
KO

Interaction Models

·2026.05.12 09:00

Key point

Thinking Machines released a research preview of Interaction Models for real-time multimodal interaction.

Details

Thinking Machines released a research preview of Interaction Models. The core idea is to build interaction into the model itself rather than bolting it on with an external harness, letting humans and AI continuously carry on conversation, listening, visual input, and tool use at the same time.

The company sees existing turn-based interfaces as blocking the bandwidth of collaboration. So the model continuously interleaves input and output in 200ms micro-turns, incorporating silence, overlap, and interruption into context without a separate turn-boundary detection step.

The architecture is built around two axes.

  • interaction model: handles real-time responses, conversation management, and visual/language mid-course intervention.
  • background model: processes deeper reasoning, tool calls, and long-running tasks asynchronously.

Example capabilities disclosed include:

  • Simultaneous speech and real-time translation
  • Tool calls and web search while speaking/listening
  • Time awareness and immediate reactions to visual cues
  • Mixing generative UI into the conversation flow in real time

For input processing, the company chose encoder-free early fusion using dMel audio, 40x40 patch images, and a flow head audio decoder, and explains that all components were trained together from scratch. The company claims this approach scales intelligence and responsiveness together.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.