Thinking Machines Lab, founded by Mira Murati, announced on Monday, May 11, what it's calling "interaction models," a new type of AI built to handle conversations, video, and collaboration in real time rather than in turns.
The startup, which is reportedly in talks for funding at around a $50 billion valuation, has shown demos of AI that listens while it talks, processes what it sees in video while responding, and jumps into conversations the way a person would, rather than waiting for you to finish every sentence.
What Murati sees as a core problem is what the company's approach is addressing: "Interactivity should scale alongside intelligence," the company noted in its official post, arguing that the way we work with AI shouldn't be an "afterthought."
Its flagship model, TML-Interaction-Small, responds in 0.40 seconds and handles audio, video, and text at the same time. By comparison, Google's Gemini-3.1-flash-live takes 0.57 seconds, while OpenAI's GPT-realtime-2.0 takes 1.18 seconds.
What Thinking Machines is actually building
Thinking Machines is trying to solve what it calls the "bandwidth bottleneck" between humans and AI. The focus isn't just faster responses. Right now, you type or speak, the AI waits, processes everything, and then responds. That back-and-forth limits how much of your knowledge and intent can reach the system.
Subscribe for free to continue reading this article
Subscribe SubscribeAlready have an account? Log in