Michael Spencer published an intriguing post this morning about “Interaction Models,” an AI development from Thinking Machines Lab. Spencer explains that Interaction Models is a step beyond the awkward “turn-based” exchanges that still define most AI use — user starts communication, AI waits for communication, AI responds, user responds — and represents something much more fluid. As this video illustrates, I can listen, see, speak, process, interrupt, wait, and respond in something closer to real time:
This strikes me as much more than a user-interface improvement.
For education, it signals a breakthrough in learning with AI that hinges on an enhanced interaction model.
When we imagine a student learning with AI we typically picture the student asking a question, and then the AI answering the question. The interaction still largely revolves around content delivery. The AI may be faster and more personalized than, say, a textbook or video, but the basic structure remains the same. The student seeks information. The machine provides it.
But what if the AI is not simply waiting for a question?
But what if the AI can recognize hesitation, self-correction, confusion, or partial understanding? What if it can ask a clarifying question before a misconception solidifies? Or, just be quiet because the student is thinking?
That’s a live interaction layer for learning. And it goes well beyond the prevailing “AI tutoring” concept.
Thinking Machines argument is that prevailing AI interactions force humans to adapt to the machine. Users stop, formulate a prompt, wait, receive a response, and then decide what to do next.
But that’s not how many human learning conversations work.
From Content Delivery to Process Observation
Say a student is learning Spanish. The student is telling a story in the past tense and says, “Ayer yo voy… fui… fui al parque con mi amigo.” A traditional chatbot might wait until the student finishes and then provide a grammatically correct version. A better AI tutor would explain the difference between voy and fui.
But a better interaction model notices the self-correction in the moment and says: “Good. You corrected yourself from present to preterite. Now try the next example and pay attention to the same pattern.”
That might seem like small difference, but it is pedagogically important. The AI is not simply correcting user output. It is addressing a learning process.
In math, a student might talk through a solution while writing steps on a screen. The AI could notice that the spoken reasoning is sound but the written notation is inconsistent. It then focuses on this inconsistency.
In science, a student might describe a lab result while pointing to the wrong part of a graph. The AI points this out and asks why.
In history, a student might be describing the importance of a cultural site. The AI immediately produces visual evidence that supports, or challenges, the student’s assertions.
In writing, a student might verbally explain an argument, while the AI targets claims and gaps.
In each case, the educational importance is not just that AI can provide an answer. The value is that AI can help make student thinking visible.
Summative assessments by themselves often lead teachers astry. A finished essay can hide the process that produced it. A completed math answer can hide the misconception. A polished AI-assisted response can hide whether the student understood anything at all.
Interaction Models could make the learning process more observable. They could explore a hesitation, a revision, a spoken explanation, a visual reference, and a moment of self-correction. That is far more instructionally useful than another polished answer.
The Giants are Moving: Google’s Multimodal Bet
Google may not be using the same terminology as Thinking Machines, and I wouldn’t claim that it is building the same kind of model. But Google’s recent education work seems to point in a similar direction. LearnLM, Guided Learning, Gemini Live, Project Astra, NotebookLM, and Google Classroom integrations all suggest that Google’s tutoring bet is not simply about creating a chatbot that explains things better. It is about embedding AI more deeply into the flow of learning.
Guided Learning already nudges Gemini away from quick answers and toward questioning, step-by-step support, diagrams, quizzes, and multimodal explanation. Project Astra points toward live AI that can see what the learner sees, respond to the surrounding context, and interact through voice and visual input. NotebookLM is becoming more of a representation layer, able to turn source material into slides, reports, study guides, audio, video, and visual outputs. Google Classroom integrations suggest that these capabilities may increasingly sit inside the environment where teachers already assign, organize, and assess student work.
And Google has “realized that as new capabilities in AI are emerging and also maturing — for example — we have these demos of live experience where it’s kind of video and audio and you essentially can just talk to AI in the same way as you would talk to your human teacher,” says Irina Jurenka, the research lead for AI in education at Google DeepMind.
Put these pieces together and it certainly points to enhanced interaction models with layers distributed across its learning ecosystem.
Amplification, not Replacement
It also changes the role of the teacher. If AI can support more simultaneous practice, teachers may spend less time being the only source of immediate feedback and more time designing the conditions under which feedback happens. Teachers will need to decide: What should the AI notice? When should it interrupt? What counts as productive struggle? Should feedback focus on accuracy, reasoning, confidence, fluency, creativity, or revision? What should be visible to the teacher? What should remain private to the learner?
Imagine a classroom where twenty students are practicing oral explanations using visual aids with an AI partner. The teacher cannot listen to and watch every interaction in real time. But the system might reveal that many students are avoiding evidence, several are misusing a key term, a few are using visual aids misleadingly, and some are simply winging it. The teacher can then respond to the class or redesign the next activity.
That’s not AI replacing the teachers. That’s AI amplifying the teacher.
Yes, the risks are real — and tricky. More human-like AI may also become more persuasive, more emotionally attractive, and harder for students to evaluate. If an AI can interrupt, encourage, redirect, and challenge, then schools need clear boundaries. Who controls the interaction? What data is captured? Are voice, facial, or behavioral signals being analyzed? Can students pause the AI? Can teachers inspect the interaction?
The more natural the AI feels, the more important these questions become.
In all, Interaction Models are not some utopian ideal, nor to be feared. It’s all part of a significant shift. AI in education is moving beyond the chatbot box. The future is more than smarter LLMs. It is better interactions.
Find AI-in-education tools, news, how-tos, consults, and more at tomdaccord.com
.


Thank you for sharing yes I think the impact on education and Ed Tech is going to be considerable with these sorts of innovations in micro tweaks to the interface.