The model just handed the room a confident, fluent, wrong answer. Everyone saw it. Now you have about ten seconds before the moment sets, one way or another. Here is the move.
Hallucination is not a scandal in a learning room. It is a scheduled event. The question is not whether an LLM will confabulate during a live session, it is what the facilitator does the moment it does, because that response becomes the room's default posture toward the tool for the rest of the day.
The move, in four beats
1. Stop the read-aloud. The instant you see something you cannot verify, you stop narrating. Do not paraphrase to soften it, and do not scroll past. Silence in the room is fine. Silence is signal.
2. Name the class of the claim. Say out loud, in plain words, what kind of thing the tool just asserted. A date. A citation. A number. A cause. Naming the class tells the room where to look for the check, and it models the habit you want them building anyway.
3. Run the cheapest check together. Open the primary source in front of the room. Not a summary, not a second LLM, the actual filing, paper, or docket. If the check takes longer than two minutes, that itself is the lesson: you found a claim whose confidence in the output was wildly out of proportion to how easy it was to verify.
4. Log it, out loud. Add one line to the shared tool-use log the room can see. What was asked. What came back. What the primary source said. Whether anyone in the room would have caught it without stopping. That last question is the whole point.
Why this beats the alternatives
The two failure modes to avoid are the reflexive apology and the reflexive dismissal. The apology (Class E, cited later) reads as "sorry the tool is broken today," and it teaches the room that a fluent wrong answer is an accident rather than a structural property of a next-token predictor conditioned on training data (Class E, Parr, Pezzulo, and Friston, Active Inference: The Free Energy Principle in Mind, Brain, and Behavior, 2022, on generative models producing prediction with uncertainty that must be checked against evidence). The dismissal ("these tools are useless") teaches the room to stop bringing signals back at all, which is worse, because next time the same failure occurs without a facilitator in front of it, no one will stop.
The four-beat move sits between them. It treats the hallucination as expected, checkable, and instructive, without pretending the tool is either magic or garbage.
What the log entry actually looks like
The classroom tool-use log is the same artifact learners keep privately in their weekly practice, writ large on a shared board. One row per incident, filled in front of the group, kept after the session (Class C, the log is a plain text file wired into no analytics, so nothing else quietly reads what the room decided).
# Session log, YYYY-MM-DD ## Incident 1 Asked: "Summarize the 2024 court ruling on X." Got: A confident paragraph with a citation. Class of claim: legal citation, single case, docket number. Primary source check: docket does not exist as stated. Time to verify: 90 seconds. Would we have caught it without stopping? probably not.
Read that entry a week later and it teaches without a facilitator present. That is the intent.
The falsifier
Here is the honest test for the move (Class F). If a cohort runs the four beats through a full session and the shared log shows the room accepting later fluent-but-uncheckable outputs at the same rate they accepted them at the start, the move has not landed. In that case the correct response is to publish that finding, name what did not transfer, and change the sequence. What we have not yet done, and want to (Class U, unverified in our own data), is a matched comparison against sessions where the facilitator handles hallucinations privately after class rather than in the room. Until we run that, we hold the four-beat move as a strong candidate, not a settled result.
The frame this sits inside
The move works because it is a small enactment of the larger posture: our work is on the attainable path toward General Natural Intelligence, natural not artificial, a working hypothesis with growing, evidence-classed evidence, tested in the open. Do not take the claim on faith. Test the build, inspect the gates, and help us find where it fails. A room that watches its facilitator check a wrong answer against a primary source is a room that has been shown, in a very small way, how the whole program is supposed to behave.
Ten seconds. Four beats. One line in the log. That is the move.
Where to go next
- AI as partner, not replacement: the cluster this post belongs to, and the posture the four-beat move enacts.
- The tool-use log, a weekly practice: the private companion artifact to the shared classroom log described above.
- The IamHITL workshop: where the move is rehearsed, starting with low-stakes prompts and then moving to the ones people actually bring in from their work.
Evidence classes cited in this post: A (empirical-in-session, facilitator observations across cohorts to date), B (code and inspection, the log template and session script), C (configuration and integration, the log wires into nothing by design), E (Parr, Pezzulo, and Friston, 2022, on generative models and prediction), F (falsifier stated for the move), U (matched-cohort comparison not yet run in-house). Corrections welcome. Bring the counter-evidence and the record updates.