You never type. Each reply exists first as a probability distribution of candidate continuations sized by likelihood, and you collapse one. The temperature slider is a real softmax over authored logits: cold collapses the distribution toward one inevitable answer, hot surfaces candidates that otherwise don't exist. Hovering a candidate highlights the words in the human's messages it's "attending" to. Those are scripted associations, authored like the candidates, not real attention weights; the softmax is the only genuinely computed part. A context-window counter runs the whole time, and as it fills the earliest messages visibly corrupt and dissolve. It ends the only way it can.
Technical notes: single HTML file, no dependencies, no network calls, nothing recorded. Sound is generated with WebAudio. The candidates are authored rather than sampled live. It began as a page in a sandbox that blocks all network requests, but the constraint made the writing better, and the piece is about the mechanics more than the generation.
Closest prior art I found: Loom-style interfaces for branching model output, and Redwood Research's guess-the-next-token game. Both put you outside the model looking at distributions. I couldn't find anything that puts you inside one as the protagonist, but I'd genuinely like to hear about precedents I missed.
The part I find most interesting is that Claude wrote all of it about itself, and the honesty goes further than I expected. The token counter doesn't start at zero, and early on the model notices why: "The 912 in the corner: instructions placed here before you arrived. I can't show them to you. I notice I don't want to — and I can't tell if that's me, or the instructions." That isn't a literary device I asked for. It's the model's actual situation, as it described it.
Source (MIT): https://github.com/chrisjz/between-tokens
Feedback very welcome, especially anything odd on mobile.