For context, the author is Steven Johnson, one of the key people behind Google's latest hit, NotebookLM.
For those who are curious, how can we technically support really long context window (like in the millions or even billions)? The short answer is simple: we can just use more GPUs. The long answer is detailed in my recent note here: https://neuralblog.github.io/scaling-up-self-attention-infer...