There is still some minor memory bandwidth issue on outputting more tokens, but the truth is that if you process e.g. 16 messages at once you wont end up being much slower than Jev even though you have to perform several autoregressive passes.
I'm quite interested in this; my current understanding is though that Jev is great when scored with response quality and latency metrics.
The prevalent idea of Jev's superiority in price, speed and accuracy seems to come from TypeSafe's marketing and their, I'd say even bad faith, benchmarking. In independent benchmarks the relative numbers tend to be very different.
Sure they are no longer memory bandwidth bound thanks to that but someone could add a similar projector to a conventional model, train with a Jev style dataset and call it a day.
Whatever they are doing on inputs must either mean they intentionally chose a Mamba successor or they suffer from the same compute costs as everyone else.
Maybe first request is unbatched, to have fast prefill, and the subsequent ones are batched.
They also don't restrict your prompt. You can have a dumb one, where you put the variable data at the front, and the details on how to process it at the back, thus you bust the user-part of the KV cache every request.