https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.
I pressed Enter, and the response was instant.
> Generated in 0.037s • 14,205 tok/s
This is unbelievable.
and it gave a very reasonable answer in non-perceptible time.
I find myself getting caught up in the sheer speed of modern computing and networking. The fact I can play an online game with 10 other people is just insane.
The other thing that I think is really interesting about all of this, is that LLMs are already perforce behind the times with their knowledge cutoff, so adding an additional ~3 months for bake into silicon isn't such a huge deal, I think, for the ~10x more efficient and faster you get.
Things like this give me hope for a system that can be fully local and private, but also with the ability to be almost infinitely extendable with tools.
I’m still trying to figure out coding agents. I can’t even begin to imagine the things it would enable. Even the most mundane ideas like LLMs-in-HiFreq-trading have huge implications.
You need to be 4 orders of magnitude faster at least
This is crazy.
"LMS algorithm in bash"
Just barfed it up lol.
Amazing.
I just tried it too and 14,098 tokens in .05 seconds, I barely blinked and it was done. There was no typing at all appearing on the screen. It just showed up.
https://chatjimmy.ai/chats/01dc66a4-4b1b-4dea-bb5f-926855e37...
Here's a copy and paste prompt if somebody wants to just test it real quick to see what I saw:
Write a story about the fastest monkey who ever lived, his name is Jimmy and he is an AI superbot monkey that is part cyborg primate. He can travel through time and is psychic.
If we get to anywhere near this speed for the equivalent of the current models... I don't even know what to think about that future.
Once speed significantly increases I think we're going to see some interesting downstream effects. The three things I currently spend the most time waiting on are LLM API requests, Rust compile times, and nix derivations. As AI latency approaches zero I think we're going to start taking a hard look at whether slow-compiling languages are adding enough value over Golang, Typescript, or even dynamic languages to be worth the slowdown.
"I need a short, 4000 word essay on the the difference between star wars and Star Trek universes from the perspective of graduate level scientific work."
Blocking Fable for sure made it very politicl a lot sooner than i expected it to happen.
and because China already has massive problems of getting access, they are pushing it on hardware too like what Huawai did without EUV.
It seems China is already able to do DUV a lot sooner than others expected.
That's the media and in particular US KOLs of all sorts driving the wrong impression of China and other places. China and many other places for example have fast public transport that the US doesn't and can't even imagine today. They're not behind.
China's DUV still isn't that production grade (mass produce-able) so don't get that hyped up the wrong way (in a different direction).
The whole China-is-behind with tech and in particular semi wasn't that they can't. The truth is they spent decades in internal politics and corruption. That all got solved with the bans, so thank the bans! Jensen even said the bans were bad.
That's about how disrupting DSPs were to the industries they arose out of (over a very long time frame).
How would that disrupt the industry?
There is other in the space, Groq and Sambanova are both private companies attempting to develop their own technology.
Or, when we will start doing this, who's going to be able to do that in scale?
I'm seeing the TAALAS example, but it's only an 8B model, suggesting some real limitations parameter wise. And for 2.5kW?
The big AI labs won't do that unless they are forced to, as they want you to spend more money on the big, expensive, frontier models (so they can live up to their valuation), so it's more likely that you will see this on smaller open weights models.
High speed SRAM is where the $$$ is
I don't think it would be that difficult to manufacture compared to other process tech. HBM is really hard to do compared to other memory types.