…but not for all models, which is pretty annoying.
13,292 karma · joined May 31, 2012
Email me at josh (at) jgirvin (dot) com
https://jgirvin.com for my blog
@girvo on Twitter/Threads
…but not for all models, which is pretty annoying.
I’m lamenting the change because the thing the love about this job is being ripped away from me.
I’m still far more adapted to this than my coworkers though. Hell I have a GB10 box I run local models on for fun.
Me? I enjoy that stuff for what it is. Reviewing code is definitely not the truly enjoyable thing. Writing code, expressing my logic in code. That is enjoyable to me.
Ahead of its time, IMO. I wonder how Jolla is gong these days.
https://au.finance.yahoo.com/quote/NET/financials/
Interesting to look at, a decent example for sure.
I do imagine it'll change, but it hasn't yet.
This is already kind of the case: the big enterprises don't really want to touch the latest Chinese models. It's a real pain, personally, I want to use them at work!
Qwen 3.8 Flash Next (what I'm running basically entirely now) sees 30 / 35.0 / 45 tk/s for prose, analysis and code respectively for actual use (not short context benchmarking) with Pi. Thinking blocks are ~35tk/s or so.
The GB10 having so much compute is great for prefill too, 2000-3000/s for 14k to 64k token prompts (cold cache too) in the quick benchmark I did. 3500tk/s for warm cache which is nice :)
When I accidentally streamed my ngrams over the 2.5Gb/s network, it cut all the throughput down in half basically. Especially notable for the time-to-first-token, which is what clued me in that I'd messed up somehow!
For Qwen 3.8 27B, I got it up to a consistent 20tk-25tk/s but 27B thinks so much that it was honestly too painful: Flash Next is as smart, as useful, but much faster for real agentic dev usage IMO
Laguna S 2.1 saw similar numbers to Flash Next if I remember right, but their latest updates means it doesn't quite fit a GB10 128GB anymore at full context which is a shame.
Note: these are all NVFP4 quants (usually a dynamic one where some tensor layers are left at full precision though)
I'm so tempted to buy a second one...
This one!
I'd recommend pointing your agent at it (after installing sparkrun), and asking it to research the absolute latest in TP=1 Flash-Next - mine grabbed particular vLLM nightlies and mods to improve performance, and it was well worth it.
It’s an NVFP4 quant, but it fits, and is surprisingly capable.