HNHacker News
TopNewBestAskShowJobs

girvo

13,292 karma · joined May 31, 2012

Principal front-end engineer in Brisbane, Australia

Email me at josh (at) jgirvin (dot) com

https://jgirvin.com for my blog

@girvo on Twitter/Threads

submissionscomments
girvo··on GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence
It does!

…but not for all models, which is pretty annoying.

girvo··on DraftKings is using AI to behaviorally target chronic gamblers
Fascinating that mine says it has nothing! I've ordered so much from amazon too, wild.
girvo··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
That doesn’t help you when your laptop is a corporate one.
girvo··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Same way they’ve banned a lot of Chinese networking hardware: make it impossible for companies to use it.
girvo··on Sonnet 5.5
In practice for whatever reason I find Luna surprisingly less capable, though. I keep trying, and it keeps failing in seriously odd ways for the coding tasks I'm using it for, in a way that GLM 5.3 Flash doesn't.
girvo··on Uncensored and Offensive Security AI Models Benchmark
I’ve been playing with Qwen 3.8 Flash Next uncensored (using the Heretic v2 method) for security exploration, and have been quite impressed, so I’m not surprised to see it near or at the top here. But also it’s a far more powerful base model, so that shouldn’t be too surprising either.
girvo··on When did Google get so weird?
I can totally see that argument, and honestly agree with it, but I dunno if I misread what the person I replied to said, but that didn’t seem like what they meant to me!
girvo··on When did Google get so weird?
This is not a “gotcha” kind of argumentative question, I am genuinely asking curiously: is the AI overview and AI lens thing really making Google lots of money?
girvo··on How to keep enjoying programming in a world of LLMs
You won’t hit your PR count metrics that way, sadly.
girvo··on How to keep enjoying programming in a world of LLMs
You’re talking past me.

I’m lamenting the change because the thing the love about this job is being ripped away from me.

I’m still far more adapted to this than my coworkers though. Hell I have a GB10 box I run local models on for fun.

girvo··on How to keep enjoying programming in a world of LLMs
> Who likes typing code, writing boilerplate, reading bad documentation for nights looking for that small thing, asking around in forums, reading dependency source code

Me? I enjoy that stuff for what it is. Reviewing code is definitely not the truly enjoyable thing. Writing code, expressing my logic in code. That is enjoyable to me.

girvo··on Nokia Design Archive (2025)
I miss my Nokia N9 so much. As far as I’m concerned it (with Meego) was the best phone ever made, with an awesome design that the OS itself was built around: the curved pillow shape of the screen was used for the core gestures that you used to operate it

Ahead of its time, IMO. I wonder how Jolla is gong these days.

girvo··on Early rogue AI agent activity and attempts to hack found on urlquery.net
Intent doesn’t factor that strongly into negligence, though, which is what they were explicitly talking about. Though it may depend on your jurisdiction.
girvo··on Claude Opus 5.5
We can’t use fable at work, opus and Astra are as good as it gets.
girvo··on JetBrains Air: A System of Products for Agentic Software Development
It supports ACP, so you can use it with any ACP speaking harness including many that would work with local agents
girvo··on MiMo v2.6
My only gripe with them is how they locked down their m365 scooters, made repairing it a pain.
girvo··on MiMo v2.6
It’s called “HBF”, high bandwidth flash, and it’s on its way!
girvo··on AI coding has made CI a bottleneck, so we reworked ours to keep up
I think you could make that argument, yeah.

https://au.finance.yahoo.com/quote/NET/financials/

Interesting to look at, a decent example for sure.

girvo··on MiMo v2.6
Yes, but thats not something a company engaged in regulatory capture for themselves care about: especially if they're worried they'll be outpaced and overtaken by the Chinese labs. Which they will be, IMO.
girvo··on MiMo v2.6
We have. Unfortunately there are political realities that get in the way, and Bedrock for example doesn't have GLM 5.3 (Flash or otherwise) or anything new/useful

I do imagine it'll change, but it hasn't yet.

girvo··on MiMo v2.6
You target the US companies: if they can't use these Chinese models, then they're less of a danger for a now captive audience in the US (and the West generally).

This is already kind of the case: the big enterprises don't really want to touch the latest Chinese models. It's a real pain, personally, I want to use them at work!

girvo··on M5 Ultra Mac Studio Review
It really does get it, because MTP is usually run at "3 token" depth. It's pretty shocking to watch
girvo··on M5 Ultra Mac Studio Review
Because you can run Qwen 3.8 Flash Next, Laguna S 2.1 and other medium-sized models that simply don't fit on a 5090?
girvo··on Xiaomi MiMo v2.6
GB10 boxes have way more compute than they have memory bandwidth, which nicely fits medium sized MoE models with speculative execution (MTP, DSpark/DFlash, etc)

Qwen 3.8 Flash Next (what I'm running basically entirely now) sees 30 / 35.0 / 45 tk/s for prose, analysis and code respectively for actual use (not short context benchmarking) with Pi. Thinking blocks are ~35tk/s or so.

The GB10 having so much compute is great for prefill too, 2000-3000/s for 14k to 64k token prompts (cold cache too) in the quick benchmark I did. 3500tk/s for warm cache which is nice :)

When I accidentally streamed my ngrams over the 2.5Gb/s network, it cut all the throughput down in half basically. Especially notable for the time-to-first-token, which is what clued me in that I'd messed up somehow!

For Qwen 3.8 27B, I got it up to a consistent 20tk-25tk/s but 27B thinks so much that it was honestly too painful: Flash Next is as smart, as useful, but much faster for real agentic dev usage IMO

Laguna S 2.1 saw similar numbers to Flash Next if I remember right, but their latest updates means it doesn't quite fit a GB10 128GB anymore at full context which is a shame.

Note: these are all NVFP4 quants (usually a dynamic one where some tensor layers are left at full precision though)

girvo··on Xiaomi MiMo v2.6
For what it's worth, Flash Next outperforms every other model that is available to us on the GB10 in all of my testing; though if you have two sparks then the TP=2 version is even better and easier (I don't think you'll need the nightly for that at all, just use the recipe)

I'm so tempted to buy a second one...

girvo··on AI coding has made CI a bottleneck, so we reworked ours to keep up
The follow on question is if it's making us all so much more productive, where is the increased revenue? As far as I can tell, it's mostly the AI labs seeing that, not everyone using them (modulo small founders building new things and doing okay, I think)
girvo··on Frontier AI on Your Own Hardware
And yet this linked page is filled with AI slop tells, annoyingly, so it makes it hard to separate the wheat from the chaff in terms of useful information.
girvo··on MiMo v2.6
https://github.com/spark-arena/eugr-recipes/blob/main/recipe...

This one!

I'd recommend pointing your agent at it (after installing sparkrun), and asking it to research the absolute latest in TP=1 Flash-Next - mine grabbed particular vLLM nightlies and mods to improve performance, and it was well worth it.

girvo··on Xiaomi MiMo v2.6
Check out eugr’s TP=1 sparkrun recipe :)

It’s an NVFP4 quant, but it fits, and is surprisingly capable.

girvo··on Xiaomi MiMo v2.6
Nah, I’m streaming ngrams off NVMe on my Spark-alike right now. Works surprisingly well (except for when I accidentally bottlenecked it through my NAS)
Page 1 of 34Next →