HNHacker News
TopNewBestAskShowJobs

Casteil

713 karma · joined October 27, 2016

submissionscomments
Casteil··on Qwen3.8-Flash-Next
Yep.. for 'general purpose' use I found qwen3.8:27b to be disappointing due to overthinking. It's brutal especially considering how slow it is compared to MoE variants. It often overthinks to the magnitude of ~10x the tokens vs a ~4x faster gemma4:26b-a3b.

As a result, qwen3.8 will churn over a prompt often for 5-10 minutes while gemma4 regularly finishes the same prompt in under 20 seconds, while giving a consistent and accurate response in my favorite test case. Qwen3.8, despite churning like that, often misses with an inaccurate answer.

Obviously, 'YMMV' depending on your use case... just sharing my two cents.

Casteil··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
It's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them
Casteil··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf.

qwen3.5:122b-a10b is significantly faster at around 60-65.

Casteil··on Firefox is now the last major browser that still supports uBlock Origin
Now? Sure seems like it's been that way for well over a decade.
Casteil··on Qwen 3.8 27B
Yep.. it's pretty obnoxious for real-world use with the default 'xhigh' thinking. Ridiculous amount of "Wait, actually.." which might help for complex coding tasks but makes it unbearable for general purpose use.
Casteil··on Qwen 3.8 27B
Given that it apparently defaults to 'xhigh', this is probably the answer.

Granted, it's still much lower tokens/s than you'll get out of many MoE models.

Edit: Even set to medium or low there's still a lot of second guessing, less consistency, lower 'acceptable response' rate, and slower/more token churn vs gemma4:26b-a3b. I think gemma4 is just a better 'general purpose' model.

Casteil··on Qwen 3.8 27B
Yeah, that's probably the answer given that it apparently defaults to 'xhigh'.
Casteil··on Qwen 3.8 27B
One thing a lot of people don't seem to factor when hyping Qwen is how much models like this tend to 'overthink' with seemingly endless 'second guessing'. 3.8 seems no different from what I've tried thus far.

As capable as it is, it's hard to justify using it when a competing model (e.g. Gemma4:26b-a3b) can consistently achieve the same or similar response with only 1/10th as many 'thinking' tokens, achieve much higher tokens/second, and take a small fraction of the time. I suppose 'YMMV' depending on your use case.

Also, I haven't used it enough yet to see if it's prone to infinite looping, but its predecessors sure were.

Casteil··on Qwen 3.8 27B
I'm hoping too that they'll put out some MoE variants.

Qwen3.5:122b:a10b can run about twice as fast as this 27b dense model.

Edit: Like its predecessors, 3.8 seems really inclined to overthinking, and on a 27b dense model that's kind of painful. I think I'm going to stick with gemma4:26b-a3b as my go-to because it runs about 4x as fast and tends to only need a fraction of the tokens in its 'thinking' stage to get the same or similar answer.

Casteil··on I stopped trusting USB-C cable labels and started testing them
I just got this one a few days ago so it's kinda weird I've seen it mentioned or posted about multiple times since. Guerrilla marketing?

Anyway, it's definitely better than having nothing, and tests resistance as well. I tossed a couple low value/duplicate/high resistance cables because I've got too many just lying around causing clutter. Also helped 'confirm' my good (e.g. TB4+) cables that weren't labeled clearly are indeed good cables - granted, this is only as dependable as the marker chip the manufacturer added.

Casteil··on So Reddit has decided that plain HTML is unsafe
They're definitely doing it based on IPs. I just hopped to a different VPN endpoint and it works for me again (same browser). For now.
Casteil··on So Reddit has decided that plain HTML is unsafe
Doesn't work for me. Seems like this login gate for old.reddit is being progressively rolled out.
Casteil··on Most Americans say "not in my backyard" to AI data centers
The difference in noise pollution comes from power density & power generation, if they're doing so on site. The power demand (and resulting waste heat) absolutely eclipses that of data centers of olde.

If they're not generating power on site & creating noise/consuming local natural gas resources, they're guzzling grid power and almost certainly contributing to electrical rate increases that affect residential customers disproportionately.

Casteil··on Most Americans say "not in my backyard" to AI data centers
There's more to it than that. They're not just obnoxious to exist in the vicinity of, they also drive up local energy prices.
Casteil··on Most Americans say "not in my backyard" to AI data centers
>No one ever cared about DCs before now.

Fewer people cared, sure.. but it's because until recently they weren't eclipsing power consumption of extremely energy-heavy industries, noticeably driving local residential energy prices up.

They also weren't so heavily noise polluting nor were they typically built in places where their noise pollution affected so many people.

Casteil··on Running local models on an M4 with 24GB memory
You can expect around 55-60t/s with Qwen3.5:35b-a3b or gemma4:26b-a4b Q4
Casteil··on Running local models on an M4 with 24GB memory
Qwen3.5/3.6 are really prone to looping and 'overthinking'. Gemma4 doesn't seem to have the same problems.
Casteil··on Running local models on an M4 with 24GB memory
Why not 35b-a3b? ...or gemma4:26b-a4b? Both will be more capable than 9b and run at roughly similar (perhaps faster) speeds
Casteil··on Ollama is now powered by MLX on Apple Silicon in preview
It's gotten significantly better with the advent of local/offline MoE models (e.g. qwen3.5:35b-a3b, qwen3:30b-a3b, gpt-oss:20b-3.6b), which offer a good balance of prompt response speed and output quality.

'Dense' models of yesteryear (e.g. llama:70b, gemma2/3:27b) tend to be significantly slower by comparison, therefore, your hardware spends a lot more time 'maxed out' for a given prompt.

Casteil··on Hetzner Prices increase 30-40%
OVH is nearly doubling their some of their VPS pricing soon.
Casteil··on Microsoft Copilot AI Comes to LG TVs, and Can't Be Deleted
100%. Roku's privacy policy is the most wildly invasive thing I've ever seen - basically everything that used to be just conspiracy theory.
Casteil··on X5.1 solar flare, G4 geomagnetic storm watch
Nice. Looks like it was peaking around 21:00 - 22:00 local time, got pretty intense for a while.
Casteil··on GM will ditch Apple CarPlay and Android Auto on all its cars, not just EVs
Yep.

Manufacturers ditching Apple Carplay/Android Auto support will, if not immediately, inevitably pursue rent-seeking behavior in the form of paid subscriptions for services people could otherwise just have for free (and likely better) via phone.

Casteil··on M5 MacBook Pro
...they're not. This is a release of a 14" with the base M5, alongside the other existing M4 Pro/Max models.

The Pro/Max rollout tends to lag behind by about 6 months.

Casteil··on M5 MacBook Pro
Good chance they'll introduce it with the upcoming M5 Pro/Max; the non-pro/max devices always tended to be a little lower spec all around.
Casteil··on M5 MacBook Pro
Pro/Max rollout tends to lag behind the 'base' by about 6 months

It used to be a little less 'weird' when the base M-chips were only available in the Air and 13" MBP.

Casteil··on M5 MacBook Pro
Correct. Sometimes much later.

Still no M4 Ultra Studio available.

Casteil··on M5 MacBook Pro
I wish they'd bring Space Gray back. Not a huge fan of Silver, and the 'Space Black' apparently tends to show smudges more.
Casteil··on M5 MacBook Pro
The non-pro/max chipped MBPs have always been a little 'lower spec' in several regards. There used to be a little more separation though, with the non-pro chips available only in the Air & 13" MBP, but back then people complained about Apple having 'too many models'...

I suspect the M5 Pro/Max chipped MBPs will bring some of these improvements you're looking for.

Casteil··on M5 MacBook Pro
Part of me misses my OG base 14" M4 Pro. The battery on that thing was absolutely phenomenal - literal 12-14+ hours of real-world use. Not so much on the 14" M1 Max (64GB) that I upgraded to after about 2 yrs.

'Real-world idle' efficiency on the newer chips is the main reason I've got the (slight) itch to upgrade, but 64GB+ MBPs certainly don't come cheap.

Page 1 of 7Next →