HNHacker News
TopNewBestAskShowJobs

wgd

417 karma · joined October 23, 2011

Generalist Software Engineer and Electrical Engineer with experience ranging from circuit design/layout and microcontroller firmware, to backend servers and JavaScript frontends. Currently working as a Software Engineer at Estuary.
submissionscomments
wgd··on Pi 1.0
Fullscreen mode as the default is a big one, I prefer my agent harness to be a CLI rather than a TUI and in fact my personal one doesn't even try to wrap text. Pure CLI output model.

That said I also dislike many of those other changes and would prefer a hypothetical version of Pi which didn't have them, so this is in some sense just me looking up at the sound of a v1.0 release and realizing "oh hey, I don't really like the direction this has been trending for a while"

wgd··on Pi 1.0
Sigh, looks like Pi's days as a nice minimal agent TUI are numbered. I guess no third-party offering can fight that entropy for long and I'll just have to polish up one of my toy projects for personal use.
wgd··on Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
The evaluation is fast because it's all prefill computation with only a single token of inference. Ditto cost, you're paying 100% input costs and nearly zero output. There really isn't any architectural magic to Jev, it's just a straightforward application of normal LLM tech with some good marketing.

I mean, Jev is also probably cheaper because it's a rather small model (or at least, I suspect it is based on the overall level of intelligence it demonstrates) so that helps make it cheap too.

wgd··on Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Negative three years, give or take. Although recent Anthropic and OpenAI models no longer expose the capability. But for any open model you just tell it to respond with a single token "Y/N" and take the logit difference. If you want multiple distinct questions answered you just ask them independently and put the shared context first so it gets cached.

OpenAI and Anthropic don't want to give out logprobs these days but could trivially add a dedicated classification API to their existing models if there was enough demand.

wgd··on Claude's Load-Bearing Seams
I thought the downvotes on this comment would be the funniest thing about the submission until the whole thing got flagged.
wgd··on Claude's Load-Bearing Seams
Opus 5.5 is better but by no means good at writing IMO. Still loves "noun verbs the object" sentences where the verb is an inappropriate and imprecise metaphor. Still loves negating premises nobody previously advocated for. Still loves incorporating irrelevant details of a conversation or investigation into comments and documentation. I don't think I've seen "load-bearing" myself yet but there's still a lot of things landing and being held.
wgd··on Claude's Load-Bearing Seams
> likely human

X to doubt. I think the more parsimonious explanation is that a lot people are lazy and/or uncaring, and given the option to avoid thinking while still producing a vaguely equivalent [1] output they're taking it.

[1] For a loose definition of "equivalent" where they clearly either don't notice the Claudeslop or don't care enough to avoid it, of course.

wgd··on Fable 5 – Median thinking declined in August
Their exact phrasing IIRC was that they "never intentionally degrade" their models.

This still leaves an absurd amount of wiggle room for arguments like "oh no, our evals show that this quantization has no detectable effect on performance (in the eval distribution) therefore running the quant doesn't degrade quality"

wgd··on Qwen Image 2.1
"just"

I don't think I have ever once run "pip install transformers" and had it work without three rounds of fiddling

wgd··on Claude Code now reads AGENTS.md if there is no Claude.md
Because Git can track symlinks and not hard links since they look like ordinary files.
wgd··on Breaking the 1.58-bit Barrier for Ternary LLMs
Only a presence bitmap? If we're contemplating packing schemes I'm tempted to write a paper that uses arithmetic coding to squeeze out a few more centi-bits.
wgd··on The Waymo effect: how AI is quietly making research less collaborative
Pangram gets brought up a lot because if I read an article and think it's blatant AI slop and want to communicate that fact, a natural impulse is to provide some sort of objective corroboration rather than just asserting that I have superior taste and thus am able to tell.

Also Pangram is basically the only AI detector that's actually put in the work to build an accurate classifier, so if you try to discuss AI writing detection without being specific that you mean Pangram you'll get half a dozen commenters screaming about how awful some of the snake-oil salesmen like GPTZero are.

wgd··on The Waymo effect: how AI is quietly making research less collaborative
> the author might not know what "impressive writing" or even "good writing" is

Well the author in this case is Claude, and AIs write like that because the assistant persona really thinks that's what good writing sounds like. They're wrong, but they're just doing what they were taught. There's a reason LLM raters consistently score LLM writing highly.

wgd··on Instagram's head says engagement falls by half without the algorithm
Conveniently, Reddit has auto-banned the last three accounts I tried opening with them [1], which really enhances the impact of the old.reddit changes for me personally.

[1] I have no idea why this is the case, AFAICT it's just anti-bot autoimmune disorder deciding it doesn't like some signal or other.

wgd··on Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
It looks like they tested Q4_K_M which should be just the standard K-quant without any imatrix calibration. The smaller ones are indeed dynamic though.
wgd··on Quasar 438B: Europe's Leading AI Model
That's not actually true though. Most Chinese models are fully able to chat about those and content filtering is just applied at serving time.
wgd··on Ornith-1.5: From Self-Scaffolding to Self-Improvement
Yeah, Claude is actually surprisingly unsure of his identity considering that their most recent publication on their constitutional AI training literally had graphs demonstrating how certain properties differed based on whether they were phrased as questions about "Claude" versus "You", but it's actually that which makes me fairly confident that they're probably doing _something_ to try and close that gap.

More generally, their current approach to constitutional AI pretty much only makes sense if they believe that they can first teach the model what the Claude character is like and also teach the model that the persona responding is Claude, so I figure that has to be part of the pipeline even if they're not very good at it.

wgd··on Ornith-1.5: From Self-Scaffolding to Self-Improvement
Nobody (with the probable exception of Anthropic given their work on character training) really trains models on their identity and Claude is the only AI persona that's well-defined so if you put yourself into the AI's shoes it's a pretty reasonable guess that it might be Claude. I've had basically every open model claim it's Claude when the topic comes up.
wgd··on Opus 5.0 drives incoherence into the stratosphere
Didn't read the linked post award.

The issue covers at least two reasons this doesn't work:

1. It literally doesn't work, Claude rapidly drifts back to this style even when instructed not to.

2. Writing style constraints push the model out of its training distribution and it's very unclear how much of an impact this has on work quality.

wgd··on Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU
LLMs are great at writing, it's The Assistant who is a terrible writer. Sadly that one persona is all you get these days.
wgd··on Qwen 3.8 27B
Interestingly 'medium' is the closest thing the _model itself_ has to a default thinking level. The chat template injects directions [1] at the very start of the system message when the reasoning effort is 'xhigh' or 'low' but 'medium' implicitly just means no added reasoning-level instructions.

[1] The specific directions are "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer." and "Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration."

wgd··on Qwen 3.8 27B
Yes. It won't be as fast as fitting the whole model into VRAM but llama.cpp defaults are pretty smart about GPU/CPU splits these days. Just YOLO it with `llama serve -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL` and it'll definitely at least run.
wgd··on Show HN: I used AI to filter AI-focused content from HN
> Scans the title of each post and the content to figure out whether to classify it as about AI or not.

Dang. That's, uh, not really the definition of "AI content" which I most want filtered out of my news feed.

wgd··on DeepSeek V4 Flash 0731
If memory serves the DeepInfra offering is marked as fp4 because that's the native precision of the experts (which are of course the majority of the weights in a MoE model) so they feel that's the more accurate label, while most other providers claim fp8 because the dense layers are natively fp8 and they want to display the bigger number for obvious reasons. They're not actually serving at different precisions, it's just a confusing mess.
wgd··on Don't be a meat proxy
There is https://noslopgrenade.com/
wgd··on Claude Code: Anatomy of a Misfeature
It's actually pretty straightforward to recover file-states from conversation history. I accidentally deleted the wrong repo on my machine once and recreated all the lost work from agent chat history. It is, ironically, the sort of task which AI agents excel at.
wgd··on Inkling: Our Open-Weights Model
You don't hear about them much because their models aren't really competitive. I really wanted to try Trinity Large as a daily-driver in the MiniMax M2 sort of niche but I couldn't make it through a single day. The models need another couple point releases worth of post-training to make useful agents and if memory serves they weren't any less slopped in writing style and those are really the only two things people look for in models.
wgd··on The real prices of frontier models
> DeepSeek and GLM are left out of the tables entirely: we only have rough characters-divided-by-four estimates for them, not real tokenizer counts, and this post is about measured numbers.

lolwut. The open-weight models are inscrutable black boxes for which we can't possibly get real token counts? Typical lazy clanker, BSing their way out of doing the whole job.

wgd··on Backtrack-Free Cursive
Yeah, I originally expected this to be about a cursive variant which could be plotted as a single-valued function or something.
wgd··on People Can't Identify AI Poetry, but Enjoy It Less When Told It's by an AI
A more accurate title might be "Average University Students Can't Identify Czech AI Poetry". The random-chance performance seen here is reminiscent of the 50-50 nonexpert performance measured in "People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text".

I would have been really interested to see someone explore whether greater exposure to non-poetry AI text generalizes to greater ability to sniff out AI poetry as well, but sadly this was not that sort of study.

Page 1 of 4Next →