HNHacker News
TopNewBestAskShowJobs

dannyw

15,328 karma · joined August 29, 2017

Hi :)

Unless stated otherwise, opinions here are my own and personal.

submissionscomments
dannyw··on Singapore govt dating app uses Gale-Shapley stable marriage algorithm
GLP-1s, while being pretty miraculous as a whole, do come with risks like every other drug. There’s plenty of 1 in 10k or 100k side effects.

Kurtzerkag has a good video about it: https://youtu.be/TYhNHX372ek?si=5ESz3ykXykyUcmYg

dannyw··on Show HN: Ledge.sh – Runnable Markdown Notes
Love this concept, simple but brilliant idea with lots of utility. I think this can fit in a decent niche. I’ll give it a try!
dannyw··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Both labs are spying? Employees hang out at the same bars, have overlapping social circles, etc. Alcohol does what alcohol does.
dannyw··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Adaptive thinking was a good idea though. The old method of manually specifying how many thinking tokens you wanted as budget was just silly. The rollout might not have been great, but the change is good.

And, is it really even an oligopoly anymore? Open weight models are incredibly competitive in every way; whether you want to use US providers, Chinese official providers, self host, etc.

dannyw··on Livenerf: Has Opus 5.5 been nerfed yet?
Official tweet from Tibo from OpenAI confirming they ran experiments tweaking “juice” (reasoning effort mapping values) and have reverted them: https://x.com/thsottiaux/status/2076495156757577895

As he confirmed, for at least a period of time, and for some users, “Sol xhigh” was actually “Sol high”, etc.

dannyw··on Livenerf: Has Opus 5.5 been nerfed yet?
I’m also often not affected by Claude outages on the API; but Claude.ai is down.
dannyw··on Livenerf: Has Opus 5.5 been nerfed yet?
Super long chats are slow and sluggish too at least on chatgpt/claude.ai; even on a semi beefy machine.

I’m sure it’s lowkey intentional, probably encourages users to spin up new chats; hence less context.

dannyw··on Livenerf: Has Opus 5.5 been nerfed yet?
Serving LLMs at Anthropic scale is very, very different. It’s not SGLang or vLLM.

If you’ve tried setting either of these up, you’ll know how various tricky settings can impact throughout and model output quality; and those are much simpler stacks.

Even homelabbers are getting into disaggregated compute; e.g. one GPU for prefill, another for decode.

Obviously Anthropic and co are using a mixture of GPUs and hardware and clusters; not everything is just GB300 or whatever; so you then get into hardware quirks, kernel optimisations that may deliver huge speedups at the cost of a tiny bit of KL divergence, etc.

And I believe they’ve publicly said they use TPUs for inference too, but I doubt exclusively; and I’m sure that’s well optimised too.

Finally, Google has publicly stated they intentionally and silently degrade/poison models in response to distillation attacks; who knows what the other companies do.

dannyw··on Livenerf: Has Opus 5.5 been nerfed yet?
Inference isn’t flat 24x7, peak hours have more usage, but you buy/rent servers; not servers only for peak hours.

At their scale, you’d have to be setting money on fire if you’re not doing dynamic inference optimisations based on load.

API and consumer subscriptions are treated differently; all trackers measuring via API won’t notice this.

dannyw··on You are no longer invited to dinner
I think ADHD is a good parallel. I absolutely have it, I’m diagnosed and have had symptoms since childhood.

But even so, I often realise I self-justify not doing something and attribute it to my ADHD, excessively. It becomes a clutch, and I gravitate away from the willpower I do have.

That reminds me, it’s time to reach out to some friends and make weekend plans, even if I’d prefer having a quiet one. I’ve had too many recently :)

dannyw··on Everybody’s home. No one’s coming over
I think the examples you listed (doomscrolling, being sedentary, junk food) are related but distinct to ‘avoiding friction’. My perspective is it’s about quality, not the form; even if they correlate.

I don’t enjoy going to the gym, but I love swimming and hiking, and love that. I haven’t worked out in a gym for years, I don’t want to, and I’m quite healthy and reasonably active.

Some people are just introverts, and I think the ‘optimal’ amount of IRL time is different for every individual. Much like my gym vs swimming example; quality social interaction can take many forms; some of my closest and most meaningful friendships over decades have been based on games (voice calls; discord chats; hanging out in person sometimes and going on trips together sometimes), and that works for me; and I don’t think I’m missing anything.

Prioritise genuine human and social interaction, and for introverts, get out of the comfort zone once in a while (being a hermit is not good). But whatever form that takes is secondary to the quality and how it fits your personality and quirks.

dannyw··on Claude partial outage
For me, commits are the atomic block; and a good commit should be self-contained anyway; where no additional context is necessary. If an agent can't figure out what a commit is supposed to do, with only the commit title/description and diff, it's a bad commit, and this has always been true in software engineering. I do not pass prompts between the two ordinarily.

After claude or codex finishes a commit, I switch console tabs and ask the other to review it; with the commit ID. I often find it helpful to inject a bit of human knowledge, and callout any areas of attention I see from a quick skim. (But that could just be me wanting to not abstract myself away from software engineering that much :)

Sometimes, I do ask Codex to read my ~/.claude/; and vice-versa. But, generally, I try to keep as much knowledge (e.g. investigations, reports, deep dives) inside the git tree as possible; so that is not necessary.

I don't use skills, but I do have ~/.codex/AGENTS.md and ~/.claude/CLAUDE.md. These are high-level instructions for what (1) I consider readable, maintainable code and patterns, and (2) workarounds for empirically observed model behavior IOdon't like, such as Astra being a bit of a "over-correct over-validation nit-picker". I keep these human-authored, and update regularly based on what I find annoying.

I do all of this before I submit a PR; but of course, for trivial stuff (e.g. CSS changes, copy/string changes, etc), I don't bother.

My practices and workflows do change over time. Back in the ~Opus 4.5 days I'd often define a rubric/criteria in a markdown file, iterate with AI to improve it, and that's the "spec". I've stopped doing that since GPT 5.6; partly because models have gotten a lot better at understanding high level intent from the context; and partly because nearly all models these days feel 'gradermaxxed' when working like that.

Finally, consumer $100/$200mo subs get you _so_ far, I get a lot of value from both. I used to have multiple Claude subs for a while, but trying to 'get full value' made me work on projects just for the sake of it; so 2x$200/mo is my cap :)

dannyw··on Claude partial outage
There's far more than zero evidence, see https://arxiv.org/html/2402.08806v1 ; or the industry-standard practice of using multiple model families as LLM judges; or even 1P implementations https://code.claude.com/docs/en/advisor (which misses most of the benefit; since you want different model families, with different pretrains and posttrains).
dannyw··on 500k facial scans at UK stations yield no arrests, 1 false positive
The BBC of today is far from what it was a decade ago. So much trust and credibility has been lost.
dannyw··on Claude partial outage
Claude models don't go out globally either. Bedrock is fine. Azure is fine.
dannyw··on Claude partial outage
If you're not using Codex and Claude at the same time, you're missing out heaps. Even without outages, OpenAI models are great reviewers for Claude; and vice versa.

Other than trivial PRs; everything I do with Opus/Fable gets reviewed by Astra; and everything I do with Astra gets reviewed by Opus/Fable.

Using only models from a single vendor, is like testing your website/webapp only on Chrome.

dannyw··on Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
So it generates logits in a 255 token output space? ;)
dannyw··on The problem is not AI code, but not knowing about system architecture or intent
Our brains are hard-wired to enjoy (at least in the short term) immediate gratification, and the removal of friction. It's a hard battle to fight.
dannyw··on Claude Code reads AGENTS.md only when telemetry is on [fixed]
Do you use chrome? Check chrome://flags
dannyw··on Claude Code reads AGENTS.md only when telemetry is on [fixed]
Don’t make too quick snap judgements. “Honeymooning”, or giving new subscriptions / upgrades extra usage or “juice” is pretty common industry practice amongst SaaS “growth hacking” for years. Sadly.
dannyw··on Claude Code reads AGENTS.md only when telemetry is on [fixed]
It’s more likely that GP encountered some bug or corner case or weird experiment conflating than anything intentionally deceptive.

There’s also the fact that LLMs aren’t perfect, and sometimes even the best models act really stupid sometimes.

dannyw··on Claude Code reads AGENTS.md only when telemetry is on [fixed]
I’d just assume good intent here. Feature flags, telemetry, and fast rollouts / rollbacks are standard practice in software. Have a look at chrome://flags perhaps.

I fully believe GP that there was zero intent to gate this behind collecting telemetry. Sounds like a little tech debt and a little oversight, and the simplest explanation is that it is.

dannyw··on Claude Opus 5.5
Models are generally posttrained to a window shorter than you think.
dannyw··on Claude Opus 5.5
IIRC it only blocks kernel development for Huawei and other Chinese chips.

Fable and Opus, since 5.1 and 5, will happily hill climb on my CUDA kernels for transformers.

dannyw··on Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
There’s enough thinking leakage from the recent paper and just generally catching things on Reddit. Claude models overthink and self-doubt itself just as much as Qwen, but the summariser hides much of that.
dannyw··on GPT-6 Sol and Luna
If you’re using the API, both OpenAI and Anthropic models will happily update you on what it’s doing in significant and frequent detail with system prompting. You’re not getting raw/hidden thinking, but what you’re describing is more behavioural quirks of the harness and its system prompts.

The other explanation is just as part of ‘token efficiency’

dannyw··on Spymarks, Not Watermarks
Much prior art here, works against both Google and OpenAI's SynthID: https://github.com/0xROOTPLS/DeSynth
dannyw··on Raspberry Pi blocks changing RAM chips
That's quite disappointing. Whose hardware is it anyways? Yours, when you pay for it.
dannyw··on Qwen-Image-2.1: Compact, efficient, and unified image creation
Try translating your prompt to Chinese first, it seems a lot better at understanding and following Chinese prompts even with the translation hop.
dannyw··on Qwen Image 2.1
You really think Alibaba is going to go around and sue in US courts for something like this?

It’s more of something to scare companies with legal teams. If you’re an individual or hobbyist doing a side project the risk is essentially zero.

Page 1 of 34Next →