HNHacker News
TopNewBestAskShowJobs

dcre

5,021 karma · joined January 3, 2013

https://crespo.business/
submissionscomments
dcre··on 1M context is now generally available for Opus 4.6 and Sonnet 4.6
One tip I have is that once you have the diff you want to fix, start a new session and have it work on the diff fresh. They’ve improved this, but it’s still the case that the farther you get into context window, the dumber and less focused the model gets. I learned this from the Claude Code team themselves, who have long advised starting over rather than trying to steer a conversation that has started down a wrong path.

I have heard from people who regularly push a session through multiple compactions. I don’t think this is a good idea. I virtually never do this — when I see context getting up to even 100k, I start making sure I have enough written to disk to type /new, pipe it the diff so far, and just say “keep going.” I learned recently that even essentials like the CLAUDE.md part of the prompt get diluted through compactions. You can write a hook to re-insert it but it's not done by default.

This fresh context thing is a big reason subagents might work where a single agent fails. It’s not just about parallelism: each subagent starts with a fresh context, and the parent agent only sees the result of whatever the subagent does — its own context also remains clean.

dcre··on I'm glad the Anthropic fight is happening now
They are changing the meaning of the original text. It does not say “workers.”
dcre··on Personal Computer by Perplexity
Sounds terrible. Seems like Perplexity is desperate to appear innovative but doesn’t know how.
dcre··on I'm glad the Anthropic fight is happening now
Do you not think a change in knowledge work similar in scale to the industrialization of agriculture would be significant?
dcre··on Helix: A post-modern text editor
Same!
dcre··on Just Send the Prompt
I think linking to the source material is great. But providing a little synthesis is a real value add!
dcre··on Just Send the Prompt
This is kind of dumb because with agentic tooling, the input to the model is whatever it looks at. Sure, I'll send you the prompt along with a design document and the entire transcript of a meeting. That'll be useful.
dcre··on We Will Not Be Divided
This is just not thinking clearly. There are bad things that are asymmetric in character, dramatically easier to do than to mitigate. There’s no antidote or vaccine to nuclear weapons.
dcre··on Cord: Coordinating Trees of AI Agents
Not exactly a surprise Claude did this out of the box with minimal prompting considering they’ve presumably been RLing the hell out of it for agent teams: https://code.claude.com/docs/en/agent-teams
dcre··on The Popper Principle
There is obviously a lot of space between the two extremes "every opinion is the author's" and "we shouldn't take seriously anything authors write".
dcre··on The Popper Principle
If two characters express contradictory ideas, which side is Plato's? And even when there is not a clear contradiction it is not at all straightforward to decide what is being claimed. It's not an encyclopedia. It is written to be interpreted.
dcre··on The Popper Principle
Comment approved by my wife, who is a Plato scholar. Your point that whether True Philosophers even exist is left open is the kind of problem she points out all the time in dogmatic interpretations. It sounds basic, but it's so important to keep in mind that just because a character says something (even if that character is Socrates), that doesn't mean it's the "view" of the dialogue. And you have to be careful to pin down exactly what is being claimed, as you point out with the conditional. Plato is a master (surely one of the greatest of all time) of creating a dynamic space to think in without settling the questions raised.
dcre··on Anthropic officially bans using subscription auth for third party use
It's the latter. It's the average use that matters. Though I suspect API margins are also probably higher than people think.
dcre··on Measuring AI agent autonomy in practice
Tokens per second are similar across Sonnet 4.5, Opus 4.5, and Opus 4.6. More importantly, normalizing for speed isn't enough anyway because smarter models can compensate for being slower by having to output fewer tokens to get the same result. The use of 99.9p duration is a considered choice on their part to get a holistic view across model, harness, task choice, user experience level, user trust, etc.
dcre··on Anthropic officially bans using subscription auth for third party use
Why do you think they're losing money on subscriptions?
dcre··on Fastest Front End Tooling for Humans and AI
Vite+ is not “this dude’s project”, it’s made by the team that makes all the tools discussed in this article.
dcre··on Fastest Front End Tooling for Humans and AI
I only use it for typechecking locally and in CI. I don’t have it generating code. Of course, what is generating my code is esbuild and soon Rolldown, so same issue maybe. If CVEs in tsgo’s deps are a big risk to run locally, I would say I have much bigger problems than that — a hundred programs I run on my machine have this problem.
dcre··on Fastest Front End Tooling for Humans and AI
I look at it and don't really have an issue with it. I have been using tsc, vite, eslint, and prettier for years. I am in the process of switching my projects to tsgo (which will soon be tsc anyway), oxlint, and oxfmt. It's not a big deal and it's well worth the 10x speed increase. It would be nice if there was one toolchain to rule them all, but that is just not the world we live in.
dcre··on Fastest Front End Tooling for Humans and AI
Bun and Vite are not really analogous. Bun includes features that overlap with Vite but Vite does a lot more. (It goes without saying that Bun also does things Vite doesn't do because Bun is a whole JS runtime.)
dcre··on Fastest Front End Tooling for Humans and AI
I like tsx for this, and it's actively maintained. The author may not know about it. https://github.com/privatenumber/tsx
dcre··on SkillsBench: Benchmarking how well agent skills work across diverse tasks
"Self-Generated Skills: No Skills provided, but the agent is prompted to generate relevant procedural knowledge before solving the task. This isolates the impact of LLMs’ latent domain knowledge"

This is a useful result, but it is important to note that this is not necessarily what people have in mind when they think of "LLMs generating skills." Having the LLM write down a skill representing the lessons from the struggle you just had to get something done is more typical (I hope) and quite different from what they're referring to.

I'm sure news outlets and popular social media accounts will use appropriate caution in reporting this, and nobody will misunderstand it.

dcre··on "Token anxiety", a slot machine by any other name
How are we still citing the (excellent) METR study in support of conclusions about productivity that its authors rightly insist[0] it does not support?

My paraphrase of their caveats:

- experts on their own open source proj are not representative of most software dev

- measuring time undervalues trading time for effort

- tools are noticeably better than they were a year ago when the study was conducted

- it really does take months of use to get the hang of it (or did then, less so now)

Before you respond to these points, please look at the full study’s treatment of the caveats! It’s fantastic, and it’s clear almost no one citing the study actually read it.

[0]: https://metr.org/blog/2025-07-10-early-2025-ai-experienced-o...

dcre··on Qwen3.5: Towards Native Multimodal Agents
Yeah, I see this in dark mode but not in light mode.
dcre··on Something Big Is Coming (Annotated by Ed Zitron) [pdf]
I think those are all reasonable worries, and many critics do a better job than Zitron of articulating them.
dcre··on Something Big Is Coming (Annotated by Ed Zitron) [pdf]
Do you think can they work for 5 minutes without guidance? Because that's something Ed said would not and could never happen, and the people who said it would were dupes and idiots.
dcre··on Something Big Is Coming (Annotated by Ed Zitron) [pdf]
I'm going with the pathologically incurious guy who is wrong in essentially every detail.
dcre··on Something Big Is Coming (Annotated by Ed Zitron) [pdf]
Everyone is free to make their own judgment about who is offering a genuine analysis that clarifies reality rather than obscuring it.
dcre··on Something Big Is Coming (Annotated by Ed Zitron) [pdf]
See edit. Tens of thousands of lines of borderline gibberish for the gullible.
dcre··on Something Big Is Coming (Annotated by Ed Zitron) [pdf]
It's a response to this: https://shumer.dev/something-big-is-happening

The post is silly, but I do not expect Zitron's commentary to be particularly illuminating as he is a charlatan himself. I could point to many examples, but here is a blog post I wrote about one case of him trying very hard to not understand a simple and familiar situation: https://crespo.business/posts/cost-of-inference/.

dcre··on Anthropic raises $30B in Series G funding at $380B post-money valuation
I can’t find any news stories mentioning a $500B target for Anthropic. You might be thinking of OpenAI’s October raise.

https://www.cnbc.com/2025/10/02/openai-share-sale-500-billio...

← PreviousPage 5 of 34Next →