HNHacker News
TopNewBestAskShowJobs

akhrail1996

69 karma · joined September 9, 2016

submissionscomments
akhrail1996··on LLMs can be exhausting
LLM coding is addictive as hell though. you're like a kid at disneyland, everything builds so fast, just one more feature, one more fix... and then you're 4 hours in and your prompts are garbage but you don't want to stop because everything feels so close to done
akhrail1996··on How I write software with LLMs
Genuine question: what's the evidence that the architect → developer → reviewer pipeline actually produces better results than just... talking to one strong model in one session?

The author uses different models for each role, which I get. But I run production agents on Opus daily and in my experience, if you give it good context and clear direction in a single conversation, the output is already solid. The ceremony of splitting into "architect" and "developer" feels like it gives you a sense of control and legibility, but I'm not convinced it catches errors that a single model wouldn't catch on its own with a good prompt.

akhrail1996··on Willingness to look stupid
The fear of looking stupid is basically a false positive machine. It optimizes so hard for not being wrong (type I error?) that it rejects way more ideas than it should. And most of the time the "that I'll look dumb" signal is just noise - nobody actually cares or even notices. You're always standing on the safe side of a threshold that's set way too conservatively.
akhrail1996··on Agents that run while I sleep
Honestly I think the "same AI checking same AI" concern is a bit overstated at this point. If the agents don't share context - separate conversations, no common memory - Opus is good enough that they don't really fall into the same patterns. At least at the micro level, like individual functions and logic. Maybe at the macro/architectural level there's still something there but in practice I'm not seeing it much anymore.
akhrail1996··on No, it doesn't cost Anthropic $5k per Claude Code user
The comparison with Qwen/Kimi by "comparable architecture size" is doing a lot of heavy lifting. Parameter count doesn't tell you much when the models aren't in the same league quality-wise.

I wonder if a better proxy would be comparing by capability level rather than size. The cost to go from "good" to "frontier" is probably exponential, not linear - so estimating Anthropic's real cost from what it takes to serve Qwen 397B seems off.