HNHacker News
TopNewBestAskShowJobs

afro88

2,400 karma · joined March 17, 2014

submissionscomments
afro88··on Sonnet 5.5
Opus 5.5 in my experience outshines Fable 5.1 anyway. May as well have Opus do plan, breakdown and review, and Sonnet implement.
afro88··on Sonnet 5.5
I've been vibe coding a game and running multiple Opus 5.5 in parallel on Claude Code Cloud, 5x Max plan, and I'm yet to hit a session limit too. Not sure when I'd use Sonnet. Though it would be nice to switch back to Pro I guess
afro88··on Prompting Claude Opus 5.5
I don't get it either. Ditto for visual design capability. It's so far above Opus 5 and yet the announcement mentioned nothing about it.
afro88··on OpenAI bots meddled with multiple US Government agency sites
Agents are out of control by default, especially during a training run, which is why when they are productionised into cloud hosts or even local harnesses they have external guardrails in place

In other words, I have a gun that shoots bullets. It's up to me to use it responsibly and legally.

afro88··on GPT-6 Sol and Luna
I've found Sol to be an excellent orchestrator, with Astra the planner and Sol again the implementer.
afro88··on Understanding ChatGPT Work
It's not the techs fault. This person chose to use it this way. They could have also carved out 20 mins at the start of their workday to do the same thing
afro88··on Apple introduces M6 and M5 Ultra
> something that can beat a turing test w/o sweating

Ok I'll be that guy. It's pretty easy to figure out if you're talking to an LLM now we know it's tics, failure modes, jailbreak techniques etc

afro88··on Claudette: Make Claude stop talking like a BuzzFeed article
Where?
afro88··on OpenAI’s head of ethics leaves less than a year after joining
> I would expect an ethics team to build frameworks that can help train/eval the model that the company spend millions of dollars and months training is going to be aligned to the ethical stances the company chooses.

Which like any company will be entirely driven by legal constraints and money. Or just money if it's cheaper to break the law for profit and pay fines. There will be no "this is what's good for humanity, economics be damned".

Much like HR isn't to help employees but just the company.

afro88··on Mark Zuckerberg attacks 'closed' AI rivals as Meta returns to open models
I'm not a fan of Zuckerberg in the least, but one area a super intelligent lawyer would be fine at is being drowned in court filings and paperwork

The bigger problem with his argument IMO is that even with open models, it's still the person or organisation with the most money / access to GPU compute winning. They run the larger model (or collection of models), they can process more tokens through them in the same amount of time etc

afro88··on How I use LLMs to learn complex topics
I'm all for this and I'm keen to try it. I certainly don't want to take away from sharing another neat use case for learning.

But

> What you get is a beautiful animation that is 100% accurate and free of hallucinations

100% free of hallucinations when you're not an expert that can check it is impossible. LLM hallucinations are an unsolved problem.

afro88··on US Military's cyber command unit grapples with cluster of deaths by suicide
Well go on then...
afro88··on DeepSeek V4 Flash 0731
> The beauty of intelligence at this cost (even if it's not SOTA) is that it opens a whole bunch of new use cases. Test failure in CI? Have the bot automatically propose a fix, its cheap enough that you can discard it w/h issues. Test coverage too low? Auto generate tests on CI for every pull-requests! Monitoring server logs, continuous security audits and investigating every received exception now becomes possible.

I don't think this is the win you think it is. It's amazing that this is possible, but it introduces so much human overhead that you can drown in reviews and it can effectively slow you down more than a quick check and fix yourself.

The models need to get a lot more consistent in what they can and can't do before you can automate this stuff and only check the things you know the model isn't good at

afro88··on Stateless MCP has recaptured my interest
You can't do that with tools either. Skills are basically prompts - they're not analogous to tools or MCPs. I'm not sure what your point is
afro88··on Ten advances in mathematics and theoretical computer science
Gary doesn't argue it's hype though. He argues 2 things: other people are getting carried away with the result, and we don't know enough about how it was reached to know where it falls on the impressive scale.

He literally says it's an impressive feat in the second article.

afro88··on Software for One
I've got a few of these.

There's one in particular that I use quite often and have for about a year, vibed for myself: it's a chat interface that walks you through processing an emotional or difficult moment, following a process / workflow. Supports Martin Seligman's ABCDE and Byron Katie's "The Work". Two techniques that I found most useful to improve my thinking and responses to difficult situations. You converse with it, and it leads you through the stages of the selected flow. More than a system prompt - it tracks and progresses through the workflow at the right times, and you download a consistent pdf of key details from the "session" at the end for your records.

The thing is, I can't sell it. I don't even think I can open source it. It's mental health (minefield of legal and ethical issues), and would probably breach copyright, trademark etc law.

Yet it's very useful to me. I guess it's the equivalent of making a system to do these things at home, like prompt cards or key points taken from the books, with a well formatted note taking structure. Except it's way easier because you just chat like you're talking to someone.

I would never have spent time building this. But now it's so cheap and easy to do.

afro88··on Flint: A Visualization Language for the AI Era
Are they in 2026? I haven't had an issue with json and LLMs in a long while
afro88··on Claude Opus 5
IMO a much better test would be designs that aren't AI to begin with. Much more useful to see how well a model can html an image design without slopping it up
afro88··on Show HN: Bento - An entire PowerPoint in one HTML file (edit+view+data+collab)
> Only an encrypted blind relay to allow for shared editing. The relay doesn't see any of the data.

Would love to know more about how this works then? Is it more or less encrypted P2P?

afro88··on Terence McKenna's Mega Bad Trip (2025)
Is this a quote from a book? Beautifully written
afro88··on Codex Resets
When did that happen with Codex? I thought that was a Claude Code thing
afro88··on Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
It's not about figuring out if it's LLM written though. The style is hard to read and annoying. With the kind of sentences GP was talking about it's actually harder to get the substance.
afro88··on New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
Curious whether you were just bare asking it questions, or whether you provided it with lessons one by one with instruction that the lesson is the baseline truth etc
afro88··on Better Models: Worse Tools
This has been the case since the early days. Aider had a bunch of code to be very forgiving with formatting of tool calls (file editing in particular at first). It's just the nature of the beast. It surprises me that Pi doesn't have a lot of this kind of stuff built in too
afro88··on The short leash AI coding method for beating Fable
Maybe I'm too optimistic, but given appropriate skills and references (not just for writing but also reviewing) and intelligent use of subagents for isolated reviews and checks, you can lengthen the leash a bit.

But you still need to properly review plans and PRs to keep a good mental model of the codebase. This effectively limits the number of tasks being done in parallel to maybe 2-3. Though you'll be mentally exhausted and probably start to make mistakes or take shortcuts in reviews yourself.

afro88··on Claude Fable 5: mid-tier results on coding tasks
We selected PRs (real ones we merged over the 6 months prior) and have an "LLM as judge" score how close the AI generated code is to the PR. Same as how other benchmarks do it, but it's with tasks we actually do and code we have decided is actually up to scratch for us
afro88··on Claude Fable 5: mid-tier results on coding tasks
Similar result on our kotlin coding benchmark at work. It measures how close agents can get to a small mergable PR (according to my team). 20 tasks of varying difficulty, with 5 attempts each, LLM as judge to evaluate accuracy (same outcome and quality but allowing for acceptable variances).

Fable 5 sits ahead of Opus 4.7, but behind Opus 4.6, Sonnet 4.6, Opus 4.8, GPT-5.4, GPT-5.5.

Fable isn't a good coding workhorse. That doesn't mean it's not good for actually complex problems and long horizon tasks (big POCs, complex research and such). But I only have vibes and Anthropics own benchmarks and marketing to guide me there.

afro88··on AI is slowing down
I'd love to read about the predictions that have been wrong (genuinely)
afro88··on LLMs are eroding my software engineering career and I don't know what to do
I wonder if there's a way to include data that's so unique you can prove it was trained on and sue later
afro88··on LLMs are eroding my software engineering career and I don't know what to do
> The dynamic of agent codes human reviews does seem like the only sane one for the foreseeable future. Even Anthropic themselves still fall back to this.

Do they? I saw some crazy stat from the guy who built claude code that he was pushing hundreds of PRs a day. There's no way you can human review that much code. It's probably closer to heavily AI assisted review and planning.

Page 1 of 28Next →