2,400 karma · joined March 17, 2014
In other words, I have a gun that shoots bullets. It's up to me to use it responsibly and legally.
Ok I'll be that guy. It's pretty easy to figure out if you're talking to an LLM now we know it's tics, failure modes, jailbreak techniques etc
Which like any company will be entirely driven by legal constraints and money. Or just money if it's cheaper to break the law for profit and pay fines. There will be no "this is what's good for humanity, economics be damned".
Much like HR isn't to help employees but just the company.
The bigger problem with his argument IMO is that even with open models, it's still the person or organisation with the most money / access to GPU compute winning. They run the larger model (or collection of models), they can process more tokens through them in the same amount of time etc
But
> What you get is a beautiful animation that is 100% accurate and free of hallucinations
100% free of hallucinations when you're not an expert that can check it is impossible. LLM hallucinations are an unsolved problem.
I don't think this is the win you think it is. It's amazing that this is possible, but it introduces so much human overhead that you can drown in reviews and it can effectively slow you down more than a quick check and fix yourself.
The models need to get a lot more consistent in what they can and can't do before you can automate this stuff and only check the things you know the model isn't good at
He literally says it's an impressive feat in the second article.
There's one in particular that I use quite often and have for about a year, vibed for myself: it's a chat interface that walks you through processing an emotional or difficult moment, following a process / workflow. Supports Martin Seligman's ABCDE and Byron Katie's "The Work". Two techniques that I found most useful to improve my thinking and responses to difficult situations. You converse with it, and it leads you through the stages of the selected flow. More than a system prompt - it tracks and progresses through the workflow at the right times, and you download a consistent pdf of key details from the "session" at the end for your records.
The thing is, I can't sell it. I don't even think I can open source it. It's mental health (minefield of legal and ethical issues), and would probably breach copyright, trademark etc law.
Yet it's very useful to me. I guess it's the equivalent of making a system to do these things at home, like prompt cards or key points taken from the books, with a well formatted note taking structure. Except it's way easier because you just chat like you're talking to someone.
I would never have spent time building this. But now it's so cheap and easy to do.
Would love to know more about how this works then? Is it more or less encrypted P2P?
But you still need to properly review plans and PRs to keep a good mental model of the codebase. This effectively limits the number of tasks being done in parallel to maybe 2-3. Though you'll be mentally exhausted and probably start to make mistakes or take shortcuts in reviews yourself.
Fable 5 sits ahead of Opus 4.7, but behind Opus 4.6, Sonnet 4.6, Opus 4.8, GPT-5.4, GPT-5.5.
Fable isn't a good coding workhorse. That doesn't mean it's not good for actually complex problems and long horizon tasks (big POCs, complex research and such). But I only have vibes and Anthropics own benchmarks and marketing to guide me there.
Do they? I saw some crazy stat from the guy who built claude code that he was pushing hundreds of PRs a day. There's no way you can human review that much code. It's probably closer to heavily AI assisted review and planning.