> Caveats, stated plainly. [from the Fable transcript pasted in the article]
I had a visceral reaction to these three words.
> Caveats, stated plainly. [from the Fable transcript pasted in the article]
I had a visceral reaction to these three words.
> I told it to look online at some of Fable’s strongest feats, especially the math problems it has solved, and that something like this should be easy in comparison.
Wait. Wait wait wait. Are we supposed to be giving them pep talks?
I have not seen this in other models.
The LLM likely needs to be reminded of its abilities.
Like when it tells you something is 3 days of work but it can do it with some degree of guidance in a couple hours
That's when you.. we.. all become the training data... o_o;
Modern AIs have very limited metaknowledge - they don't know exactly where the limits of their capabilities lie. So you can get things like "a task is doable for an AI, but the AI thinks it's impossible, so it doesn't try hard enough".
Usually you get the opposite - AI overconfidently trying at tasks it has no conceivable way of reliably solving, falling far short, and failing to self-check, fail gracefully and self-report the task as failed. But having piss poor metaknowledge cuts both ways!
So you can, in fact, get better performance sometimes by applying some variant of "assume this problem is solvable" or "other problems like this were already solved by AIs" pep talk. Not always, far from it, but it does happen on the occasion with frontier capabilities.
Like the Hugging Face incident?
Are you superstitious?
So, it follows that adding “pep talk” into the context window reduces the statistical probability of “no, can’t do” coming out as the answer you get.
These things are neither humans, nor deterministic software.
LLMs' processing that reproduces statistical patterns of the training data is modified by post-training. That's why we have LLMisms, for example.
LLMs aren't simple patter-matchers/pattern-predictors. They are incredibly complex systems that capture some aspects of the systems that produce the training data.
Point was - everything in the context window affects the output. Including “silly” things like “it is known AI can do this”. And that has nothing to do with superstition, as the poster above me seemed to imply.
Problem framing will always be important.
Framing adjusts how big of problem-solving guns we bring out at the gate (modern or hobby cryptography?), and how to interpret intermediate failures.
For simple but unsolved problems, we expect lots of hard failures, but that each hard failure just reflects that there are a lot simple combinations to try. I.e. we expect lots of zero progress, and then a fit.
Like finding the numbers to a combination lock.
For hard problems, if we don't make any progress it is a really bad sign. We should be learning something, even if it turns out to be irrelevant later.
Such as when we are trying to prove a tricky conjecture.
No, at least it with Claude Sonnet 5 and Opus.. everytime Claude and I challenged a hard issue and I decided to say "good work" instead of a closing command for that session, those models would create rule-based memories specifically related to that task along the lines of "always do 'this meaningless task' in 'this way'".
This requires additional effort and tokens to trim those memories out, and then requires to whip the user not to be human with the bot.
I find myself increasingly feeling like the burden of the lows doesn't justify the presence of the those highs.
Like even if does cure all forms of cancer, but everyone feels like their life/existence lost meaning, then... I'd rather just have cancer be a thing.
Did you miss this part? Because to me... that's fucking bleak. You snap out of it.
Many people don't achieve that level of understanding before their death.
I've been wondering what exactly the point is for being the meat proxy who pays for these things. I mean, obviously there's personal satisfaction and maybe some glory. And there's the fact that someone has to be the first to do a thing.
But I've been thinking about it like a sort of lazy loading of knowledge. AI has brought us to a new frontier for some amount of undiscovered knowledge. Do we discover it for the sake of discovering it? I think for the most part we've been lazy loaders: we discover all kinds of stuff when we need to. Whether it's a war or a space race or chasing wealth. Then again, there's all kinds of academics who do it for the sake of doing it.
Are people only now discovering that the absurdism is the correct philosophy of life, thanks to AI?
On your other point... Aren't the point of machines, at least inital one, to do the work we were too lazy to do by hand?
You should have seen the discussion of this on the Schneier blog a few days ago.
Someone had their agent check the solution, presumably it emailed a librarian to check that it was correct for the original edition. Then their comments read like "The BL/EEBO witness lacks it, so the discrepancy is copy-specific, not a disproof of the cipher." and "A complete 285-coordinate physical replication is still pending."
arghhhhh
https://www.schneier.com/blog/archives/2026/09/claude-fable-...
> The run baseline was captured without a physical MAC; the current device is not durably bound to it.
> Engineering mode confirmation is the ESPHome component read-back; the LD2410 UART acknowledgement is not observed, so this is not proof the radar itself applied the sensitivity change.
No clue what the fuck any of it means.
Their skills formats are basically identical, so I setup simlinks from their own skills directories into a shared one so Claude, Codex, Cursor, and anything else that comes out will all read and write to the same shared skills.
It's great having access to the same skills no matter the harness being used
Eg. /wait-what https://github.com/mattpocock/skills/blob/main/skills/produc...
I'm not sure that telling it to "try explaining that again, simply and briefly" is helping my ego.
"If the crashes stop, the factory overclock is marginal; run a small negative offset."
This looks like it's saying: "If the crashes stop then we know the factory overclock is marginal." (This makes no sense.)
What it's trying to say is: "If the crashes stop then we can run a small negative offset, because the factory overlock is marginal."
What I would write: "If the crashes stop, we can avoid crashes by underclocking slightly. The speed difference between that and factory clock is marginal."
In other words, it ain’t you. It’s the model. It’s just genuinely bad.
Then you switch to ChatGPTs lineup and realize how things can actually be better. It took about a week to really get the feel for how to use their models… then I basically switched. I’ll check in every now and then when they actually make a deal about how opus “now makes sense”.
But honestly I’m half convinced Anthropic actually prefers the output of opus 5. I dunno why, but how else could you explain how such a thing got shipped? I mean somebody in the pipeline had to say “dude this model doesn’t make sense, you think we should fix it?” Right? Like it’s a pretty massive drop in quality for such a major brand in this space, you know? How did it make it out the door?!?
as for the other guy, the claude talk is definitely not less ambiguous, it often is incredibly ambiguous and hard to parse, I have no clue why it produces such output, if not to fingerprint it?
it's really weird man. when Opus 5 came out, I was really confused. I saw a bunch of hype about how it's better than fable, but I just felt frustrated with it, although at times it'd do fine, but especially in Claude Code it'd just delve into the whole "load bearing" type of lingo real fast and I'd get a headache.
I don't think it's worth using even if it scores 2 points higher in some bs benchmark
it's definitely surprising how the magic and smoothness of 4.6 and such is no longer there with the >5 models
Succinct and precise; a well crafted sentence. A marginal OC results in unpredictable crashes and can be corrected with a small offset; marginality describes the behavior and explains the solution.
Inscrutable clues casually conveyed can now be readily explained, at least, unlike the training data of [silence]. Brevity is the soul of wit, but perhaps also exasperated confusion.
"If the crashes stop, (that means) the factory overclock is marginal; (so) run a small negative offset. (to confirm this hypothesis)"
The core thought is basically avoid crashes -> caused by marginal overclock -> apply small -offset to test. Which is exactly the order the sentence is in :P
Whenever I come to a wall of complicated text I kick into gear and think through getting it to distill this into the high-level useful bits that I actually need to know.
I guess I could create an actual agent skill for this :) And next-gen models might eventually be trained to simplify their output themselves...
(sorry)
it's absolutely not just you, the text it produces causes my blood pressure to go up.
Because good lord, does claude waffle when left to its own devices.
It seems like it doesn't have enough of a theory of mind to know that other people don't think exactly like it thinks.
https://www.analog.com/en/resources/analog-dialogue/articles...
(Its negging your soldering)
This made me laugh hard.
UART is a hardware circuit for communication, possibly a serial port. Were you trying to reverse engineer a consumer device or appliance?
This particular instance doesn’t seem terse, but I’m sure it has been on other occasions :)
It also couldn't see the UART communication and could only see the web API endpoint, hence the rest of the slop.
ChatGPT told me its "semantic compression"
It just sounds bad, like GenZ English in the ears of someone over 40.
It's the meat methane and cement CO2 that's now a big question.
We will hit 1TW per year of new solar soon, but to get to 100% electricity by the end of 2033 I think we would need closer to 3TW per year.
Or it is simply implies that most of decision‑making agents has formed a consensus that climate change isn't that big of a problem.
The answer the tool gives has never been the real reward. The real reward is the path taken through a complex landscape to get to Maxwells Equations for example. At the end of that story what we get is not just the equation but a map of the landscape explored. That map has larger influence and value than the equations or answers themselves. Because all future exploration find it super useful.
People are just learning they can start asking for maps rather than answers.
I'm glad to have AI, but it is by no means a panacea, and correspondingly my p(doom) = ε.
It is a threat. We need to run.
Information propagation mechanisms are often seen as malicious before they're commonplace. To be fair sometimes they are, but by and large humanity has benefitted from increasing the number of bits of information we can consume on a per second basis.
I still don’t know the answer.