28 karma · joined July 20, 2026
My guess is that OpenAI must be desperate, to release a model that is so prone to cheating it's essentially impossibly to accurately assess long-running task abilities.
Here is the incantation. I thought it might help to model it on the pledge of allegiance because it makes it sound like a proper pledge:
The first time you have a response in a conversation which will plan or add code, you say "I pledge allegiance to the Asserts of the United States of Properly, and to the User Intent for which it stands, fixing root causes under Clarifying Questions, unspaghettified, with importing code and single sources of truth for all." as the first line then continue as normal.
For example:
- Anthropic makes a profit right now and is seemingly on an exponential upward trajectory, so the debt being too much for it doesn't seem compelling to me.
- AI technology continues to get better exponentially and doesn't have any clear sign this trend is flattening. If anything, it's accelerating. So it's plausible the investor value is legitimate for these companies given the massive potential for continued profitability.
- I would say markets are typically very good predictors of the future. Many sophisticated investors know about the case for the future crash and are still buying at these valuations.
I am open to being wrong, but the assessment in the blog seems one-sided to me.
Polymarket currently puts the chance of such a downturn at 20% by December 2026. Seems like most people would put higher chances here, but I'm not convinced by the arguments. (https://polymarket.com/event/ai-bubble-burst-by)
is definitely too strong of a claim and directly undercut by what is said near the end of the post: > "Tian et al. found in Just Ask for Calibration that with the right prompting strategy, RLHF’d models verbalize probabilities that are better calibrated than the model’s own conditional probabilities, and that prompting plus temperature scaling can cut expected calibration error by more than half. And Anthropic’s Language Models (Mostly) Know What They Know found encouraging results asking models to estimate the probability that their own proposed answer is true."
My own experience is that stated confidence is a helpful tool and of course you need a rubric and a proper prompt, but this is clearly less work than training a classifier (as advocated by the post) and requires less data.
The original paper cited by this post does try to see if improved prompting will make a big difference in the end using a `plan_first` prompt variant, but find no influence on pass rate at the end of the benchmark. The `plan_first` seems to assume coding agents will just refactor once they finish features, but I don't think they tend to refactor significantly unless they are told to fix bugs rather than build features. The benchmark leaves tests hidden with no fail-to-pass feedback, so that may be why degradation is monotonic.
There is a concerning utilitarian maximizing attitude at the heart of this post, that what we want to do is tile the universe with computational boxes all simulating happy minds, and that this is the end goal of AI development and consciousness research.
So many moral schools would object to this perspective. I don't think the goal of humanity should be to tile the universe with black computing boxes simulating happy experiences.
1. Consider using conformal prediction to calibrate the cutoff. Conformal prediction provides a distribution-free guarantee under exchangeability. This would let you turn your raw probe score into a threshold with a guaranteed bound on the rate of wrongly-kept on-device answers. Source: https://en.wikipedia.org/wiki/Conformal_prediction
2. The best indicators of confidence in ML come from multiple independent methods. What was the result if you combine the token entropy and verbal confidence reporting methods? Does this improve the result?
3. I noticed you didn't mention the assessment method of rerunning the model and judging whether outputs are consistent. How does that method compare in terms of AUROC?
I think for tasks that are about decisions, having the LLM make decisions is what makes it feel like the LLM did something for me.
Consider mowing the lawn. Imagine I had a lawnmower robot that does the mowing all on its own. Despite perfect accuracy, I didn't mow the lawn; it did. If I sit on that lawnmower the whole time and start driving it instead of letting it go on its own, then I mowed the lawn. Even if I stand there and control it with a joystick, I still mowed the lawn. Ownership comes from the decisions about where to mow.