2,957 karma · joined September 15, 2020
Even if the input is in plain English, the model never sees any words, tokens or glyphs to begin with. It's vectors all the way down.
OpenAI could have done this same experiment with GPT-4, with possibly even worse results, depending on the quality of the sandbox. Even if the techniques used were not as sophisticated, the natural language output could still easily contain more unhinged sequences of words that lead to the techniques being used.
If the system generates strange conclusions as to when the task is done, or should be stopped, it wouldn't speak to the intelligence inherent to the system.
Not that the techniques used by the LLMs in the actual incident weren't unexpectedly sophisticated, but the outputs of each and every one of these processes could've been read at any time during the run. They just weren't.
Should be treated like a digital cousin of gain-of-function research.
> In this report, we argue that frontier performance can be achieved by a wide range of institutions through Continual Learning on readily available open-weight models.
> As opposed to existing limited approaches such as small-scale fine-tuning, prompt engineering, or tool-augmentation with a frozen model, our Continual Learning approach takes advantage of the effectiveness of a modern mid- & post-training stack while introducing safeguards preserving both plasticity and stability at each training stage and seeking to make the minimal number of high-impact interventions on the parameters.
For the large model, Thomson is utilizing the fine tuning stack they describe in the article, running it on Snowdon 1.0-Large, which in turn is a fine tune of Qwen3.5 397B. Same thing for the small model, but it's a fine tune of Snowdon 1.1-Small, which is a fine tune of Qwen3.6 35B.
As for the small version's run:
> The full pipeline consumed approximately 1.63 × 10²³ FLOP over 35,207 B200 GPU-hours, showing that these results are achievable with compute and personnel budgets substantially lower than commonly thought.
That would amount to around a quarter to half a million dollars of spend on that run. 100k minimum, if they got a great deal.
Building an LLM from scratch has a hard split between a tutorial project you can complete in a weekend (that's useless for actual usage) and then a solid 1km high brick wall if you want to create anything actually useful from scratch.
Modified open models have a very active community around them, without the need to look much further than Hugging Face.
When the post training run is nearing its cutoff point, there's a massive amount of data on how long coding tasks take to complete by that model in the golden format of "task -> time task took to complete", separable to whatever amount of subtasks, in the same format. With the parts from the end of the dataset being useful for evaluating the finished model's capabilites, whether that data is then fed back into another step in post training or not.
Completely separate from even the actual training: If you have a model proactively giving estimates that are an order of magnitude wrong, you can already fix the worst of it as of this moment by just changing the system prompt. It's a dirty fix, but it's the type of fix that has been used by Anthropic and OpenAI since forever when a model is dishing out blatantly wrong outputs.
Not sure if I'm misunderstanding your point?
What I was going after with "aware", is that the actual people working at the companies, training the models, are aware that people aren't mostly going to be implementing the plan by hand, if they've already made the plan in Claude Code or Codex. As for a specific Claude / GPT instance, "aware" would definitely be the wrong word choice there, but the instance does have its stats and environment information in its context window, unless you specifically remove it.
Either way: Training does include estimates on working with the model, and adjustments of the model itself based on that. That's literally what RLHF is.
It's straightforward to have a portion in post training that aims specifically at the model being able to give better estimates on how long that model takes to complete a certain type of task.
It's like every time they make a plan, there's something about things taking "a week or two", "month of focused work", or whatever.
This is something that would've been RL'd out a long time ago if it wasn't great for business.
Anthropic goes to insane lengths to block other labs from training off of their models' output, as it's been done over and over again in the past. But the models that have used synthetic data from Anthropic's models aren't distilled versions of whatever model(s) they got the distilled data off of.
Crazy how time flies, given that talk is seven years old at this point. But could've been given today, with the points being more relevant than ever.
So, not a distilled version of Mythos or Fable, but those models likely helped a lot in the post training phase of Opus.
Both GPT-5.6 and Fable will happily run 3+ hours straight off a relatively simple goal in that mode, burn through millions of tokens in the process.
Not saying it's great bang for your buck, but just that those modes are there and being pushed by the companies.
Optimizing models to be fine-tuned is an amazing direction, but just makes me wonder how much better this actually is at being fine-tuned compared to other models. As none of the modern models are great at being fine-tuned afaik. Basically looking for some sort of benchmark showing that it's resistant to overfitting / catastrophic forgetting, etc.
Would be very interesting to see concrete demonstrations of different fine-tunes of the model. I'd imagine they've done hundreds of those internally.
Either way, inference is very much where the money is made, training is where the money is lost.
At least for the segment of 20$ subscribers who actually use Claude Code it seems that it wasn't being profitable, as a couple months back they were testing out a pricing model where Claude Code would've not been included in the 20$ plan.
https://arstechnica.com/ai/2026/04/anthropic-tested-removing...
Using a full Claude Max 20x plan to 100% of weekly usage would easily cost you 2k through the API. While the Claude Max 20x plan is 200 a month.