This is basically the answer, they generate A LOT of synthetic task rollouts in parallel, then use RL on the resulting reward signals to improve the model. Add scale to this and you have a Fable class model.
Does not explain timing
keep in mind fable = mythos which as been "done" since february. so the gap is not 2 months, it's more like - techniques probably started "working" in late 2025, now are trickling down to 2nd tier labs 9 months later.
Yeah that would make more sense, it's probably a tight community and word gets around when something starts working.
Maybe because frontier labs buy the same RL tasks from task producer companies.
Who are these task producers? Are you saying that Anthropic, et al delegate the RL part to third party companies that do it for pretty much every other AI company as well?
Turing, etc. And yes.
There are companies that will pay you $$$ for technical challenges that stump frontier models. I’ve met these people. They make good money.
Yes they're called RL gym companies and there's a whole ecosystem of them. You hardly hear about them because their only customers are AI labs and RLVR is where the improvements are coming from at the frontier right now.
Note that RLVR is incredibly compute expensive but it's CPU as much as GPU.
Yes it does, it just means all the companies come out with similar models around the same time. If what they were doing was completely novel, it would take a long time to repeat. As it is now each company releases a new model every few months, and every couple years the "leading" company changes.