We’re seeing a massive slowing in the value of all that additional training. Folks don’t like to talk about that, but absent a completely new break-thru the current math of LLMs has largely run its course.
We simply don’t need massive training forever and ever. We’re getting to the point that “good enough” models will solve most use cases. The demonstrated business value is also still broadly missing for AI on the level required to keep funding all this training for much longer.
I don't think the argument is that isn't true, it's that the gains from those massive training runs is diminishing. Eventually, it won't be worth it to do the run for each new idea, you'll have to bundle a bunch together to get any noticeable change.
I'm adapting all of this to Rust+WGPU with compute shaders if you want to follow along.
See this repo: https://github.com/tmzt/shady-thinker
Goal is Qwen3.5 27b on a Pixel 10 Pro running GrapheneOS.