Just skimming through here but I think you have the wrong ideas with llms, I’d recommend Andrew Ngs course (correct me if you’ve already seen it or something similar).
If this base is then trained using RL towards a different objective (maths and coding), the model becomes fundamentally a different thing and the recent models are clear evidence of that, regardless of they fact they remain autoregressive.
If you modify an engine to increase it’s output by adding sensors and an ECU, you don’t change the fact that is powered by gas.
If you use RL to increase the accuracy, it’s still a next token prediction, just more accurate.
An aside, I finally do appreciate single column format now, makes it easier to convert to epub.