No one is arguing about the architecture of the model. It's the objective function and optimizer.
If this base is then trained using RL towards a different objective (maths and coding), the model becomes fundamentally a different thing and the recent models are clear evidence of that, regardless of they fact they remain autoregressive.
If you modify an engine to increase it’s output by adding sensors and an ECU, you don’t change the fact that is powered by gas.
If you use RL to increase the accuracy, it’s still a next token prediction, just more accurate.
An aside, I finally do appreciate single column format now, makes it easier to convert to epub.