It used to be self-expression in an oddly entertaining way but that Nikita Bier ruined the whole thing with his metrics chasing algorithmic shifts.
653 karma · joined February 8, 2025
Pretraining + RL works, there is no clear evidence that it doesn't scale further.
I find it funny that Microsoft is scaling compute like crazy and their own products like Copilot are being dwarfed by the very models they wish to serve on that compute.
I suspect that attention is naturally tuned to work towards genuine interests which may be orthogonal to conventional value producing tasks
Earlier, I had to only keep my phone away and not open Instagram while studying. Now, even thinking can be partially offloaded to an automated system.