1,438 karma · joined March 19, 2019
On another note: I'm a bit paranoid about quantization. I know people are not good at discerning model quality at these levels of "intelligence" anymore, I don't think a vibe check really catches the nuances. How hard would it be to systematically evaluate the different quantizations? E.g. on the Aider benchmark that you used in the past?
I was recently trying Qwen 3 Coder Next and there are benchmark numbers in your article but they seem to be for the official checkpoint, not the quantized ones. But it is not even really clear (and chatbots confuse them for benchmarks of the quantized versions btw.)
I think systematic/automated benchmarks would really bring the whole effort to the next level. Basically something like the bar chart from the Dynamic Quantization 2.0 article but always updated with all kinds of recent models.
(Mar 5 2022) TinyGL 0.4.1 is out (Changelog)
(Mar 17 2002) TinyGL 0.4 is out (Changelog)
"our plans are measured in centuries"
They for sure did not anticipate that the user would backflip into their robot and knock it (and himself) out :D
It would bear a good lock-in effect for OAI as well
For a very narrow range of professions, like ATCs, time is absolutely critical but for most it does not really matter that much. Especially in many STEM fields. I think people in a broad IQ range can build abstractions and acquire intuitions about pretty complex matter. From this view-point ability to concentrate for long times, curiosity etc. seem more important than "raw-compute".
"if you value intelligence above all other human qualities, you’re gonna have a bad time" - Ilya
Timeless statement imo, even in the absence of AI
At some point, the rot in institutions can not be covered by dumping more money onto it. Luckily, NYC is not there yet, but also I don't think we have a good recipe on how to reset these kinds of institutions back to excellency