They also have incredibly smart people that can understand which improvements and directions are even worth pursuing for, I think these people are being paid in millions
73 karma · joined May 7, 2026
They also have incredibly smart people that can understand which improvements and directions are even worth pursuing for, I think these people are being paid in millions
> Kimi K3 $3.00 $15.00
> Qwen3.8 2.4T-A95B $2.00 $6.00
Why is kimi priced at twice that of qwen 3.8? If I remember correctly both of them are of the same 2.5T param family. I wonder kind of optimizations or hardware they have that make such pricing differences for the same sizes of models
I can see some parallel between how arduino gave hobbyist a mini computer they can use in any way imaginable, I can see this kind of thing being used to solve unimaginable problems by hobbyist or RL computer scientist or heck any UG student with claude subscription
I wonder how much of these tricks actually work? As a side effect does it make the model produce thinking tokens to remember not to do that, “ohhh wait the user instructions says I should not use PG style in my thinking switching back to my …”
I just didn't find the SIA paper to be testing this out rigorously or either provide sufficient evidence that it works