451 karma · joined May 3, 2017
you didn’t have to use China as an example, the US clearly does not care what its citizens want as the most popular policies are never even discussed or proposed in congress
meanwhile, China destroying their housing market to decommidify it so everyone can have housing…they seem to care about their people more
the sparks have much slower memory bandwidth is the trade off
i can imagine insane amount of capital is wasted on these two companies compared to the efficiency elsewhere
You add requirements and make previous tests invisible to see how pigeon brained the model is - Sol and Fable seem to rank the same as Opus tends to fall behind
still don’t think anthropic models are worth the money
I run a lot of SlopCodeBench - https://github.com/michaelasper/benchmarks
Fable/Sol/GLM 5.3/Kimi are its league (in that order) Deepseek/Opus is solid Qwen 27B is the floor - there's no reason to use Sonnet/Terra/Haiku
For everyday activity - I don't think you need to be using Sol (xhigh) for everything - unless you're made of money - I've found using Luna from OpenAI to be more than enough - it'll outreach to Opus/Sol when it needs to
Haven't had access to Gemini 3.7 but we're getting it at work soon, will give it a go!
Codex CLI is pretty bare bones in a bad way (at least Pi is extensible). Claude code is vibeslopped to the extreme
(talking about leadership, not my lovely msft engineers reading this)
Also if your thing doesn't work with `pi` out of the box, then low effort
it’s more akin to 3d printing to me, i get the design all setup and let the machine do 3) and then get to play with it in 4)
these LLMs are great are generating arguments but they don't ask questions, we will need mathematicians to shepherd them into more discoveries
i really want to see open weight models crack some breakthroughs