7,346 karma · joined January 20, 2012
Model Score Cost / Task Output Tokens / Task
-------------------------------------------------------------------------
Muse Spark 1.2 (xhigh) 56.8 $0.40 30,430
Gemini 3.7 Flash (high) 56.0 $0.40 36,847
GPT-5.6 Terra (max) 56.6 $0.51 20,838
GPT-5.6 Sol (high) 57.3 $0.52 7,545
GLM-5.2 (max) 53.0 $0.56 32,200
GLM-5.3 (max) 59.5 $0.68 41,107
GPT-5.5 (xhigh) 56.3 $0.69 16,893
Grok 4.6 (high) 60.9 $0.84 21,735
Kimi K3 (max) 59.7 $0.84 25,474
GPT-5.6 Sol (xhigh) 59.0 $0.87 11,098
Qwen3.8 2.4T A95B 57.7 $0.95 32,472
Claude Opus 5 (medium) 58.6 $0.98 12,459
Qwen3.8 Max 58.1 $1.13 38,287
GPT-5.6 Sol (max) 60.9 $1.23 16,879
Claude Opus 5 (high) 61.5 $1.52 21,353
Claude Opus 4.8 (max) 57.3 $1.65 33,557
Benchmark score: Model Score Cost / Task Output Tokens / Task
-------------------------------------------------------------------------
Claude Opus 5 (high) 61.5 $1.52 21,353
GPT-5.6 Sol (max) 60.9 $1.23 16,879
Grok 4.6 (high) 60.9 $0.84 21,735
Kimi K3 (max) 59.7 $0.84 25,474
GLM-5.3 (max) 59.5 $0.68 41,107
GPT-5.6 Sol (xhigh) 59.0 $0.87 11,098
Claude Opus 5 (medium) 58.6 $0.98 12,459
Qwen3.8 Max 58.1 $1.13 38,287
Qwen3.8 2.4T A95B 57.7 $0.95 32,472
Claude Opus 4.8 (max) 57.3 $1.65 33,557
GPT-5.6 Sol (high) 57.3 $0.52 7,545
Muse Spark 1.2 (xhigh) 56.8 $0.40 30,430
GPT-5.6 Terra (max) 56.6 $0.51 20,838
GPT-5.5 (xhigh) 56.3 $0.69 16,893
Gemini 3.7 Flash (high) 56.0 $0.40 36,847
GLM-5.2 (max) 53.0 $0.56 32,200Spending money via Google Pay on Android is extremely easy, Google does know how to accept customer's money (in the consumer space)
There are more plausible explanations for why the models are similar - all the labs are buying the same datasets from third parties
So Model N/N+1 might literally have had the exact the same pretraining run and only differ on how much/what kind of postraining they got
So, yes, I would say agents are pretty good at working with Bluetooth on Linux
I recently priced out getting a PCB done in China and its insanely cheap. My design isnt complicated, a few ICs, a few connectors, an a dozen or so passives. PCB manufacturing, parts, and soldering/assembly is on the order of $10-20 total per board and that is with zero volume discounts. The parts alone would cost that much in the US, and I suspect the sort of companies that you contract out to also aren't really interested in tiny orders so they would probably quote huge setup fees (looks like your average order so far is ~$7000).
On the other end of the spectrum, if you grow in fertile native soil from seed your costs can be close to zero
If you scroll very slightly farther there is a benchmark that includes Sol, showing it outperforming Terra (as expected) and Spark 1.2
Unified memory has existed for decades in the PC space, Apple didnt invent it.
And the test you linked to has nothing to do with unified memory, its a web browser benchmark (almost entirely constrained by single threaded CPU performance that Apple better than competitors at).
Several hundred million Indians would disagree
I expect if they add Kimi 3 to Go the limits are going to be really low since 2.7 is already one of the most limited models and 3 is much larger.
By definition there is no model that is both cheaper and as intelligent or better than another on the frontier.
My guess is something to this effect was in the prompt and the LLM made a point of correcting it
Its cool to see this implemented in a tiny amount of code without dependencies, but does it actually bring more performance?
The way I read this is: yes, if the third party harness uses Anthropic's Agent SDK. Most of them do not, AFAIK, and are still against ToS (though maybe its not enforced for now)