It's hard to not see this as a gut punch for OpenAI. They're lead was largely captured by scoring on value (by way of reset after reset) and now they're getting eaten up on price and being bestes and equalled on performance. I'll still pay a premium for Opus 5.5 right now because it's nearly unlimited use, but Google is the quiet sleeping king
Everyone is happy to watch everyone else, but I'd wager google burns more tokens through their search product than basically anyone else and now they're just quietly pacing the frontier...
The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice.
How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.
Accuracy is truly what people die for in the OCR game. Price isn't the primary function here.. it's an equation of price, accuracy, speed, and in mayn cases regulation.
Can you explain what you mean? These days courtesy of an addictive reset game OpenAI is playing, I can't find anything with frontier intelligence that's more cost efficient...
What genuinely disappointing result. Long time pixel user here and I've been routinely buying these phones with the argument that you're getting the most value of any modern smart phone. Now? I'm just waiting for the pixel 10 family to drop in price. Happy to just wait.
There is a more serious question in here that's not being answered. How effective is the retrieval in finding buried needles in larger and larger haystacks. And there's a correlary question, how effective could you be in finding paired needles in that haystack where you need to hold a needle to unlock finding another needle.
I'm trying and failing to find value running a potential Qwen 3.8 27b dense model on a 16 core, 128 GB of ram, 2080ti box. Yes, the GPU yells for help, but the problem is that no math works to upgrade this machine even when pouring $200 in rent every month into the large model providers...
How are you all justifying economical use of these local models right now? What's the cost efficient way to do this and do better (even with models evolving over time and losing now vs later) than the big labs?
Honestly, I have a 2080ti that I use to play and I can tell you the math isn't there to upgrade it. It's much easier to just find a 3090/4090/5090 and keep pace with the software and hardware simultaneously.
I was just using infinity parser 2 (flash, to be fair) for pennies self-hosted to run through thousands of pages of documents with remarkable confidence. I decided to use https://huggingface.co/datasets/allenai/olmOCR-bench to determine what was the best OCR tool, yesterday, but I've got no idea what the best is now. What is the dominant OCR eval right now? Between Baidu and Mistral this morning, I wonder if there's a new tool to switch to..
How does this compare with infinty parser 2 which seemed to be running the table on every other OCR tool (https://huggingface.co/datasets/allenai/olmOCR-bench). To be fair, there's no single winning OCR benchmark and this isn't showing up anywhere yet..
This sounds incredible. Have these models effectively solved the problem of trying to use a fast-processing network to predict the world's state ahead? For example, to catch a ball?
The problem here is always the cost-benefit. For $200/mo, you're receiving subsidized best of breed access. There's no model competing for that price anywhere. If a 27B param model is what you choose, show me your hardware! I would love to be wrong...
I'm really running into this deep at the edges of content creation. Take, for example, a need to general some kind of legal work. The cost of painstakingly checking and rechecking each case cited is reducing the value of these frontier models immensely.
Coding, however, is solved like magic. Easier to add tests, to be fair.
Is there a good benchmark tracking hallucinations? The models are all incredibly good now, even the open ones, and my hope is that the rate of hallucinations is something that's falling off in concert with larger and larger context lengths.
This is a tough moment. Claude is simultaneously becoming substantially more expensive, substantially less reliable (single 9 of reliability), and substantially less performant. It's really hard to justify the cost of a subscription over there right now.
I wonder how this kind of response from Anthropic is actually being read by the community at large. If you consider the rough sentiment of the r/ClaudeCode subreddit against the r/Codex subreddit, you can see that there is a definite loudness among the folks departing ClaudeCode for Codex. Something big is shifting on the ground, I think.