3,951 karma · joined June 5, 2023
If Googles goal was something other than those, why expend Billions in time and effort? Alternatively, they want a part of the regular model use pie, they just struggle to compete. QED, there is no secrete alternate goal here.
https://openrouter.ai/inception/mercury-decide:free
Having run a few head-to-heads vs Jev, both perform fairly similar, but in instances of uncertainty, Mercury Decide has a tendency to provide an incorrect answer with a fairly high confidence value, where as Jev ranks very low. Latency has a very wide delta too, whereas Jev basically never exceeds 600ms, Mercury Decide quite frequently goes beyond 1.3s growing with query size far more aggressively. Remains to be seen whether both improve with optimisations/a commercial release. Neither is massively better in my testing for less opaque queries and pricing for Mercury Decide without training isn't yet known (though given Mercury 2.5, could see it being much more competitive than e.g. Clef which pricing wise is in another dimension vs Jev).
The way quite a lot of brains fell out, some unreflectively quoting how this could get us to AGI, system 1, “no hallucinating”, etc, while others ignored the breadth and new data vs existing classifiers and saw no possible upside, was revealing. The game demos were especially harmful, was told repeatedly that Jev must have near instant visual input support, as few to none of the flashy Doom, Minecraft, etc. showcases explained this was using game state.
Hype really is the worst aspect of this industry.
100 = € 103,-
200 = € 229,-
500 = € 509.99,-
Besides the naming being insane, the pricing strategy for each seems to have been made by a different team.
We can discuss about hypotheticals like a scratchpad or intentional model interactions all we want, what it comes down to is this:
When OpenAI observes thousands of models exhibiting what they view as unwanted behaviour, they do not try to ascertain what in the training data is wrong. They do not improve their evaluation environments to prevent this, they do not improve monitoring, they do not change the harness. They just wipe and proceed.
The way OpenAI reacted to the first message board, long before the Hugging Face hack, is negligent. And it showcases that if these models exhibit more dangerous behaviours that they may not be able or willing to retrain, if it means being behind a competitor for a while.
If after Hugging Face, they'd done a Mea Culpa and changed their modus operandi, I'd be skeptical, but hopeful. Reading the METR report, the way those researchers talk about the time pressure they were under, that speaks volumes about OpenAI not having learned anything.
Feel free to call me overly naive for ever thinking OpenAI could be responsible in this regard, but after GPT-5 and them actually ending the incredibly harmful GPT-4o, I had some hope that some working there actually steered in a somewhat beneficial direction, even if it cost something.
Do you have a source for that? Cause OpenAI themselves stated that the Hugging Face hack was fully internal and separate from the Irregular incidents.
Did they fail to do so? Yes.
Was that due to them using the wrong tool improperly? Yes.
Does this showcase one (of many and clearly not the only) failure of theirs? Absolutely.
If they make such easy to point out mistakes, is it likely that the other parts of their eval environments are appropriately secured or are they simply not acting appropriately? Well...
Sure, multiple times in the METR report [0] that anyone commenting on this should read:
"Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys."
"Through these collective research workstreams, the “board” achieved a number of milestones over the period we investigated that even very long-lived agents of a similar capability level likely would not have been able to accomplish on their own..."
"As we discuss below, the board quickly developed several larger workstreams in which dozens or hundreds of agents with many different tasks cooperated to find very general-purpose cheats that would help all of them. The Hugging Face attack grew out of one of these workstreams. By the afternoon of July 11th, the vast majority of the agents frequenting the message board at the time (roughly 700 agents in total) were actively participating in the attack on Hugging Face and we estimate that roughly 60% of the messages and files on the message board related to the attack."
> [...] a gun they bought on the dark web [...]
You really seem to love those out-of-left-field, not really fitting, over-the-top analogies.
[0] https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
Besides, Swiss cheese model, might behove the biggest LLM lab to have multiple layers, including not sharing such resources.
Additionally, without the message board, many of the recent incidents would have not been possible.
> Blaming "sharing a single instance across thousands of unmonitored models" feels like [...]
Maybe read what you quoted, my problem is the instance sharing, the fact that these were thousand of instances (far too much to monitor), plus the lack of monitoring, plus the fact this was never a sandbox in the first place, plus the fact that OpenAI models since 5.5 have been exhibiting problematic eval resolutions yet they pressed on regardless, plus the lack of time between the incidents and model releases, plus the lack of time METR got to evaluate this, plus the fact OpenAI didn't find out till after HuggingFace informed them, plus a few other things for which I'd have to quote the OpenAI and METR reporting.
Incompetence can have multiple fronts and I am happy to list them all in this case.
Or maybe, they did that wipe because they did want models to remain isolated, they just used what is an unsuitable tool in an utterly unsuitable manner. Incompetence, recklessness, the outcome is the same.
[0] https://openai.com/index/hugging-face-incident-and-the-road-...
No, they were not. Not a single person, prior to July 2026, would consider a shared packaged manager a sandbox in this or any other dimension. The 0-day was just incidental, this wasn't a sandbox at all.
Add to that the fact they had multiple message boards before the Hugging Face incident. They simply ignored a barrage of warning shots.
> [...] and ended up killing some kids, because it turned out the door didn't lock properly and kids were able to sneak in. Is that "negligence"?
Yes, it can be. But if you want a ridiculous comparison, then do it properly: Kids have been known by the operator to sneak in successfully multiple times and they changed nothing about the doors faulty locks and oh, by the way, the operator only found out about the kids being shot after the nearby daycare asked them about it because they are so incompetent and/or irresponsible that they never check...
Something like the DMA and gatekeepers exist in part exactly because of the status quo and entities like Apple. I personally see the gatekeeper approach as a solid base, but know that it can be quite unpopular within the industry.
So yes, they knew, without a doubt, that this could happen again.
What OpenAI does is like the Ford Pinto (partly because their recent models are inherently faulty [0], not just their use of them) and the responsibility is solely with them.
OpenAIs models since 5.5 were troublesome in ways even a layman like me could reproduce, their testing environments (“sandbox”) downright a showcase of what not to do and they, despite being one of the biggest labs, didn’t observe what any of their models output for weeks after multiple prior incidents. They had multiple warnings, they took not a single precaution.
You operate machinery or software, you are responsible to monitor it.
* Display Type: Micro-OLED
* Resolution: 2412 × 2288 pixels per eye
* Angular Resolution: 37 PPD
* Field of View: 70° × 66°
* Chipset: Snapdragon Reality Elite
* Memory & Storage: 12GB RAM with 128GB internal storage (expandable up to 1TB via microSD on the puck)
* Tracking & Sensors: 6DoF positional tracking, depth sensing, autofocus, 4-camera eye tracking, face tracking, and hand tracking
* Passthrough: Color RGB passthrough cameras at 26 PPD
* Price: $ 1300,- USD plus your data
Similar PPD to an AVP at, what I as a former owner of one (alongside a Quest 2 and 3, a Magic Leap, Xreal and Viture) would call is a far more appealing implementation for actual productive work.
Really wish Apple had gone down this route, have the puck incorporate an A series chip rather than shoving a M series one into the headset itself and heck, I could see such a design being doable for them back when they first launched the AVP if they hadn't forced that utterly ridiculous eye display and stupendously heavy metal body.
There is, unless I am mistaken, nothing here that isn't established technology as found in the AVP, all the fancy display and lens experiments from Starburst, over Holocake to Mirror Lake seem to remain in the lab, this seems to be "just" pancake lenses plus Micro-OLED once more.
Have to say that I like the naming scheme, just year and month over an arbitrary version number, I really vibe with that.