2,429 karma · joined September 6, 2018
See: https://github.com/lovasoaHardcore
But that means speed does matter after all. The renting company will need to clean as many flats as possible with a single robot in a single day.
Maybe when HF published their blog post on the 15th, OpenAI already knew something had happened, had started to investigate, and already reported the issue to JFrog ? But looking at your other comment in the thread, I agree that CVE-2026-65925 and CVE-2026-66014 are better candidates.
Taking a step back, so many basic vulnerabilities in a security-oriented product just makes the headline "agent autonomously escaped containment" sound a little less spectacular.
* were the ExploitGym solutions actually available somewhere inside huggingface's private datasets ?
* was the model really trying to extract the solutions ? or had some sub-agent drifted enough from the original context that it was not even trying to solve the initial challenge ? that would look much worse for OpenAI, PR-wise.
You can see that 17% of answers come from India alone and that software developers got below average results, for instance.
Running inference for a model, even when you have all the weights, is not trivial.
Linux has bugs, bug MacOS does too. I feel like for a dev like me, the linux setup is more comfortable.
* [1] https://sql-page.com/
There is no fire at all in the painting, only some smoke.
https://en.wikipedia.org/wiki/The_Trojan_Women_Set_Fire_to_t...
ChatGPT (4o): Noticed "a pattern"
Le Chat (Mistral): Noticed a "cartoonish figure"
DeepSeek (R1): Completely missed it
Claude: Completely missed it
Gemini 2.0 Flash: Completely missed it
Gemini 2.0 Flash Thinking: Noticed "a monkey"