I've played with explorer agents giving exploration summaries to help the implementer agents use more of their context for implementation, but it doesn't work as well. There's always something lost in the handoff.
895 karma · joined May 9, 2018
I've played with explorer agents giving exploration summaries to help the implementer agents use more of their context for implementation, but it doesn't work as well. There's always something lost in the handoff.
It's very easy to answer this without my help by trying to get access to Mythos.
Do you see requirements clearly listed anywhere?Can you even apply?
What you'll find is maintainers of large open source projects and analysts' reports with vague statements like - "should follow strict security requirements":
"Trinidad also noted that the Anthropic announcement pointed out that each of the 150 new participants, in Anthropic’s phrasing, “will need to meet our security requirements before they gain access.”
Trinidad said the security requirement claim doesn’t build confidence, because “nobody knows what those security requirements are.” [1]
It's also some random rich companies like Hitachi or Dragos [2]
Do you trust that Hitachi and hundreds of other random organizations will be able to contain Mythos and not accidentally attack your project or your bank? I don't.
> Yes we do. That's why there is the saying "regulations are written in blood"
We absolutely don't. We have already learned with blood that gating access to security based on the number of zeroes in bank account and authority is a horrible model. We can apply this knowledge to LLMs, we don't have to spill blood again.
[1] https://www.csoonline.com/article/4180265/anthropic-grants-p...
[2] https://www.bankinfosecurity.com/anthropic-limits-on-ot-acce...
And I don't doubt that, not in the slightest. But I've seen exceptionally smart people in one field being dumber than a random kid from around the block in another.
This incident is clearly at least 2 failures that could've been easily avoided: failure to communicate, and failure to investigate the logs after letting the "most dangerous" roam free.
No, it doesn't require creating a mock internet with an alert as a side effect. Their own "most dangerous" model could have probably told them this happened if they supplied logs to it.
What, Anthropic didn't know model could escape sandbox without OpenAI reporting it?
Who knows how these companies are using it. If Anthropic can't effectively contain their own models, can the partners?
While the rest of us get fallbacks and warnings, not even being able to defend against the attacks they themselves are causing.
Do we really have to re-learn all the industry's knowledge the hard way?
> In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations
> we identified three incidents
> The incidents involved three different Claude models: [...] and an internal research test model
This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index.
I may be too cynical, but the well of benefit of the doubt is running very dry towards AI labs that like to engage in this game.
It's very taxing, especially since these are usually multi-paragraph texts. I noticed I've started doing a lot of "hey, you're talking gibberish again" a lot with 5.6 Sol.
- Models that can be launched as subagents are hardcoded (can only be another Sol or Terra, but not Luna). Most of the time it'll just launch same model as parent anyway.
- They encrypt initial task delegation from root agent to subagent, for whatever reason
- You can't switch into subagent view at all, despite the fact that apart from initial root>subagent task handoff, all session is visible in transcript.
GPT 5.4 is/was a very capable model.
It seems like we all tried to contain and organize a system that simply prefers to select its own organization.
Which makes me to think that these skill packs of workflows are really made to make it easier for humans rather than agents.
- Ways to obtain cheap guarded-AI tokens that are not linked back to me and with no danger of getting my legitimate accounts banned
- Ways to get rid of guardrails and have models work on things they wouldn't otherwise work on.
The attackers were already in these communities long before I knew they existed, they already had the advantage. Ones with enough reputation probably have access to even more information and tools than I do.
It is true that these communities exist because guardrails were put in place, so yes, it is slowing them down too - as in they can't just put in their CC on claude.com and hack a hospital. But attackers are much better at finding these communities and utilizing resources available there than defenders.
Personally, I don't have any ethical concerns of utilizing these resources when I put them to actual defense, but I know many people that would, leaving them at a disadvantage.
My point is that there's only one guardrail that will effectively contain the threat the models pose, and it's in direct conflict of the big 2's goals - pull the models from worldwide access completely. Strict KYC and all. And it would only last for so long anyway.
Yes, but didn't it always? Hence why my position is that this will get us back to relatively where we were pre-LLMs.
And I don't know what Trusted Access programs give to defenders, because as a defender who has credentials, connections, but no deep pockets and no high ranking passport, it only gave me silence. I fail to see how this is better than total access.
I don't think the world where defense is given to those that "deserve" it is the world that we all want to live in. Which brings me back to the starting point - attackers are almost completely unaffected. If I masquarade as an attacker, I get way more capabilities already.
Open/closed doesn't matter that much. You can get closed models to do a lot of cyber harm, even with all the guardrails, which currently are heavily skewed towards more false positives.
The only effective control is to level the playing field. If both offense and defense have access to the same capabilities, then we're relatively back where we started.
If you want to ensure chaos, then you do what Dario is proposing to do - create gates that attackers can bypass and defenders can not.
So my question is: is this by design (they know nobody's buying this), or is Dario simply so out of touch with reality?
If it's the former, then why publish this?
Yes, please. We don't know whether we'd have open weight models today, had the chip-prohibition not been in place. Nor would we see the more optimized models such as DeepSeek or qwen.
We also would not see new players entering RAM market after you and your pals in Silicon Valley hoarded the entire world's hardware.
So by all means, double, no, triple down on this.
> We should crack down on industrial-scale distillation operations
And let's apply this retroactively to Anthropic too. You industrial-scale-operation-distilled all of humanity's knowledge. Let's have some of that crack down on you too.
Good times.
Not in US, but had my phone seized by authorities and was asked to unlock it. I'm walking free, but I think what really saved me is that I genuinely complied with investigation and the only place where I drew the line was me giving away my password.
If I used my right to remain silent AND not give away the password, I'd probably be charged with something. I'm sure its very similar in many countries, including US.
I usually just start by preloadig context with plausible legitimate use, have it work and obviously fail, and then ask to figure it out without ever mentioning any high risk words. Model offers to RE itself and classifiers are happy.
Where do you think the principle came from? I've used claude code for a year, and stopped February this year.
It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.
Why is Anthropic is so hell-bent on this auto/silent downgrade? Do they have a single user who prefers an auto-lobotomization instead of a refusal? Have they learned nothing from the backlash the first time?
On the rare occasion that I DO want to go there, they greet me with an impossible gate that takes 15+ seconds to pass on Brave and gets invalidated quickly.
If anyone from SE is reading this: you already failed to protect your data from the LLM crawlers, SE is no longer that much interesting to them. The only visitors you're gating today are not bots, they're humans.
Then I saw the announcement that they'll be merging the bloatware Chinese version of their OS and that was the last day I held my OnePlus phone.
Had they not taken that path, I'd likely have bought at least 3 of their phones since then.
How's Kimi in this area?
It's effectively just a completely hidden thing now.