Certainly no existing harness today, nor a year ago, could achieve a fully autonomous e2e own with a ‘25 class open model. Even with Opus 3 series I don’t buy it!
Certainly no existing harness today, nor a year ago, could achieve a fully autonomous e2e own with a ‘25 class open model. Even with Opus 3 series I don’t buy it!
There are multiple AI network pentesting and redteaming startups that have been on the market for at least 3-4 years now and have conducted similar actions.
Horizon3 and Pentera off the top of my head, but I remember Crowdstrike, SentinelOne, and Wiz had similar capabilities on roadmap around 18 months ago (and in Wiz's case GAed).
And that’s all just “from memory” without something like RAG or web search access…
The newer models are still more capable, but there were people doing this and writing about it (e.g. XBOW).
Look, a lot of energy from third parties goes into demonstrating the maximally bad thing a model can do (and the standard response around here is always to throw shade on frontier capabilities).
It has nothing to do with the guardrails; Pliny jailbreaks those within hours. It has everything to do with task horizon coherence and overall IQ.
It’s simply revisionist IMO to claim this stuff was latent all along.
Automated dynamic exploit chains and credential discovery has been something every red team worth their salt does at a lesser scale for 20 years, why would you ever think that the capability didn't exist until it only cost a few thousand dollars to do?
Existed does not mean was easy, or was cheap, or was thought of as reasonable to attempt.
Also, IQ is not a term that applies to language models, and the term you are looking for is "Long Horizon Coherence".