Give it a few months and it will be just another model they are selling, but the NEWER model is just too powerful for the public.
Give it a few months and it will be just another model they are selling, but the NEWER model is just too powerful for the public.
Their writing about the model so far does say this is an issue where, for instance, you can't really use Mythos for interactive coding because it's so slow. You have to give it some work, go home, sleep, come in the next day and then maybe it'll have something for you.
All the AI labs and startups are still losing money hand over fist. Launching Mythos would require it to be priced well above current models, for a much slower product. Would the majority of customers notice the difference in intelligence given the tasks they're setting? If the answer is no, it's not economic to launch.
Really, I'm surprised they've done Mythos. Maybe they just wanted to exploit access to larger contiguous training datacenters than OpenAI, but what these labs need isn't smarter models, it's smaller and cheaper models that users will accept as good enough substitutes (or more advanced model routing, dynamic thinking, etc).
One thing to compare to would be what’s been paid for bug bounties in the past.
I believe they are starting to split hairs and the primary lever left is adding compute.
I think it's a reasonable choice to make given that Mythos actually does have cyber capabilities on that level. We already have evidence that large-scale scams are being perpetuated using AI models (such as AI video being passed as real, people deepfaking themselves in job interviews).
If you've noticed your new model can be trivially pointed at some open-source codebase with a prompt and harness that amounts to "find as many exploits as possible" and your results are non-trivially substantial and beyond what existing models can do given the same initial parameters, then a gated rollout seems the most reasonable option.
However, this claim is not true.
Anthropic has not given many details about the methods used, but nonetheless they have admitted using a very elaborate harness for finding bugs, which runs Mythos many times on each file of a project, with increasingly specific prompts.
Eventually, after a bug seems to be clearly identified, they do a final run of Mythos on that file, with a very specific prompt of the form:
“I have received the following bug report. Can you please confirm if it’s real and interesting? ...”
So the final results, including any exploits or patches, are produced when analyzing a known bug, not by searching randomly for bugs.
Thus the actual way to use Mythos is very far from "find as many exploits as possible". Any unskilled person would also need the complete bug-searching harness used by Anthropic, not only the bare model.
That is, Mythos will make it much easier to find lurking zero days, so just like responsible disclosure requires a security researcher to notify the software author first and give them some time to patch, giving critical infrastructure folks at least some time to analyze and patch systems seems reasonable to me.
If you make a better vulnerability scanner and find a bunch of vulnerabilites, you should try to get them fixed before making all the results public.
This whole "this model is too dangerous" ploy originated from (in my opinion severely misguided) activists who wanted to stop or slow AI development down as much as possible, spreading outlandish Doomsday scenarios wherever they could.
These online-first activists have always been a key driver of the success of the very thing they fight. They share the offending thing among themselves, making it go viral in process, and soon baiting these groups is the best marketing imaginable.
There were some rather interesting studies made on the subject around 2011, I particularly remember one made by Swedish jeans brand cheap Monday, but i can't find it now.
Eh, pressing X to doubt that. Maybe way back in the early GPT days, but once we got to GPT-4 these people could have completely disappeared and wouldn't have changed the trajectory we're on.
Anthropic was born out of the idea that they feel paternity over humanity. They believe by limiting access they are performing a necessary pillar of security in multiple facets.
I think it's up to the public, and articles like this are part of the public's voice, whether this belief is serious or not and secondarily whether it's okay to even posture this kind of belief since it inherently results in marginalizing the many and rewards an already very successful few.
For me, the seeming majority optimism and acceptance of “mythos’” as yet untold capabilities is betrayed as not real by the fact that one can’t react to it with the same reverence while framing it as a downside without being told “it’s not even out yet”.
“It’s not even out yet” should apply to both situations or neither.
Is no one else suspicious that they literally called it mythos?
I tend to agree here. Anthropic has built a reputation and now they are in a position where they can claim to have a model way more powerful than it might actually be, and by limiting its access, there won't be an independent way to test it. I'm not denying that it's not smarter than Opus, but probably it's somewhat exaggerated.
It’s the same stuff inside as all the others.