1,957 karma · joined January 9, 2012
True. With 5 choices you need at least 126 people before you can guarantee that two lists are the same.
Which is: When you think about Google, true, old, "don't be evil", tech excellence Google, you don't think about Demis. You think about Jeff's and Sanjay's geeky, technically uncompromising, faces.
This makes sense. At most it converts the question of ability to a question of cost (i.e. fire up a prompt for each file).
This article reduces the hype about Mythos in my mind: A new model that can find 9 new bugs while no previous model can identify them, is a whole different story from what this article demonstrates: that only 2/9 of the detected bugs are new for Mythos.
Great work.
This should have probably been surfaced at the text. To answer my own question: If one uses gemma4-26b-a4b, mimo-v2.5-pro and gemini-3.5-flash then 7/9 bugs are covered, while no model can uncover the remaining two.
> And, you have misunderstood what the benchmark does. It tells the model to audit the file, and it is allowed to look at the rest of the repo. It is not pointed at the bug.
I obviously meant that the model was pointed to the bug report. What would have been the point of providing the actual bug. This is still easier than telling Mythos to generally look at a codebase and find any bug.