The problem isn't the 20% actual bugs found (that's great), but the 80% false positive rate which are reported with high confidence and (most likely) misleading reproduction code. An experienced programmer familiar with the code base first needs to validate
all reports and throw away 4 in 5. That's a
massive waste of time. If a traditional static analyzer had an 80% false positive rate nobody would take it serious.
In my hobby projects I use LLMs in my development workflow mainly for passive bug scanning, reviewing and helping to maintain tests, they are definitely useful for catching some bugs early and noticing unhandled edge cases, but they're also definitely no silver bullet (they sometimes ignore quite obvious bugs, and the fewer 'obvious' bugs remain the more one has to be careful about false positives). E.g. the funny thing is that now I'm actually slower than before due to the intense 'rubber ducking' with LLMs and cross-checking their results, but I still want to pretend that the resulting code is more robust out of the door.
Eg everything that Greg KH says in the video sounds very familiar, and it's very disappointing that Mythos still suffers from the same issues (or maybe even worse) as older models.