Like, "4x as powerful as a team of engineers" is still really quite impressive
In my hobby projects I use LLMs in my development workflow mainly for passive bug scanning, reviewing and helping to maintain tests, they are definitely useful for catching some bugs early and noticing unhandled edge cases, but they're also definitely no silver bullet (they sometimes ignore quite obvious bugs, and the fewer 'obvious' bugs remain the more one has to be careful about false positives). E.g. the funny thing is that now I'm actually slower than before due to the intense 'rubber ducking' with LLMs and cross-checking their results, but I still want to pretend that the resulting code is more robust out of the door.
Eg everything that Greg KH says in the video sounds very familiar, and it's very disappointing that Mythos still suffers from the same issues (or maybe even worse) as older models.
Okay, but again, even with all that extra effort, they found 4x as many bugs, so it seems like the effort is clearly worth it.
And each model gets more reliable, we get better at building proper reproduction code, etc. - this was mostly a comment about the cycle, direction, and velocity we should expect from the future, given this has already happened twice.