If it's just 20% that is really low. Especially for complicated long chain potential bugs.
Most other detection tools have much higher rates of FP, or much higher rates of false negative.
Then you have humans that miss bugs for 20+ years. Or, they don't tell you about the things they thought were bugs they wasted hours on themselves. Because of this it's really hard to measure how bad/good the AI really is.
It would be interesting to know why the more SOTA models are getting the FPs. Is it from a lack of understanding of C? Is it complex code with deep branches? Is it code smell and convoluted logic?