That's a really weird way to test Jev. I don't see any mention of latency or cost in https://github.com/Gaurav-Gosain/jev-sec-bench
If you don't care about that why wouldn't you use the flagships??
If you don't care about that why wouldn't you use the flagships??
I'm arguing against the idea that Jev fixes a fundamental deficiency in LLMs by showing it's confidence.
It doesn't.
Jev can do a strict subset of what LLMs can do, but incredibly quickly and incredibly cheaply, which is certainly valuable.