There are benchmarks for these things, which Jev and others have been put through; it's not a marketing boast!
There are benchmarks for these things, which Jev and others have been put through; it's not a marketing boast!
* With jev-sec-bench the confidence numbers reached roughly the same as Qwen2.5-7b-instruct (confidence was off by about 6% on average in both cases).
* On the OpenRouter Banking 77 benchmark Jev did the worst out of all the models, being about 25 points off from reality.
I also think the fact that Jev hasn't published any real papers or actual technical details should make people a bit more skeptical than they have been.
If you don't care about that why wouldn't you use the flagships??
I'm arguing against the idea that Jev fixes a fundamental deficiency in LLMs by showing it's confidence.
It doesn't.
Jev can do a strict subset of what LLMs can do, but incredibly quickly and incredibly cheaply, which is certainly valuable.