But I've seen nothing to indicate that the upper bound on classification tasks of a Jev-like model can exceed a frontier LLM with reasoning tokens. That seems nearly impossible even in principle (since Jev-style models are still based on LLM pretraining).
So while they're definitely on the Pareto frontier, which is valuable, they're at the "cheap" end of the spectrum more than the "good" end and I don't expect that to change.