I don't think that makes it a dead end. Laya is designed to be fine-tuned for a specific task, and the site reports the fine-tuning gains. The 0.766 number comes from fine-tuning on the benchmark's train split, not from the base checkpoint. They also report that fitting a single temperature scalar per question type cuts expected calibration error from 0.466 to 0.081. That's a large gain, and it only shows up after you specialize the model.
"Peer" is doing a lot of work in that comment. Jev can take on a new task without retraining because it starts with far more knowledge; Laya trades that away to stay small and trainable for a fixed task. So I wouldn't compare base Laya to Jev and stop there. Compare Jev to fine-tuned Laya on the same task and test set, then look at accuracy, latency, cost, calibration, and robustness, depending on which of those matter for the deployment.
My main point of disagreement would be that I fundamentally see a different use for a jev sort of model (generalism is applealing), but if you're finetuning, a bert base is not bad.
Anyways, thanks for the vouching!
However at this point I talk to LLMs more than anyone except probably my wife. As a multiple times immigrant, I can absolutely believe I'm adjusting my speech patterns to its vernacular.