Same finding from a different angle. Zero-shot Laya tracks Jev on 2-4 labels and falls off a cliff on 77 (38% vs 76% on banking77 in jevbench). What closed the gap for me was not a bigger model but a small head trained on the task's own rows on top of the frozen encoder: a 12-label intent site went 89.5% -> 100% agreement with its teacher on 3000 rows, support triage 69/34/66% -> 99/86/98%, with the head answering 90% of requests at a 0.99 agreement target and the rest falling back to the provider. The catch is that a frozen encoder learns what the text says, not arithmetic over fields: the same risk rule scored 0.42 as "amount > X" and 0.94 restated in words. I packaged the loop (record -> train -> shadow -> live with fallback) as a proxy that also speaks the Jev API:
https://github.com/bladedevoff/stuntd