I hear what you're saying. But this is not such a model. The article says that various internal teams had to fine-tune Clef to suit their specific needs. They're even rolling out a RL training product line for it.
Fair. Though in this case my actual classifier (dino v3 features fed into a very shallow DNN) works surprisingly well and is insanely faster than using an LLM -- so it really was the one shot performance I was hoping for.