Thanks. How reliable is the world knowledge? Say, I use it as a delegation router for picking the best model for a task - how can I be confident it's doing that with enough intelligence?
When I use traditional LLMs I get some sense from the flagship-ness and regular usage. How do we get such confidence for Jev like models?