How is calibration of Jev or Jev-inspired models being evaluated?
I tried this with Bobcat, a Jev-style model I trained on Qwen. A single temperature lowered development ECE from 0.021 to 0.010; its sealed final ECE was 0.012. That is calibration against our own labels, not a measured claim about Jev's calibration. On TypeSafe's 20 published workflow examples, we can compare probability assigned to a reference answer, but that reference is model consensus, not ground truth, and the sample is too small for a general calibration claim. Details and the evaluation code: https://huggingface.co/sanghwa-na/bobcat-1.1