I think you hit on my core criticism of this and similar analyses: LLMs are a tool, and using them to draw from a distribution is misuse of the tool. It's not what it's designed/tuned for. Asking follow ups like "is Jev calibrated when I mis-use it?" is asking the wrong question. The right question would be, "is Jev calibrated for expected use cases?". And based on some initial exploration, I do think Jev is calibrated for common natural language questions