HNHacker News
TopNewBestAskShowJobs

cannedbread

5 karma · joined October 9, 2018

submissionscomments
cannedbread··on How accurately calibrated is Jev?
I think you hit on my core criticism of this and similar analyses: LLMs are a tool, and using them to draw from a distribution is misuse of the tool. It's not what it's designed/tuned for. Asking follow ups like "is Jev calibrated when I mis-use it?" is asking the wrong question. The right question would be, "is Jev calibrated for expected use cases?". And based on some initial exploration, I do think Jev is calibrated for common natural language questions
cannedbread··on How accurately calibrated is Jev?
IMO by asking Jev underspecified questions like this, you're essentially using it as a random number generator (similar to the dice example). On actual NLP problems (including ones with uncertainty under human review) it does appear to be well calibrated: https://leonardgrazian.com/blog/jev-calibration/
cannedbread··on Launch HN: InspectMind (YC W24) – AI agent for reviewing construction drawings
When I upload my drawing set, how often should I expect it to hallucinate? And how much of the real stuff does it flag?