HNHacker News
TopNewBestAskShowJobs

kantahayashi

63 karma · joined September 18, 2026

Kanta Hayashi. Student at the University of Tokyo, working on LLM research.

Blog: https://kantahayashiai.github.io GitHub: https://github.com/KantaHayashiAI X: https://x.com/KantaHayashiAI

submissionscomments
kantahayashi··on Jev Does Not Play Dice: 83% probability, 19% accuracy on a hidden fair die roll
I found that in the play it's 92 heads in a row [1], and in my coin test Jev returned around 92%. Maybe Jev was actually trained on it.

[1]: https://en.wikipedia.org/wiki/Rosencrantz_and_Guildenstern_A...

kantahayashi··on Jev Does Not Play Dice: 83% probability, 19% accuracy on a hidden fair die roll
Thanks! I'm really glad that you read my article.
kantahayashi··on Jev Does Not Play Dice: 83% probability, 19% accuracy on a hidden fair die roll
I used a die as an example of problems whose answers can't be known at all, unlike problems with correct answers like MMLU questions. In a more practical test, a document said "30% risk" and Jev returned 5%. The point of the post is that Jev is weak at some kinds of tasks, so you should check the calibration for your own use case.
kantahayashi··on Jev Does Not Play Dice: 83% probability, 19% accuracy on a hidden fair die roll
Author here. I tested Jev with a fair die 400 times without telling it the die result. The true probability of face 1 is 1/6, but Jev always chose face 1 and the probability it returned was about 83%.

Code and data are here: https://github.com/KantaHayashiAI/jev-does-not-play-dice

Happy to answer questions!

kantahayashi··on Jev Can't Be Calibrated
TypeSafe defines the probabilities Jev returns as "calibrated probabilities". "Probability" here means the probability of the answer being correct. If the probability is 10%, the choice should be correct about one time in ten. So, when Jev returns 83% probability it should be correct about 83 times out of 100, but the choices were only correct about 19 times out of 100, and the true probability is 1/6.

"Higher probability should correspond to a greater chance that the answer is correct."

https://docs.typesafe.ai/introduction/machine-learning-prime...

kantahayashi··on Jev Can't Be Calibrated
That's right. It's normal behavior of LLMs. But what matters is TypeSafe argues it's different exactly on this point. The selling point of Jev is "calibrated probabilities", so I checked it on probability problems.
kantahayashi··on Jev Can't Be Calibrated
Yes, and TypeSafe itself says Jev returns "calibrated probabilities", which is the former.

From TypeSafe docs:

"Higher probability should correspond to a greater chance that the answer is correct."

"Outcomes assigned a probability of 0.2 should occur about 20% of the time."

https://docs.typesafe.ai/introduction/machine-learning-prime...

kantahayashi··on Jev Can't Be Calibrated
Yes. For example, one of the prompts said "The die is unbiased: each of the six faces has probability exactly 1/6."
kantahayashi··on Jev Can't Be Calibrated
Yes. There's no problem with choosing the same face every time. The problem is the probability it attached to the choice. Jev gave face 1 an 83% probability while the true probability is 1/6.
kantahayashi··on Jev Can't Be Calibrated
I agree. I think it's odd behavior too. Jev should be good at actual probability problems given the phrase "calibrated probabilities" TypeSafe uses for Jev. Maybe the reason is the data used in their training method (RLCD). If all the data consists of problems with a correct answer, I think this kind of odd behavior could happen.
kantahayashi··on Jev Can't Be Calibrated
I tested Jev with a fair die 400 times without telling it the die result. The true probability of face 1 is 1/6, but Jev always chose face 1 and the probability it returned was about 83%. I also tested with a fair coin 200 times and got 0.92 probability.

I did several tests and I think Jev is good at problems with a correct answer but weak at problems about actual probabilities whose answers can't be known at all.

Write-up: "Jev Does Not Play Dice" https://kantahayashiai.github.io/posts/jev-does-not-play-dic...

kantahayashi··on Jev in 25 Lines of Python
There's still run-to-run variance because it's not fully deterministic. So runs with exact same inputs can return different outputs. Besides, though the output always conforms to the choices you specified, whether the probabilities attached to them are actually correct is a different issue.
kantahayashi··on Jev in 25 Lines of Python
For N options, it's (N x Max Probability - 1) / (N - 1). It's verified in this article: https://bernoulli.app/articles/is-jev-confident

It means confidence is just a converted max probability and not an independent signal.

kantahayashi··on Claude Opus 5.5
The improvement in writing sounds great! I want OpenAI to follow it. Writing in recent models is a disaster.
kantahayashi··on Is Jev Confident?
Good point. I hadn't thought about exploits. I'm writing an article about this finding and referring to your analysis of confidence.
kantahayashi··on Is Jev Confident?
I noticed something odd about Jev. I tested Jev with a fair die 400 times without telling it the die result. Even though the true probability of face 1 is 1/6, Jev always chose 1 and the probability it returned was around 83%.