HNHacker News
TopNewBestAskShowJobs

medicalthrow

128 karma · joined October 14, 2025

submissionscomments
medicalthrow··on GPT-5o-mini hallucinates medical residency applicant grades
It's a bit complicated. Each school has their own grading system (some pass fail, others four tiered, others full letter grades). Additionally, there are reported distributions for each grade. Lastly, there's sometimes a summary statement at the end that usually says "X student was 'superlative'" and then a table at the end that says 'superlative' means top X% of class. On top of that, students may not get their full dean's letter that says all of this stuff. Basically, self reporting is very difficult to do given the amount of variability in grade reporting.
medicalthrow··on GPT-5o-mini hallucinates medical residency applicant grades
Hi HN, submitting from a burner since I'm an applicant this current medical residency admissions cycle. I thought it was interesting to show the real world implications of using LLMs to extract information from PDFs. For context, thalamus is a company that handles the "backend" for residency programs and all the applications they receive (including handling who to invite for interviews, etc). One of the more important factors in deciding applicant competitiveness is their medical school performance (their grades), but that information is buried in PDFs sent by schools (often not standardized). So this year, they decided to pilot a tool that would extract that info (using "GPT-5o-mini": https://www.thalamusgme.com/blogs/methodology-for-creation-a...). Some programs have noticed there is a discrepancy between extracted vs reported grades (often in the direction of hallucinating "fails") and brought it to the attention of thalamus. Unfortunately, it doesn't look like the main company is discontinuing usage of the tool.

Regardless, given that there have been a number of posts looking into usage of LLMs for numerical extraction, I thought this story useful would be a cautionary tale.

EDIT: I put "GPT-5o-mini" in quotes since that was in their methodology...yes, I know the model doesn't exist