LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes
arxiv.org
arxiv.org
EDIT: And Pangram agrees that the abstract is 100% AI generated.
I know arXiv has taken some measures to combat spam like this, but it seems like they’ll need to do more. There’s just very little barrier now to creating giant slop papers like this and then dumping them anywhere that won’t reject them. It is an insult to everyone’s time, and I can’t imagine they expect people to actually read this. If the expectation is that everyone will use an LLM to interpret it, then maybe they should have at least had a few more rounds of tightening and polishing the paper, even via LLM, to save the redundant token use.
The whole thing is nasty partly because it isn't just AI problem. When checking the human labels that we used in our evals 40 out of 142 answer keys claiming absence were wrong. Tricky one.
Do you have an example?
I'm in healthcare and ambient documentation is obviously a huge thing now but I don't have any experience with it. We have anywhere from 5-10 companies reach out a week trying to sell us on their product and the demos are mostly okay (though you can tell they're rely on the happy path through a lot of it), but we haven't actually pulled the trigger on anything. Thanks in advance.
Receipts where there's no subtotal line, only a total. One model returns a subtotal anyway, "63.000", "88.000", numbers it copied or computed from elsewhere in the document. Ours invented a 2% discount on a receipt that has no discount line.
TV ad contracts where the station's address isn't on the contract. Several models fill it with a real address that is on the page, just the agency's or the advertiser's. Given your industry this pattern should worry you most, since these aren't made up from nothing. Just the model tripping and taking it from the wrong place, which might look plausible on the surface.
Here's what I'd suggest for your vendor demos. Run them on a handful of your own documents where you know a field is genuinely absent, and count how many come back filled. Their own samples won't tell you that.
All the raw outputs are here if you want to look: https://velrim.com/research/fabrication-on-absent-fields
TV ad contracts where the station's address isn't on the contract. Several models fill it with a real address that is on the page, just the agency's or the advertiser's. Given your industry this pattern should worry you most, since these aren't made up from nothing. Just the model tripping and taking it from the wrong place, which might look plausible on the surface.
Here's what I'd suggest for your vendor demos. Run them on a handful of your own documents where you know a field is genuinely absent, and count how many come back filled. Their own samples won't tell you that.
All the raw outputs are public, so you can check the examples above yourself. My reply to you got filtered so I'm omitting the link. If you want to read the writeup, search for "velrim" and "fabrication on absent fields" for the article plus the repo.
People using AI like this should be run out of their jobs.