Edit: I don't think its exaggerated and I think its important .
1. they invented a new disease and published a preprint (with some clues internally to imply that it was fake)
2. asked the Agent what it thinks about this preprint
3. it just assumed that it was true - what was it supposed to do? it was published in a credentialised way!
It * DID NOT * recommend this disease to people who didn't mention this specific disease. Edit: I'm wrong here. It did pop up without prompting
It just committed the sin of assuming something is true when published.
What is the recommendation here? Should the agent take everything published in a skeptical way? I would agree with it. But it comes with its own compute constraints. In general LLM's are trained to accept certain things as true with more probability because of credentialisation. Sometimes in edgecases it breaks - like this test.