> The size of the training data far exceeds the size of the model's weights.
I'm not expecting strict recall but I imagine there is a non-insignificant amount of text in the training data about cysts relative to other entities in the "medical genre" as cysts of all form are one of the most common medical conditions and would be discussed, so I would have expected more probable sequences than the one generated.
Doesn't this alsois then also raises the question of what parameter-corpus size works best? At 120b parameters to 106b tokens Galactica was still hallucinating quite a bit.
> What do you mean? It got _very_ close to the correct answer. One useful thing to know is that GPT-3 cannot "go back" and correct early mistakes. This may have happened here.
I meant that with the negation it became a physical impossibility and therefore very incorrect, but if not negated it would be a correct statement. Your explanation sounds right in this instance, at some point it decided to negate the sentence and it went from being correct (although not relevant to the prompt) to an incorrect statement.
This also suggests to me that ChatGPT doesn't have a good enough understanding of negation. This is a challenge for many models in the medical domain as the frequency of negated statements is much higher than in general texts, but very intuitive for any human.
I think part of the problem is ChatGPT works so well most of the time I get surprised when it fails in seemingly obvious ways, granted to someone with expertise in the field. It's interesting to probe.