For the same reason why it might produce a bogus answer: it's a language model, not a general AI. Prompt it for footnotes and it will generate footnotes for books that don't exist or that don't actually contain the citation. They'll match the
form of a footnote, because that's what the model was trained to do, but the citations won't resolve.
For example, I asked who the first president of Argentina was. It gave the wrong name (Juan Manuel de Rosas) with a half-true bio. I then asked it for books where I can learn more about this person.
It produced five books. None of the book-author pairings were correct, though three of the books were real and four of the authors actually wrote books. Of the four authors that existed, three actually did write about Argentina's history, while one is a professor of Spanish and gender studies.
This is my point: ChatGPT was trained to produce text, not to produce facts. What I'm seeing here isn't just something that can be optimized away, it's fundamental to the model's design. It was basically told to produce text that looks like it could have been written on the internet. Its text is really convincing: when I saw the list of books I thought for a minute it had actually done it! But it's still just inventing text that reads well, because that's what its training was optimizing for.
It could be paired with a different system that searches for facts, but then it's just a friendly abstraction layer on top of Google or some other system, not a revolution in search.