I consider code to be non-fiction, and ChatGPT will generate stuff that compiles and works most of the time. There's no need to triple-check the code output.
"Code" is really a much much much smaller and much much much more structured output than "English words".
Presumably, the system was trained with a very small amount of "untrue code" in the sense of stuff that just absolutely could never work. And also presumably, it was trained with a lot of free-form text that was definitely wrong or false, and highly likely to have been originally created to be purposefully misleading, or at a minimum, fiction.
That the system outputs reliable code tells us nothing about its current ability to output highly reliable free form text.