In most ML cases (and ChatGPT likely), "confidence" would generally just correlate how closely the query matches data and patterns it's seen in its dataset and inversely correlate with how many conflicting matches and patterns it sees.
Humans are subject to the same problem of course. If you asked how confident a person living many centuries ago was that the Earth was flat, they'd probably say "very confident" because there was nothing in their training data / lived experience to conflict with that view. But they'd be wrong.
But humans still have a significant advantage in that they report lack of confidence when they sense logical inconsistencies and violations of reasoning to a level that ML models can't (at least not yet).
Maybe a fan-out of the possible ways it could answer would be interesting, but really we more need a disclaimer next to every answer that says "this thing that's answering in fully formed language does not have human reasoning capability and can't be trusted (yet)"
https://en.wikipedia.org/wiki/Myth_of_the_flat_Earth
"The earliest clear documentation of the idea of a spherical Earth comes from the ancient Greeks (5th century BC). The belief was widespread in the Greek world when Eratosthenes calculated the circumference of Earth around 240 BC. This knowledge spread with Greek influence such that during the Early Middle Ages (~600–1000 AD), most European and Middle Eastern scholars espoused Earth's sphericity.[3] Belief in a flat Earth among educated Europeans was almost nonexistent from the Late Middle Ages onward ... Historian Jeffrey Burton Russell says the flat-Earth error flourished most between 1870 and 1920, and had to do with the ideological setting created by struggles over biological evolution"
I asked ChatGPT the same question and it prevaricated:
"There is evidence that some people in medieval times believed the Earth was flat, while others believed it was round. The idea that the Earth is round, or more accurately, an oblate spheroid, has been around since ancient times. The ancient Greeks, for example, knew that the Earth was a sphere. However, the idea that the Earth is flat also has a long history and can be traced back to ancient civilizations as well. During the Middle Ages, the idea that the Earth was round was not widely accepted, and there was significant debate about the shape of the Earth. Some people continued to believe in the idea that the Earth was flat, while others argued for a round Earth. It is important to note that the medieval period was a time of great intellectual and scientific change, and ideas about the shape of the Earth and other scientific concepts were still being developed and debated."
But from what I know, it's wrong, at least as far as we know the historical record (of course there may have been peasants who believed otherwise but their views weren't recorded). The fact that the Earth is a sphere is obvious to anyone who watched a ship sail over the horizon, which is an experience people had from ancient times.
For example:
> I'm going to share some information, I want you to classify it in the following JSON-like format and provide a responses that match this typescript interface:
> {
> "isXXXXX": boolean;
> "certainty": number;
> }
> where certainty is a number between 0 and 1.However, I got either 0 or 1 for the certainty every time. Not sure if it was because they were either cut-and-dry cases (certainty 1) or not-enough-information (certainty 0).
I'm actually trying to think of a good example of text I could ask it to intuit information from and give me a certainty
For example, ask it to subtract 2 20-digit numbers. It will come up with an answer X where the first couple of digits are correct, and everything after that is wrong.
It gets better.
Ask it to correct itself. It will come up with a different wrong answer Y.
If you then ask it to explain why the answer is right, it will give you an explanation. At the end of the explanation it states the answer is X again, and then in the very next line concludes by telling you that is why the answer Y is correct. :)
It's apparently really hard to objectively measure/report the "truthiness" of LLM results
Allowing an LLM to "improvise" and be a bit fast-and-lose is unfortunately a necessary ingredient in how they currently work.