Thanks for the link!
I've now taken the time to watch John's talk and I have some thoughts. It's not only difficult to solve truthfulness due to disagreement (and subjectivity), but it's also very difficult because different contexts have different standards of evidence.
In some contexts, like programming, we'd rather have the model output its best guess for what the program should be, no matter how low the confidence, because we would like a starting point and we can debug the program from there. The answer "I don't know how to write that program" is not a useful starting point and it may even be an example of the model withholding information it does have due to low confidence.
In other contexts, such as scientific or historical questions, we want a high standard of evidence. Asking the question "what year did Neil Armstrong land on Mars?" should not produce a hallucinated response with fully unhedged language complete with fictitious date of landing. This problem may be solvable by training the model to hedge or even to question the premise when the confidence is low. Of course, this also suffers from the garbage-in-garbage-out problem of having falsehoods buried in the training set.
A more subtle and difficult problem with scientific/historical questions is with long-form answers. Currently, models tend to produce long-form answers that fairly consistently contain a mixture of true facts and falsehoods, and it can be quite difficult for even expert readers to spot all of the mistakes every time. Furthermore, the human labellers were given very sophisticated tools for highlighting sentences in long-form output but the information this produced had to be reduced down to a single bit per example since the detailed information did not improve training very much.
Personally, I think it's going to be very difficult to teach the model how to recognize the appropriate contexts and associated standards. This is a very subtle problem and one of the issues is that it relies on information the model does not have access to, for example: the identity of the question-asker. If a child asks an astrophysicist about black holes they're going to get a different answer than if an undergraduate student asks the same question in class. Yes, this additional context can be included in the prompt, but at some point it becomes a pain to have to copy-and-paste the context for every prompt.
Perhaps people will create a tool to save this additional context in the form of presets but this imposes additional effort on humans. At some point I think the amount of human curating and feedback that goes into these models will cause a collapse and backlash. We saw the same thing happen in the early days of search engines, when Google (fully automated) trounced Yahoo (human curated), leading to Yahoo's abandonment of human curation. We also see the same problem manifest itself at the Patent Office, where human review is policy. The entire patent system has become grossly dysfunctional at least partly due to the overwhelming complexity of this problem.
One thing I really liked was the "inner monologue" of the model performing a sequence of steps to answer a question by doing a search. If this could be generalized to other tasks it could be a home run for automated assistants (Google/Alex/Siri).