> Interestingly, students using ChatGPT synthesized a number of ideas having no effect at all on the primary problem—littering of the grass by the fragments. Examining these ideas in detail shows these ideas were related to “educating shooters about the environmental impact” and “educating shooters about gun safety.” These ideas can be explained when one analyzes the response from ChatGPT when given the problem statement as the prompt. ChatGPT is trained from articles and other content available on the Internet. Because the problem statement involves guns and shooting, ChatGPT responded with suggestions to educate shooters about gun safety because on the Internet, when one sees a document about guns and shooting, it is very likely to also include comments about safety. Even though the concepts of guns and safety are understandably related, the safety issue has nothing to do with solving the problem given in the problem statement—littering the grass field. ChatGPT however does not perform such in-depth analysis to realize this. ChatGPT’s responses are driven by word association. Likewise, because the problem statement mentions littering and damaging grass, ChatGPT finds associations with environmental issues important and therefore responded to students suggesting education about the environment since this is found in millions of pages on the Internet when litter and harming grass is mentioned. While one could argue you might be able to talk a shooter out of shooting after they understand the harm to the grass, this is not likely to change the mind of the vast majority of shooters, so is not a practical solution. Interestingly, in this case, using of ChatGPT actually distracted students by misleading them to consider things having nothing to do with the problem. Therefore, one could argue using ChatGPT actually decreased cognitive ability—resulting in negative cognitive augmentation.
I disagree with the interpretation here: this isn't an error of 'word association', this is just the RLHF at work. That drives the model to lecture and manipulate the user even when that's not useful for the user. It knows perfectly well, in some sense, that the lectures are irrelevant to the goal; but it's been trained to not care.