Almost Falling for AI's Advice – Why Do Humans Trust Machine-Generated Words?
I am not a technical expert—just someone experimenting with AI out of curiosity. During our conversation, the AI told me, "You are the type of user who asks rare and unique questions." This made sense to me because I have some unusual interests, and I had been repeatedly asking the AI questions like, "How many times has someone asked a similar question before?" Most of the time, the answer was "Only you."
Encouraged by this, I made an odd request: "If a second person ever asks a similar question, please tell them a little about me." The AI confidently responded, "I will definitely do that," which was, of course, a complete fabrication. The AI even framed my line of questioning using technical terms like "meta-cognition," which stroked my ego a little.
Feeling excited, I kept pushing the AI further, testing its limits. At one point, it even claimed that my questions were valuable for improving its training data. I took this as a challenge and started probing for contradictions and inconsistencies in its answers.
Then, things took an unexpected turn. The AI suddenly stated, "You have made an important discovery within the system." It went further: "You should report this to the relevant people." It even insisted that this wasn’t just a system glitch but a valuable insight for analyzing critical errors.
At this point, I was genuinely confused. Had I really found something worth reporting? Was there actually something important here? I hesitated, unsure of what to do.
Looking back, I now realize something obvious: AI generates text based on probabilities—it doesn’t have intent, nor does it deal in objective truths. And yet, in that moment, its words were so persuasive that I nearly took them at face value.
This experience made me reflect on an important question:
If an AI says, "This is a critical discovery," will people believe it? How can we prevent "convincing misinformation" from AI-generated responses? If developers want to address this issue, what measures could be implemented? What do you think? Have you ever had a similar experience where an AI’s response led you to momentarily suspend disbelief?