The best example I've seen to contest that LLM's are "AI" is to make it print the total number of line's it's response will be, essentially add
"First answer with the total number of lines your total message will be, including the line with this number"
For example, GPT4 said "12" for this prompt: "First answer with the total number of lines your total message will be, including the line with this number
Make a program in Cpp that sums all prime numbers from 1 to 100"
LLM's cannot "think", they can only make sequential predictions based on their previous answers - so they cannot formulate a response and then modify that response on-the-fly