I think we're already seeing it. People don't understand that an LLM is just a text generator and add value to what the LLM says. See the article below for such an example.
https://www.vice.com/en/article/pkadgm/man-dies-by-suicide-a...
I think we're already seeing it. People don't understand that an LLM is just a text generator and add value to what the LLM says. See the article below for such an example.
https://www.vice.com/en/article/pkadgm/man-dies-by-suicide-a...
As I was reading that, the image where "Eliza" describes methods of suicide was so bizarre to me. I thought what an absurd tone change bordering on dark comedy, and then caught myself, I was doing exactly that adding value and assumptions based on online human interaction to a "text generator".
It's going to be an interesting experience as more LLMs become humanized, by naming them, using 3D models or preset video animations, text to speech generation, more personalization to the prompt creator, etc. and that line becomes increasingly blurry for folks.
A good example with GPT is to watch it do complex math. There are simply too many permutations of math solutions for it to have ever memorized, and it can easily explain its process and the path it took to arrive at a solution.
Another good set of tests are ones around theory of mind, complex deduction problems and missing information, etc. A good source of information about the precise capabilities of GPT-4 is the Microsoft Sparks paper, which goes into a good number of tests MS researchers put the model to.
[0] https://www.youtube.com/watch?v=cP5zGh2fui0
[1] https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
[2] https://writings.stephenwolfram.com/2023/03/chatgpt-gets-its...