>
It lectures me in a quite self-satisfied tone that I should commence to use the function "AXObserve" that does not exist, as it turns out. This is completely made up. There is no such thing. If you google the exact word, it doesn't even bring up anything related to programming let alone the macOS APIs.I've seen this happening too, but it still feel more like a case of a generalized "there is something like this, or at least should be". In some cases, it was clear to me it was guessing at some reasonable patterns - like, upon seeing function calls like InitializeFoo() or AddBar(), it would assume existence of functions UninitializeFoo() and RemoveBar(). But in other cases, more similar to yours, it does invent stuff that doesn't exist (though perhaps should).
I find the latter most common when asking for Emacs instructions or Emacs Lisp code. GPT-4 is prone to inventing functions based on the prompt, such as e.g. `org-timestamp-diff-day' (or something similar), when I asked it for the command that computes difference between Org Mode timestamps in days. This function does not exist, but there are a few functions like `org-timestamp-[something]', and even a few like `org-timestamp-[sth]-day' - and more generally, Emacs/ELisp code tends to be full of functions named like the very thing you're trying to do, like `kill-whole-line' or `move-end-of-line', `kill-comment' and `duplicate-line', so it's not that surprising GPT-4 would, the other day, give me something like `kill-whole-comment-and-duplicate-line'.
> Now I will admit that my mind is capable of doing the same. Under severe intoxication :)
Mine too. It's not a bad analogy. I've observed in myself that moderate intoxication does magic when it comes to letting my "inner voice" drive the conversation. I sometimes experienced the feeling of my conscious mind being too late to catch sentences produced by my "inner voice" before they exited my mouth.
Another analogy I particularly like is, GPT-4 is actually quite like a ~4 year old kid in many ways. One that has somehow consumed half the internet, but still with child-like attention span and propensity to continue talking and making shit up even long after crossing the point of lack of knowledge.
My daughter turned 4 lasts week. The way she makes stuff up is quite similar to LLM "hallucinations" (except it's easier to tell, because she's only heard so much in her life, so she's less good at making up stuff that sounds believable). And earlier than that - between the age of 2.5 and 4 - I could observe what's best described as a "rolling context window" growing. A little more than half a year ago, I could tell hers is about 30 seconds long - anything she invented and said once would be forgotten after about that time, unless it was repeated in between by someone, possibly herself. And believe it, she talked in an extremely repetitive way back then - ~every third sentence was 50-100% repeating the critical things said earlier - as if, intuitively, she was trying to overcome her own "context window" limit.
> Maybe we could agree that LLMs can probably serve as part of a foundation for a future model that actually can output a result of some "real" understanding of a subject?
Sure. I suppose we should also agree on the same understanding of "real understanding". I'm only proposing that LLMs pick up the same kind of "understanding" your unconscious/subconscious mind does, and produce output of similar nature. This implies that, to replicate human reasoning/performance, we'll need to layer some additional models/systems on top of the LLM.