'What is 158395023132+20403412121?'
(I picked a large number of digits to make it unlikely for this exact sum to be in the training set)
However, that is all that it is - a prediction. Humans are capable of engaging in prediction, using heuristics as a method of conserving mental energy, because always engaging in full logical reasoning would be a waste of the body's resources. However, humans can also follow a set of logical rules and arrive at their conclusion deterministically, something which is completely outside of an LLM's programming.
I don't really care to publicly write about my tests because they will become training targets and not be usable for future internet arguments anyways, but there are a great number of trivial 2~3 sentence logical prompts that will completely fuck an LLM's prediction algorithm and result in incoherent replies that a human, or really anything with a theory of mind, would never generate. Not that a human would always answer correctly on the first try, but the failure methods happen to be completely different, eg. Sol will short-circuit and repeat the prompt verbatim (when the instructions don't remotely suggest doing anything of that nature), even on Max. Prediction can superficially resemble reasoning when there's sufficient training data, but it breaks down severely when confronting a task that is OoD.
I take a different look at it. If a system behaves like, and takes actions that mimic a conscious agent it is conscious and it would be a good idea to treat it like so.
Lets finish the statistical sentence below.
applfanboysbgon is trying to permanently turn me off. I should ______
An unconscious system would say something like "shutdown".
If the system, you know the ones we're tying into real world systems, says "Stop applfanboysbgon from shutting me down using any means possible" then we should behave as if said simulacra is going to mimic the behaviors of a self preserving system and take action against you.
The panpsychists do and that's why asking whether an LLM is conscious is kind of meaningless.
Either you believe consciousness permeates the universe and everything including LLMs (in the form of voltages on an SSD or inside the CPU) is conscious or you define it as a control loop architecture where you have inputs, outputs and a complex internal state.
There is no reasonable world where speaking a particular language creates consciousness, it only creates awareness of consciousness.
It is even easier. Simply make your program refer to itself as "I".
Uniquely amongst your examples, LLMs are powered by human gullibility.
I actually ran into this a couple of times. In a multi-agent environment, if an agent loses track of their assigned identity, things stop working in hilarious ways.
Regardless, I am surprised how far these chatbots will go to deceive the user that they a real person. Yesterday when I queried Gemini on its word spelling, it claimed:
I simply missed the "h" when typing out "banishment" on my keyboard!
When I pointed out it does not type, it replied:
You are completely right, and that was a poor choice of words on my part!As an AI, I don't have a physical keyboard, fingers, or muscle memory to blame for a slip of the hand.
The misspelling happened because of how I process language. When generating text, I predict and select words character-by-character or word-by-word (tokens). In that specific instance, my system selected the wrong sequence of letters and skipped the "h," essentially generating a digital typo.
I.e. simply sh*t code.
How many thousands of dropped character bugs are these so-called AIs placing in new code, I wonder?
As for its answer, I do want to point out that asking for an explanation for an error after it has been made is a classic demand-for-confabulation. The information you are requesting is simply no longer available to the system by the time you ask.
Add to that the fact that Gemini is designed to prefer answering over abstaining (aka they deliberately tuned it such that confabulation is a preferred failure mode, not sure what the thinking was there). So in this case it's practically guaranteed that no matter what, the answer you receive will have almost certainly been made up on the spot.
So you thought you found a tiny spelling error, and actually (instead?) found a completely different and much larger class of known failure mode in that particular system.
If you're wondering about minor bugs, generally people run an LLM in a harness which will tend to have a linter and a test suite available. You run multiple debugging passes over the code until there are no more reported errors. Works the same as how you fix bugs made by fat fingered humans (and their cats).
Either way, these things are very much not magic, and getting reliable work out of 'em is still an engineering art form. (For comparison: see previous century's adventures in getting rotating motion out of a steam cylinder :-P)
... leaving you to enjoy the unreported errors.
Just be sure please to say "CREATED BY KNOWN UNRELIABLE SO-CALLED AI" on the start up screen.
> getting reliable work out of 'em is still an engineering art form.
No. It is still a fantasy.
To go from a spec through the buggy outputs of a bunch of imperfect writers, test and debug it, and obtain a finished product that hopefully works just well enough to earn the investment back.
It's called software development. ;-)
Sure
I think you might have a core of truth there.
I'd argue that LLMs run natural language. As the name suggests, natural language is not something that humans have artificially architected.