Analysis: ChatGPT is great at what it’s designed for. You’re just using it wrong
pbs.org
pbs.org
It's amazing at giving general answers and outlines though, ask it what you should take with you on a trip and it can give you something you wouldn't have thought of.
Most claims about chatGPT's behavior do not match its actual behavior.
Let's consider the claim on chatGPT's home page:
> The dialogue format makes it possible for ChatGPT to answer followup questions, admit its mistakes, challenge incorrect premises, and reject inappropriate requests.
This statement is correct from the perspective of utility: ChatGPT does, at the end of the day, exhibit these behaviors successfully.
But there's a problem here: utility isn't the perspective that the average reader is likely to infer with. In fact, they are expected not to!
The expected perspective is ability. After all, we are talking about what ChatGPT is able to do.
This is the most subtle, yet most important distinction. ChatGPT is, itself, entirely not able to do any of these things!
But how then does it exhibit these behaviors? If it isn't doing the work, who is, and how (and where) does ChatGPT leverage that work?
To answer that question, let's talk about what our expectations are for what "doing" means.
Traditionally, to do a thing, you must identify the subject (the thing), and apply a verb (do). ChatGPT is explicitly capable of neither.
ChatGPT does not recognize subjects. ChatGPT does not implement arbitrary behaviors from the meaning of verbs.
Before ChatGPT can exist at all, let alone exhibit behavior, it must be trained on a dataset. Here's what that looks like:
ChatGPT's neutral net is fed a stream of text. That text is tokenized into short groups of characters. This is our first divergence from expectation: why not tokenize words and punctuation? Because ChatGPT explicitly does not interact with words or grammar. It does something else entirely, and short groups of characters are a good structure for that thing.
Next, the groups of characters (tokens) are organized as "neurons". So far, it's a graph without edges.
Next, duplicate nodes are found and linked together. This creates the foundational model that will later resemble (but not be) a knowledge graph.
Next, duplicate patterns of nodes and edges are recognized, and connected together. Same as the last step, just higher order. Our neutral net model is complete.
Finally, the neutral net is trained. It's given a series of tests, each with an expected result. A successful test is weighted positively.
At no point did we define categories like subjects, grammar, or verbs. We didn't even define words. What we did define is a model of how text relates to itself.
And that's where the magic lies. What does text mean to itself? Literally everything. Words and grammar are the symbolic representation of ideas, and that representation is explicitly defined.
The only reason that isn't enough on its own is that natural language is imperfect, and full of ambiguity. In other words, natural language lies to itself about itself.
Humans compensate with logic and context. We explicitly define the meaning of text by building up to it with backstory (context), and when we are uncertain, we filter out any construction that isn't logical.
By providing a dataset of text that was logically constructed by humans, modeling how that dataset relates to itself, and explicitly training a preferred continuation whenever a human recognizes a bad continuation, what we end up with is a fuzzy algorithm that implicitly maps the semantics of an expression to the most familiar semantic continuation.
Because semantics are very nearly a 1-to-1 mapping from symbol to meaning, the emergent effect of this system very nearly mirrors the logic-based interaction of human writing.
But every ambiguity that exists in language is also an emergent effect. It becomes a lie or a mistake. The logical relationship that was captured by symbolic representation when a human wrote it is broken. Natural language has lied to itself about itself. A human could recover, but ChatGPT must be explicitly trained. This is done either by explicit training (to define a preference/expectation), or by filtering the original dataset to be less ambiguous.
Language itself can also recover, by providing explicit clarification after the fact. This is implemented often in technical writing.
So what can ChatGPT do? Generate a semantically likely continuation from its training dataset.
What can a semantically likely continuation be? Either a logically coherent response or a semantically coherent lie.
How hard is it to guarantee the former? Extremely hard.
But hey, things can just work and the golem can overcome it's creator or surprise it. All of the things you said don't matter.
The writer of this post is just another ape sitting with his phone in his hand making funny gestures on a touch screen. He got to the point of communicating with you after hearing similar apes make similar sounds, somehow connected those sounds to funny shapes as a little ape, and then learned to click those squiggly shapes as they show up on his "phone" in belief that another ape somewhere will recognize them.
Some apes do this so that other apes will click a triangle. But you can't guarantee that other apes really mean what you think those squiggly lines really mean, or that they aren't all doing it for the triangles.
Those triangles really are a mystery, given the design of those apes didn't have any evolutionary reason to seek triangle clicking.
To make my point instead of writing more squiggly lines, if you deconstruct something and then don't understand it, that's on you. Things can just work, they don't owe an explanation to you. If you have good explanation of when they work, I'm interested. If you don't, that's only reflecting on you.
I also try to heavily imply that if you generalize a human to an ape and then get astonished at the sounds he make, maybe that was an over generalization that missed something important.
Humans are interacting with the symbolic representation itself.
ChatGPT cannot do that.
ChatGPT instead models the semantic relationships between symbols, and interacts with that.
True AI may still be potentially founded on neutral nets, but the context problem definitely needs to be resolved.