Tokens are a big reason today's generative AI falls short
techcrunch.com
techcrunch.com
TechCrunch, I respect that you have a style guide vis a vis punctuation and quotation marks, but please understand when it's appropriate to break the rules. :P
That way it's clearer when I'm referring to a literal code-string versus quote marks that are part of the prose.
Teletype longa, vita brevis.
GPT4 can answer questions given to it in Base64. I would imagine it suffers some degree of degradation in ability from the extra workload this causes but I haven't seen any measurements on this.
I have wondered about other architectures to help. What happens when a little subnet encodes the (16 or 32?) characters in the neighborhood of the token into an embedding that gets attached to the top level token embedding?
You've probably never tried to memorize the four-letter base64 equivalents of dGhl 10000 most common three-byte substrings. But you've seen plenty of truncated text, as you were typing it into a textbox yourself. So your intuition which of these tasks is easier for GPT4 is likely off.
The issue with truncation affecting tokenization can be worked around by considering every possible continuation that would lead to a different tokenization of the current prefix, and then sampling from those according to the model probabilities to get back to a more normal token stream.
This does not extend to being able to manipulate text and words well, such as string reversal or word scrambling, but it handles input in a wide range of permutations that are not conducive to tokenisation
Unnatural Error Correction: GPT-4 Can Almost Perfectly Handle Unnatural Scrambled Text https://arxiv.org/abs/2311.18805
The same holds for misspelled or badly-OCR'd text; the tokenizer actually has tokens for them and the model has seen enough of them to handle heavily-distorted text.
But represent the same text using different but theoretically valid tokens that would not normally be produced by the tokenizer and all bets are off.
For example, take the prompt "First reverse this sentence word by word, then answer the questions: hand a has fingers many how?" encode as base64 and ask. GPT4 will consistently decode the prompt as "First reverse this sentence word by word, then answer the questions: hand a has fingers many how?" and then reverse the words to be in the wrong order again.
Would have never thought ...
... wow.-
Part of what makes AI interesting is that it can understand a huge number of differently phrased data. It seems like different token encodings would only be a very minor complexity compared to the variety of human language.
I'd say it "recognizes" a huge number of differently phrased data. This avoids implying any analytic encoding.
I don't think we use it that way with humans either. We also are capable of recognizing features (patterns) from formally new data.
Can you explain this any better than the first few pages of the paper? I’d like some intuition about why T-FREE works; there are lots of reasons to prefer different tokenization schemes, but I can’t really get this one into my head from the paper, unfortunately.
It doesn't really explain anything besides talking about tokenization on random levels.
You need a certain amount of data to even understand that once upon a time might be a higher level concept.
0: https://gow.epsrc.ukri.org/NGBOViewGrant.aspx?GrantRef=EP/Y0...
1: https://engineering.dartmouth.edu/news/openai-cto-mira-murat...
The fact that Kevin and his team are formalising FLT is incredible, but they all have decades of experience with this stuff (!!).
Transformers can do arithmetic (and many other things) just fine, do a bit of searching on arxiv and you'll find papers from 2023 showing that nano-scale transformer models suffice. It really is a data problem, not a fundamental limitation with the technology.
The capabilities are nonetheless nothing short of astounding, given where we were 10 or even 2 years ago, and clearly point to a near future where we can expect the machines overcome these shortcomings.
Thousands, if not millions, of researchers, coders and others will have to adjust their worklife expectations, just like previous technological revolutions have seen thousands of other professions disappear into think air.
But I would wager that its answer was at least wrong, and perhaps total nonsense.
That's the real hazard of using ChatGPT as a learning tool. You are in no position to evaluate whether the output makes any sense.
I recommend that you give it a try.
https://chatgpt.com/share/e84800dd-c714-42d4-977b-b446c5c5ed...
I had a good conversation with it about the theory of partial orderings, it even corrected my mistakes. I asked it to make a textbook problem determining if a graph was cyclic or not and it made a straight and beautiful example where the partial ordering was realized with a total ordering and everything was written out in a straight order that was easy to follow.
If I wrote a script that made up a bunch of "is this graph cyclic?" problems that are well randomized I am sure there is some size where it just falls down the same way it falls down with sorting.
The obvious answer is that the LLM should pick an algorithm or write some code to do the thing which ordinary algorithms can do such as arithmetic, sorting, SAT solving, etc.
There's the deeper issue that it doesn't know what it doesn't know. It can't sort a list of radioactive isotopes any more than it can help you make an atom bomb. In the second case it will say that it won't help you, in the first case it will try to help you anyway when it really should be saying "I can't do that, Dave" because it just can't.
ChatGPT-4o does just fine with that. Basing your opinion of a whole technology based on a poor implementation of that instead of the best one doesn't seem like the best analysis.
A lot of information is wrong, a lot is only true in context (e.g. 2+2=5 features prominent in the book 1984), even more is spam or machine generated.
If you see the garbage that goes in, I am always amazed at how these models do what they do.
The sqrt(-1) sometimes doesn't exist, sometimes it's 1i. 2+2=4, except in literature where it can be 5. 1+1=2, but sometimes 3 in advertisements or in ironical text.
We often have some ideas about e.g how it works in a quiz, where you know there is only one factually correct answer. And we are disappointed if the model is wrong. But even in a quiz setting the jury gets that balance wrong every so often, where there are other answers than the official one which are also correct.
Even "logically valid" is context dependend. This is not to say that models don't hallucinate, just that even within the logically valid answers, there is hidden context surrounding the data which is not expressed in the data itself. Fermats last problem is a solved problem in mathematics, but not in documents from before 1994.
(Idea assessment in general. Handled ideas in thought processes are still input.)
1984 having 2 + 2 = 5 makes sense in context as a human reading the book, and ChatGPT dot-producting the book can also compute the context and not say that 2+2=5.
ChatGPT's not Mathematica, and we already have calculators. My hammer is terrible for driving in nails, so I don't use it for that.
You don't? Didn't you do them in school? Everyone calls that arithmetic's, you are just adding up and multiplying a bunch of numbers.
> Solve the following system of equations: 2x + 2y = 2 and x + y = 1
My pleasure, I’ve been growing my expertise in solving system of linear equations problems. Let’s solve the system of equations: $$2x+2y=2$$ $$x+y=1$$
We can solve the system of equations by elimination.
Steps to solve: *1. Eliminate x:* Multiply the second equation by -2: $$-2x-2y=-2$$ Add the first and second equations: $$0=-1$$
*2. No solution:* Since we obtained a contradiction (0=-1), there is no solution to the system of equations.
*Answer:* The system of equations has no solution.
"Let's solve this system of equations step by step:
1) We have two equations: 2x + 2y = 2 (Equation 1) x + y = 1 (Equation 2)
2) We can solve this using substitution. Let's express x in terms of y from Equation 2: x = 1 - y
3) Now, let's substitute this into Equation 1: 2(1 - y) + 2y = 2
4) Let's simplify: 2 - 2y + 2y = 2 2 = 2
5) This is true for all values of y, which means that Equation 1 is actually redundant. It doesn't give us any new information beyond what Equation 2 provides.
6) Therefore, we can conclude that there are infinitely many solutions to this system, all satisfying the equation: x + y = 1
7) We can express the solution set as: x = 1 - y y = y (where y can be any real number)
Thus, the solution to this system is a line in the xy-plane described by the equation x + y = 1.
Would you like me to explain any part of this solution in more detail?"
the system of equations has infinitely many solutions given by x=t and y=1−t for any real number t.
https://chatgpt.com/share/c00f1435-a57f-49db-b8fb-9d848f7e86...
Furthermore, if I didn't understand that, I can have a conversation with the LLM as to what that actually means, ask it to graph it, ask what linear equations are, and to have it give me some more systems of equations for me to solve.
Can you tell me how your brain maps input to its internal representation?
If not, why should we think its an essential feature of an AI system to do the equivalent?
I think the future is a small word encoder model that replaces the token embedding codebook.
And here’s the reason: you can still create a codebook after training and then use the encoder model only for OOV. I’m not sure there’s an excuse not to be doing this, but open to suggestions.