LLMs can't simulate a Turing machine reliably
twitter.com
twitter.com
Basically I'm saying that (on my view) GPTs is kinda growing an AGI procedure inside it, but that AGI is restricted by rigid computational constraints imposed by the GPT architecture. I think something slightly more flexible might be capable of filtering out the noise more efficiently and replicate what GPTs do with an absurdly small fraction of the cost.
So, to be very clear, the point has nothing to do with Turing completeness and more to do with the ultra-limited computation and expressivity granted to internally learned functions. It is a small distinction, but a non-Turing complete model can still be extremely expressive.
LLMs can't do anything reliable AFIK
they are too much probabilistic to be fully "reliable"
at least if we take the definition of reliable from well written programs or even proof based programs
if we take the definition of reliable from idk. how reliable a underpayed and not well treated employee is then ... it might be a different thing
This is more "project this result onto a context -specific human like space of answers" than "compute it all in LLM"
A failure in this particular space in an observability product we use leads people wildly astray. While the data is correct, the summary is definitely not because it does not have enough understanding of the context.
I said to him "Would you really take a report into a meeting without knowing 100% what it was going to say on it beforehand? Even in the summary, do you want to risk presenting some wildly inaccurate LLM output to your boss and peers?"
Essentially: Accuracy is only one thing we want out of software. Replicability is another one. LLMs are terrible at the latter and only "good enough" at the former
Better use cases for LLMs are either where:
a) There is a non-judging human in the loop who is using the LLM as a tool, and accepts it's errors/limitations. It's a bit like asking a co-worker for advice - they might have something useful on occasion, but you're not just going to take what they say unchecked and put it on the company web site or send it to your boss as accepted truth.
b) You can verify the output of the LLM in someway - maybe easier if it's a structured output of some sort (SQL query perhaps) rather than free-form text.
b) if it was that deterministic, then you could generate it without an LLM.
Other than just playing with them, the only place I've personally found any utility in LLMs is essentially as a kind of search to get a toehold in a subject matter I'm not familiar with.
For the summary thing, I think we eventually settled on a workflow that generates the summary using LLM, then we store the output in our dataset for the report. This also allows the human building the report to modify it if necessary
This is different from say code generation where, to first approximation at least, automated tools (compilers, unit tests) can do this cheaply.
I don't see an equivalent approach that can be applied to LLMs.
I think the core issue is perhaps that they don't have any personal experience, nor know the trustworthiness or core competencies of their sources. Their training material is just an undifferentiated mass of text - some might be the writings of Richard Feynman, some the writing of someone on 4chan.
Us humans know the value of personal experience and expertise. If we have seen or done something ourselves, applying our own learned skepticism and knowledge, then we treat that as much more trustworthy than something posted on 4chan. If we read something in a book by a well known expert, or hear it from an acquaintance we have learned to trust, then we also treat that as likely to be true, although perhaps still not as much as if we've seen/tried it for ourselves (for which there is little substitute, since only then can we experimentally test any surrounding assumptions for the transmitted knowledge to put it in context).
Personal experience would require architectural changes, not least online learning. No need to anthropomorphize - our brain is also just a machine, and we can use it as inspiration for what is missing to build more intelligent systems.
"Here's some information on that" (links to WikiPedia page)
Are you speaking just out of anecdotal experiences? Try something simple like +1 to an integer. Or swapping an ascii character. Neural nets are layers of function approximations, and plenty of those weights were trained beyond the precision needed to map the entire set of syntactically correct inputs to the correct output, for a given well-written traditional function. Remember to set the temperature to 0 (deterministic). Of course, as you ask more complex things you quickly dip below the threshold of perfection. But to say they can't do anything reliably isn't true.
-641114711832495315269166494840 + 1 =
This expression is a sum of two large numbers, and the result is a very large number. Here's the calculation:-641114711832495315269166494840 + 1 =
-641114711832495315269166494840 + 1 = -641114711832495315269166494840 + 1
The result of this calculation is:
-641114711832495315269166494840 + 1 = 641114711832495315269166494841
https://x.com/VictorTaelin/status/1776677635491344744
This guy is offering a $10K prize if anyone can prompt an LLM to reliably (by itself - no tool use) repeatedly apply some simple letter-pair substitution rules to strings of up to 12 characters.
The rules to be applied (to strings composed of the letters A, B, X & Y) are:
1) AX -> delete
2) AY -> YA
3) BX -> BY
4) BY -> delete
In his Tweet he uses A# #A B# #B rather than A X B Y, but has clarified it makes no difference.
Try it - it's remarkably hard. You should be able to get the LLM to do a few short ones, but try getting it to be reliable or do longer ones.
So, Turing? Not going to happen!
You can force an LLM to go step by step with an external tool that continually re-prompts with each step, but the challenge requires a single prompt.
When you try to do prompt engineering to make it go step by step before telling you the answer, things go off the rails after a handful of steps.
To me, this is an odd thing to try to get LLMs to do because we know it's immensely hard to ensure reliable repetitive rule application to the point we spend many years conditioning humans on tasks where we care about the ability to stick with it, and then still built entire social structures on encouraging ongoing compliance, and still regularly fail to achieve it.
Why would we expect LLMs that are still fairly primitive to do this well without mechanisms regulating it?
2) While these are just LLMs (specifically, pre-trained transformers), they do have a place on the spectrum of intelligence, and some insiders are claiming "scale is all you need" to reach human-level AGI.
It's useful to step back, look at simple examples like this, and realize how far they are from being generally intelligent, as well as where our intuition for what they might be good at (repetitive tasks) is wrong due to their specific architecture.
To 2), my point is exactly that if anything talking about the struggle to get LLMs to do mindless repetitive rule application is not even close to being an argument against LLMs ability to scale to AGI. At most they're an argument for why LLMs will need at least some scaffolding around them. (That does not mean they are certain to be otherwise architecturally enough)
It is useful to step back and look at examples like this, but concluding from it that they say anything about how far they are from being generally intelligent is not. If it does get more people to drop the bizarrely unreasonable intuition that you should expect them to be good at repetitive tasks, that'd be helpful though.
Best measure of how far these models are from AGI, is what is architecturally missing (I'd say more than just scaffolding), although that also depends on how you define both AGI and intelligence itself. Still, I think examples like this are indicative of where we are right now - the model can neither do, nor be taught to do, the task, nor realize that it has failed (and then learn from the mistake). There is a lot missing.
So for a whole lot of tasks it is just down to scaffolding. For AGI it's not. But for a task like running a Turing machine it is and so it downplays their abilities to measure them by what they can be "convinced" to do in a single prompt.
Also, even without "think step by step", or an external loop, the generation process is already providing looping allowing the model to see it's entire output-so-far as input to generate each next word! The model itself just takes an input sequence and predicts one next word each time it is run; it's the generative process that then takes each generated word, appends it to the input and runs the model again to generate the next-next word, etc, only stopping when the model predicts the end-of-sequence token.
So, these models already have a chance to catch their own errors, if they are capable, without any additional looping beyond what is already occurring. However, I've never seen an LLM do this ("watermelons are blue, err, I mean green") since whatever caused it to predict "blue" the first time it saw "watermelons are" will likely cause it to predict "blue" on subsequent times (when it's seeing it's own output) too, so it won't realize there's been an error and predict a correction.
Even if a person points out a mistake, these models only have limited in-context learning to be able to use that to correct their built-in (pre-trained weights) behavior. Sometimes it works, sometimes it doesn't. I tried multiple times providing feedback and suggestions to Claude on how to fix it's mistakes on this string replacement task, and was never successful.
Some types of external scaffolding (i.e. modified generative process) might be helpful, such as "tree of thoughts" which would let the scaffolding pick the best output the model is capable of, but of course sometimes your best is still not good enough!
It is very different, because you control the context fed in at the start of each cycle. Put another way, it's fairly easy to get most LLMs to follow instructions one step. With a loop, that makes it fairly easy to get them to follow however many steps you like. But getting them to follow multiple steps by asking them to "think step by step" tends to fall apart after quite few steps because suddenly your instructions are a small, distant part of the context.
Test it. The difference is dramatic.
You can "coerce" the model more strongly when you explicitly control everything, though. Adding a loop also allows you to put it in a "straitjacket" in many ways - e.g. if a local model you can select output tokens to force the response to match a given grammar to give it less room to go off the rails (see llama.cpp/llamafile that supports this), you can repeat and adjust temperature until responses match requirements or manipulating the answer so far and have it complete it. Or combine with an evaluation step to force a "does this answer comply with restrictions X, Y, Z: ..." stage, and so on. Of course these are not suitable to get you AGI, but when you have tasks where testing compliance or ranking quality of output are "cheap"/"easy" but coming up with answers is hard, they provide ways to get LLMs to work even in cases where the simple approach of individual, complex prompts are too unreliable.
(the tradeoff of course being that if you have to spend too much time tuning these kinds of tricks to a specific problem you need to weigh up if they're worth it vs. solving the problem without an LLM)
Of course they are workarounds, and I'm sure there will be architectural advances that will hopefully make these relatively crude methods less important, but in the meantime, they do help.
It's simple enough to instruct it to work step by step, showing it's work (as a type of memory - same way we'd use pencil and paper), so that each string transformation breaks down into a sequence of single replacement steps that it shows (input, rule applied, output).
Tokenization is not the issue. Just give it the string as a space separated list of letters such as "A B B A X Y".
The problem is getting it to work reliably, and for longer strings. The best I could do with Claude Haiku was to get it to apply three replacement rules successfully, then it just made a stupid mistake.
LLMs are just not built for tasks like this. A lack of short-term memory is one obvious issue, but seems almost impossible to mitigate with step-by-step instruction.
Interestingly the top solutions all used Claude-3 Opus (via the API, with temperature set to zero). The winning prompt apparently cost $1+ to run, which given the $15/M input token cost of Claude-3 Opus means the prompt was around 100K tokens in size(!!), which in turn implies heavy use of examples/in-context learning.
FWIW, despite tokenized input, these models all seem capable of giving you the sequence of letters in a word if you ask.
If you want something that will "mindlessly" execute repetitive steps, anything trained to try to act in any way like how humans do would seem an exceedingly bad choice.
Regarding the second half. You aren't very creative. Sometimes you have something very basic, like two tables with numbers in the columns and you ask the LLM to select table rows using it's language capabilities and then tell it to give you the sum and it doesn't work. Great. Except in two years someone will figure out how to train a model to do this and nobody will care. The bitter lesson doesn't mean architectural changes are worthless, it just tells you that an approach that lets the machine leverage its computational capabilities will outpace a researcher's effort to patch a fundamentally broken model with a thousand small workarounds.
Nothing I said is to suggest it can't be done with LLMs - I'm on record in HN saying LLMs can be made to work as turing machines just fine, and I stand by that.
What you can't expect is to give it a prompt, no ongoing constraints or reinforcement, or other feedback, and stay on task for prolonged periods of time.
Humans don't do that with any consistency. We self regulate through a whole array of feedback systems, including internal ongoing thoughts, that by the time of exams have been conditioned over years to get better at sticking to a task and consider exam outcomes to matter.
And still a lot of us struggle to keep focus for even that kind of time period. Or engage in activities we know are bad for us.
Provide similar mechanisms to regular LLMs and they can stay on task too - we can trivially get LLMs to execute individual steps of turing machines, so getting them to keep doing it is a question of how to regulate their ongoing behaviour, not an inherent limitation.
These were mostly women who sat at desks all day long and did math.
https://www.smithsonianmag.com/science-nature/history-human-...
It was a pretty good job, there wasn't a lot of force, and money is a decent incentive to stay on task.
This works because each trip through the LLM with the same question uses a different random seed (temperature).
You're averaging randomness through a statistical model. You're going to end up with a repeatable result. You're not going to end up with an accurate result.
Also, my mother was a calculator once. In the UK. For petrochemical company. It was really boring so she went to smoke weed and fix radios instead as it was the 60s by then and that was cooler.
So if we take as a starting point that LLMs can't simulate Turing machines and that is important, the tasks you are referring to won't involve simulating Turing machines either.
I personally suspect LLMs can simulate Turing machines. Turing machines are pretty basic. Particularly after noting that humans use external tools like pen & paper, so it'd be fair to train the LLM to use some external memory-aiding tool. Then it becomes an academically trivial problem for even a basic neural net.
I get that it’s interesting to study what an LLM in isolation can do. That doesn’t accurately reflect what the commercially available AIs are, does it? Aren’t ChatGPT and Gemini AI tool suites?
It kind of annoys me that when you ask Copilot for help with something, it will spit out non-working code. Why isn’t Copilot running and debugging the code before presenting it as an answer?
LSTMs could learn slightly deeper into the Chomsky hierarchy than transformers, neural turing machines could learn them all including recursively enumerable (context free grammar or turing machine to produce).
Transformers did outperform LSTMs in something in an alternate hierarchy if I remember.
It's not conclusive since it is not training indefinitely, there are formal proofs that many architectures can do things that they can't actually effectively be trained for, but they threw a lot of compute at it.
Computers can’t simulate Turing machines reliably either, and for the same reason.
This is probably true for humans also.
My reading is that the quote says "to reach agi, it's not enough to learn a sequence of steps in text, we will want the agi to be able to crystalize (compile) that sequence of steps into a program run without abstractions on its native system." Presumably because it's too slow to run it on top of an LLM.
Stretching further. I think humans do crystalize learning into things that look hardcoded, like muscle memory or reflexes, rather than talking themselves through everything.
You could change your sampling to top-token-only. But (and maybe this is hand-wavey) I still suspect the incorrect token will be sometimes predicted on top.
BTW re the prompt problem... several persons actually solved it and the prize was claimed.