Am I wrong to think that? Are LLMs in the future going to be able to “think through” actual logic problems?
Am I wrong to think that? Are LLMs in the future going to be able to “think through” actual logic problems?
They are literally Markov chains according to the mathematical definition. The code is complicated. Having complicated code doesn't mean it's not literally a Markov chain.
> I implemented Markov chains in BASIC in about ten lines of code in the 1980s on a 1 Mhz 64K Apple II after reading about the famous Mark V. Shaney hoax (https://en.wikipedia.org/wiki/Mark_V._Shaney). No neural nets or fancy GPUs required.
I don't doubt this. You can make a Markov chain by just counting the frequency of letters that follow each letter giving one that has a context window of one or two characters. That is a very simple Markov chain. You can make it by hand. You can make ones with more context window like a dozen characters or a few words, using sophisticated smoothing and regularization methods and not just frequency counts. Those are also simple Markov chains that you can do without neural net or GPU. Then you can also make a Markov chain that has a context window of thousands of tokens that is made from neural nets and massive training data and differentiable tensor computing libraries with data centers full of hardware linear algebra accelerators. Those are some even bigger Markov chains!
> LLMs are way more complicated than simple Markov chains.
That's true, they are more complicated than simple Markov chains, if by simple Markov chains you mean ones with small context window. LLMs are Markov chains with large context window!
Maybe you are used to Markov chains being shitty at language so you are confused how an LLM can be a Markov chain even though it's good at language and has some amazing emergent cognitive capabilities. That's a problem with your conception of Markov chains, it's not an argument that LLMs aren't Markov chains.
Finally, a Markov chain with a context space that cannot be practically iterated over (e.g. all possible 10k token contexts) can still be useful in ways that are shared with smaller Markov chains, even though if this weren't true it would still be a Markov chain. For example you can greedily generate tokens from it, calculate likelihoods, do some beam search, select multiple choice tokens, etc.
User: Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have?
Llama: Hello there! I'd be happy to help you with that question. Let me think... If each of Sally's brothers has two sisters, then together they would have 6 sisters in total (2 x 3 = 6). So, Sally would have 6 sisters as well. Is that the answer you were looking for?
User: no. Think about it. Sally is a sister to the brothers
Llama: Ah, I see what you mean now! In this case, since Sally is one of the siblings, she only counts as one of her own sisters. Therefore, Sally has only one sister - herself. Does that make sense?
User: yes
Llama: Great! I hope that clears up any confusion. Please let me know if you have any other questions or need further assistance.
User: yes
----
Um... no!
Specific logic questions like in this post, or long math formulas seem to still be a struggle. But I get the impression it is just a matter of time before it is scaled up to handle them.
You can argue whether it is 'reasoning' about them or not. But if it becomes indistinguishable from a human, then does the word 'reason' even really apply to the human anymore? Both are black boxes, giving answers.
In the end, if the result is indistinguishable from a treat people call reasoning, let's call the behavior reasoning.