The goal was to use XOR and PARITY as simple tasks and formal languages that exhibit a larger problem with the architecture. The unfortunate side-effect of this is that the most straightforward thing to relate them to are human reasoning tasks like arithmetic, for which the most straightforward thing to say is then: "We use calculators too!". Or worse, over-focusing on the PARITY problem, specifically.
The PARITY problem requires a kind of circuit that flips a bit on and off alternatively whenever we see a '1'/ signal. Is there a subconscious process that we employ in our brains that implements this circuit? Might some aspects of linguistic processing require this? (Michael Hahn brings up 'iterated negation' here: https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00306...)
All this to say, maintaining and updating latent state may be an implicit requirement for certain tasks, and that would be more likely learned from data than from an explicit prompt or outsourcing to a calculator.
Our brains are limited to using neurons as building blocks, but why should an AI be limited in the same way?
[1]: https://www.sciencedaily.com/releases/2022/02/220214121241.h...
* Classify the computational power of Transformers (when it stumbles on certain easier problems but can solve harder ones)
* Find a "minimal" change to the Transformer that would allow it to compute these problems.
Solving these 2 problems by giving LLMs arbitrary access to external plugins is a cop out. You would not: * Call youself a chef just because you own a restaraunt (you need to cook too!)
* Or (more program-y), say that C code meets Rust's memory safety standards simply because you can write the main function in C and write the rest of the program in Rust
Allowing arbitrary external plugins seems absurdly overkill and not 'minimal' (although that doesn't mean it isn't interesting from a practical perspective!), which is what I assumed that the rain1 was originally pointing out.edit: I don't mean to dismiss the work trying to figure out what they can do. That seems reasonable and valuable.
It's just, we're not trying to figure out how to tweak QR decomposition to solve arbitrary equations. It's a tool, a powerful tool, but it has some clear limitations.
So we can either study the LLM strictly from single-token outputs, or from how it combines into larger constructs. The former excludes not just CoT, but a large part of the benchmarks on which they are tested.
Looking at it another way, how would humans solve PARITY? I don’t think anyone would read the string, close their eyes, then give an answer without a single thought. At the very least, I would place my finger on each digit successively, effectively using a tool that acts as a memory storage (storing a cursor through the string). The reason I would do that, is that I would have devised an algorithm that requires going through digits one at a time.