Right - but this feels like being half-way between transformers and lookup tables; the latter have enormous expressive power too, as long as you're willing to handle even more state. I'm curious what else could be put on that spectrum, and where. Maybe there's something weaker than transformers, but good enough for practical applications, while being more resource-efficient?
Related thought: understanding and compression seem to be fundamentally the same thing.
I wish! Gboard is terrible at predicting the correct form of a word.
It generally knows the best word, but it doesn't get from context that (e.g.) a gerund form of the verb is the most appropriate next
Basically, for most forms of complexity, Markov Chains are a strictly bad model. Perhaps if there was a way to "virtualize" the multiplication of states, you could do something reasonable. Idle speculation though.
Both are fundamentally prediction. :)
To quote a friend of mine: If I can't just trust the status quo I would have to question everything????
#ItsSoMetaEvenThisAcronym
Where the state space would be proportional to the token length squared, just like the attention mechanisms we use today?
Eg imagine input of red followed by 32bits or randomness followed by blue forever. Markov chains would learn red leads to blue 32bits later. They’d just need to learn 2^32 states.
A few more leaps and we should eventually get models small enough to get close to information theoretic lower bounds of compression.
as long as you don't care about the quality of what they're expressing. there's a reason they never did anything better than the postmodernism generator.
putting paint in a cannon has enormous expressive power too, but if you aren't rothko, nobody's going to care