These sentiments are pretty close to my own. I read a paper that claimed that llms are General Pattern Machines and could be used to complete small games in gym environments. It seems to me that if these things really are General Pattern Machines all we have to do is figure out a way to represent any data as a pattern and try and predict the next step in the pattern right?
The multi token[1] project which allows you to take any type of data and turn it into a token it's pretty interesting and seems like it's going in this direction.
I would really like to see a framework where you can take any modality of any type turn it into a series of tokens and just cram it into a language model and effectively turning into a multimodal model with almost no effort.
[0]https://general-pattern-machines.github.io/ [1] https://github.com/sshh12/multi_token