Q-Transformer
qtransformer.github.io
qtransformer.github.io
[0] https://en.wikipedia.org/wiki/Themes_in_A_Song_of_Ice_and_Fi...
My human brain doesn't use the same algorithm for learning to play a song on a piano as learning to play a new board game. I'm not an AI person, but it seems reasonable to imagine we'd have different "modules" to apply as needed.
AlphaGo probably sucks at conversation. ChatGPT can't play Go. The part of my brain writing this couldn't throw a baseball. The physics engine that lets me throw a baseball couldn't write this. Is there a reason we'd want or need one specific AI approach to be universally applicable?
I think the transformers architecture, or something very similar with eventually-on-policy time series forecasting in a markov decision process, is the right answer actually and was what I have been trying to make progress on for a long time[1].
Theory of AI is just going to be some weird network with a shit-ton of compute power, where the latter is more important to the outcome than the former.
The difficulty is the scale. Every synapse of a neuron is effectively a neuron itself, and every synapse acts on the synapses around it. So before you’ve even got to the neuron as a whole you’ve already got the equivalent of thousands of neurons and logic gates. Then the final result gets passed on to thousands more neurons.
I don’t know how you would recreate such complexity in programming. It’s not just the scale, it’s the flexibility of the structure.
>Encoded in the large, highly evolved sensory and motor portions of the human brain is a billion years of experience about the nature of the world and how to survive in it. The deliberate process we call reasoning is, I believe, the thinnest veneer of human thought, effective only because it is supported by this much older and much more powerful, though usually unconscious, sensorimotor knowledge. We are all prodigious olympians in perceptual and motor areas, so good that we make the difficult look easy. Abstract thought, though, is a new trick, perhaps less than 100 thousand years old. We have not yet mastered it. It is not all that intrinsically difficult; it just seems so when we do it.
Still working on closing the drawers afterwards, though...
https://diffusion-policy.cs.columbia.edu/
Looking at the figures and videos it seems... worse? Sort of surprised they didn't compare it but I guess they're trying to limit the discussion to purely reinforcement learning methods.
EDIT: Ah, I see this is an older result and was published concurrently to the diffusion policy paper, so it's likely the authors didn't know about it in time to add the extra comparisons.
It feels like opening a drawer once learn't, the bot has a (world?) model of what all drawers might look like, so it can open different drawers in any world. But this might not generalize to web interfaces and more specifically, how to do those actions on those interfaces?
Not to take away from what this paper's scope and achievements.
First it was the LK-99 hype, now it is Q. Chasing hype like flies to a light bulb.
[0] https://www.theverge.com/2023/11/29/23982046/sam-altman-inte...
It's Noam Brown's (former Deep Mind) "I'm joining OpenAI" thread from Back in June, and he talks about how he wants to bring some of the work from AlphaGoZero etc into OpenAI's AGI architecture.
It looks like this Q-transformer work shares a lot of the same lineage of what Q* is supposed to be. This is Google's attempt at mushing Q-learning into transformer land, which is probably what Q* is as well. It's not the same thing, for sure, but they are at least sibling bodies of work.
What's the letter?
What's the letter?
The letter of the day is...
Q!