- RL is extraordinarily sample-inefficient.
- distribution shift/catastrophic forgetting aren't solved. only off-policy learning with giant decorrelated batches works.
- the breakout success of transformers as an architecture doesn't neatly translate to robot motion policy models.
the field is missing fundamental breakthroughs.
I also find it very interesting that conversational AI has taken this long. where are the models with good turn-taking? passive listening? the ability not to respond in paragraphs? has Anthropic simply not gotten around to it?