Why reinforcement learning plateaus without representation depth (NeurIPS 2025) | Hacker News Reader