98% GPU Utilization Achieved in 1k GPU-Scale AI Training Using Distributed Cache | Hacker News Reader