Moe inference optimizations: 15% lower expert load by request reorderingblog.doubleword.ai3 points·mezark··0 commentsOpen articleSaveView on HN