How we made one of our largest inference workloads 4.7× more GPU-efficientdecagon.ai1 point·aray07··0 commentsOpen articleSaveView on HN