ParentFull threadmmoskal·Their tech report says one inference deployment is around 400 GPUs...View on HN