Processing 24T tokens for LLM training with 0 crashes (what made it possible)daft.ai1 point·DISCURSIVE··0 commentsOpen articleSaveView on HN