Graham: Synchronizing Clocks by Leveraging Local Clock Properties (2022) [pdf]usenix.org·59 pts·mlerner·13
Resiliency at Scale: Managing Google's TPUv4 Machine Learning Supercomputermicahlerner.com·1 pts·mlerner·0
Gemini, Amazon's system for fast failure recovery in distributed model trainingmicahlerner.com·2 pts·mlerner·0
Defcon: Preventing overload with graceful feature degradation (2023)micahlerner.com·237 pts·mlerner·95
Gemini: Fast Failure Recovery in Distributed Training with In-Memory Checkpointsmicahlerner.com·3 pts·mlerner·0
Gemini: Fast Failure Recovery in Distributed Training with In-Memory Checkpointsnewsletter.micahlerner.com·4 pts·mlerner·0
XFaaS: Hyperscale and Low Cost Serverless Functions at Metanewsletter.micahlerner.com·3 pts·mlerner·0
XFaaS: Hyperscale and Low Cost Serverless Functions at Metanewsletter.micahlerner.com·3 pts·mlerner·1
Blueprint: A Toolchain for Highly-Reconfigurable Microservice Applicationsmicahlerner.com·2 pts·mlerner·0
Gemini: Fast Failure Recovery in Distributed Training with In-Memory Checkpoints [pdf]cs.rice.edu·50 pts·mlerner·13
Efficient Memory Management for Large Language Model Serving with PagedAttentionnewsletter.micahlerner.com·3 pts·mlerner·0
Efficient Memory Management for Large Language Model Serving with PagedAttentionnewsletter.micahlerner.com·1 pts·mlerner·0
Blueprint: A Toolchain for Highly-Reconfigurable Microservice Applicationsmicahlerner.com·1 pts·mlerner·0
Blueprint: A Toolchain for Highly-Reconfigurable Microservice Applicationsmicahlerner.com·5 pts·mlerner·2
Towards an adaptable systems architecture for memory tiering at warehouse-scalemicahlerner.com·28 pts·mlerner·4
TelaMalloc: Efficient On-Chip Memory Allocation for Production ML Acceleratorsmicahlerner.com·52 pts·mlerner·1