Really interesting work. we’ve been building a container-native snapshotting system too, but focused on cold start reduction and multi-model orchestration for LLM inference.
Different use case (sub-2s loading for large models), but very similar challenges around memory, device state, and restore reliability.