HNHacker News
TopNewBestAskShowJobs

summerevening

2 karma · joined July 17, 2026

submissionscomments
summerevening··on Scaling to 1M concurrent sandboxes in seconds
> While we do need to write sandbox metadata and results to durable storage, we do so largely asynchronously.

How do you guarantee durability of sandbox task metadata if it’s written to durable storage async? What if the node it’s scheduled on goes down right after scheduling completes - what service durably knows about the intended state of the sandbox and retries scheduling?

summerevening··on Scaling to 1M concurrent sandboxes in seconds
Makes sense thanks!
summerevening··on Scaling to 1M concurrent sandboxes in seconds
What was the hardest part/most unexpected design challenge in getting this to work?
summerevening··on Scaling to 1M concurrent sandboxes in seconds
Do you binpack containers such that you overcommit cpu/ram on the machines to drive up utilization?

Did you do any simulations to see if this optimistic distributed scheduling approach maintains on-par utilization and low preemption rates to a non-distributed scheduler?

summerevening··on Scaling to 1M concurrent sandboxes in seconds
Every scheduler node has cached view of whole cluster and optimistically makes a scheduling decision, retrying on conflict?

Any tricks you did to reduce conflict rate? Is there a certain cluster saturation threshold (little free capacity) where conflict rates would get too high?