Giving the programmer control over the checkpoints make sense. Also, limiting it to “workflows” (not necessarily entire programs) and requiring that these be deterministic makes sense.
I did some working with Folding@Home in grad school. All of the simulations would save their state to disk every N iterations to allow restarts. The state was relatively small (lots of compute, not a lot of data).
The DBOS papers focused on implementing core OS kernel functionality on top of a distributed relational database. Trying to make a true distributed OS.
This product seems like a pretty different direction (reliable execution and enhanced observability). Do your future plans tie back to the original project goals or are you going to keep going in a different direction?