Parallel Search Speeds Up Time Travel Debugging by 4x
undo.io
undo.io
I'm a front-end dev, so I'm not nearly as familiar with the magic happening on the backend, but this looks like it matches what I know about how we manage jumping to specific points in time: many forks of the browser process, run from the beginning up to various pause points.
We started by recording languages without a VM - C/C++ - so we've implemented things at a lower level (which has advantages and disadvantages).
Coincidentally that means our tech could probably be used to record yours - and now we're doing more web stuff we ought to be using yours :-D
Likewise, if you're building more web apps, we'd love to hear how Replay.io works for you. We're on a shared journey to bring TTD to the world.
I also like the technical blog posts, etc you guys produce.
My thinking was that if you were simulating an embedded device you could record all state changes, and then be able to step forwards AND backwards (or jump to any earlier point).
There were academic implementations of this sort of thing around quite a while ago - and VMware's record/replay tech - but these days we have Undo LiveRecorder, rr / Pernosco, Microsoft's WinDbg / TTD, replay.io
The idea's time seems to have come...
We're not actually logging changes to variables, we're only logging outside information that changes the course of the computation. The variables are a result of that computation and not something we store directly.
So there's no explicit saving of A when we set it to B - and hence no opportunity to save the previous value of A either.
The reasons are both speed and space, as you suggest, but also fidelity of recording.
* If we recorded all variable values explicitly it could become very slow to run real world programs due to the extra work being done. * Recording all changes could also use a lot of storage for even trivial behaviours - e.g. for (i = 0; i < 10000000; i++); * Even if you log normal variable values you still have to worry about uninitialised memory, stray pointers, use after free - i.e. sources or destinations of assignments that aren't captured clearly in C language semantics. If we want to catch arbitrary bugs we do need to act at a level below normal-case language behaviours.
A side effect of this is that the underlying low level engine can record other languages with a different layer on top to handle language-specific semantics - that's what we do for Java.
At recording time we're collecting info about all the program's interactions with the outside world. In replay we prevent it from actually running any operations and instead just reinject what happened last time.
So if you had a program that did some IO and compute then we've have recorded all the system calls it did. When it's reading CSV data in replay we're feeding in the came data it read before. When it's doing things without outside effects we just say "sure, you've done that" and give it the original return code back.
So many opportunities for improved debugging on iOS.
More time travel debuggers is good for everyone - it spreads awareness and the techniques are pretty applicable across languages / platforms.
Does Undo use different techniques?
It has the advantage that we work fine in virtual machines, even where they don't allow performance counters. For what it's worth, we work fine on Mac Docker Linux containers for x86.
For native MacOS support it'd be more of a proper porting effort - we'd need to implement a different set of system calls.
The principles are very similar - each has some advantages over the other but as an Undoer I have an obvious bias as to which I prefer ;-)
rr uses snapshots and deterministic repay in the same way, though they ensure determinism differently. I don't know if rr can do parallel reverse ops but evidently Pernosco does parallel pre-processing to build its database (which is magic as it then allows very fast queries about program state).