Deterministic Replay of QEMU Emulation
qemu.org
qemu.org
https://github.com/jbreu/jos?tab=readme-ov-file#reverse-debu...
So much of my qemu work spent on randomly changing options, with no change documentation, discovering features, with no documentation, options with no reason or indication why, manpages out of date, READMEs not updated, changelog not there, etc.
However, all incompatible changes are documented and also announced at least 8 months in advance.
https://www.qemu.org/docs/master/about/removed-features.html
https://www.qemu.org/docs/master/about/deprecated.html
It may seem like there are many, but in practice they are in very old, mostly unused or very badly designed corners. For example configuration of audio was overhauled last year, and is now the same as basically all other backends (e.g. -audio pa,model=sb16; compare with -nic user,model=e1000 for a network card).
Generally we've been moving command line towards a scheme where each option describes an aspect of either the guest (a device, the board type, the CPU model) or the interface to the host (a file holding the contents of the disk, the network bridge to attach to, how to show graphic contents), with some options providing both as a shortcut (for example -nic, -audio, -serial).
I can think of using this for testing, and as a vehicle to change a programming paradigm of existing/legacy software (run a thing, and roll it back aggressively from outside of a vm)
IMHO, PANDA [1] remains a better/more practical choice for whole-system record/replay analysis. It already offers quite a bit of tooling (including a python interface), as well as hooks to build your own. It does have its own shortcomings (speed and not being in-sync with the latest QEMU), but at least you're not limited to gdb-based debugging.
While it might seem like a small feature, it opens a huge door. It's similar to what reproducible build infrastructure has done for finding bugs, attestation that binary matches source, immutability, etc.
Can imagine this is useful for finding bugs in hardware designs too.
Say capturing a Qt application as it corrupts its internal state during startup, in order to work out what's corrupting its internal state?
> Record/replay system is based on saving and replaying non-deterministic events
> The following non-deterministic data from peripheral devices is saved into the log: mouse and keyboard input, network packets, audio controller input, serial port input, and hardware clocks (they are non-deterministic too, because their values are taken from the host machine). Inputs from simulated hardware, memory of VM, software interrupts, and execution of instructions are not saved into the log, because they are deterministic and can be replayed by simulating the behavior of virtual machine starting from initial state.
So, it's probably not much, you can probably comfortably save minutes of qemu sessions.
Also note the existence of the rr debugger [1], which allows you to reverse debug applications with a ~10% performance hit while recording. To achieve this, it records results of syscalls (only). It will serialize thread events, so have the effect of running applications like on a single core CPU.
If it were a problem, you can skip recording your emulated machines bootup process, and simply take a snapshot when you're about to start your QT application. That snapshot probably only takes about 10% extra RAM because most of RAM contents wont change between the snapshot and the live system.
1 . Would something like this replace packer for creating machine images?
2. Curious how quickly the replay log grows and how it compares to a CoW snapshot.
3. Will be interesting what the log looks like and what doors could open up creating or generating it by other means.
https://github.com/qemu/qemu/blob/v2.9.0/docs/replay.txt
That said, I was not aware of it until I saw this post, and I definitely want to play around with it.
> That said, I was not aware of it until I saw this post, and I definitely want to play around with it.
You could almost say it was too casual and low-key. ;)
Does this sound right? I’m trying to figure out where uncontrollable randomness would come in during a compile phase, and coming up blank.
I have not followed the progress recently, but https://reproducible-builds.org/ is a starting point if you are interested.
There is a sane path forward for reproducibility on bare metal, no custom emulation is needed.
Both your causes seem trivially fixable here - the QEMU builds could have a standard system clock time they start with, and an ‘unsorted’ file listing made in a deterministic OS environment will keep the same file order, no?
By comparison the rb.org site says you need to start with stripping all that stuff out of your build process, for the reasons you refer to.
You'd be amazed about the amount of indeterminsim lurking in the guts of depencies all the way into libc and os ... Like locale, fs
Deterministic replay with QEMU is a "power tool" in the larger picture of these efforts.
So it'd be good for cases where you otherwise wouldn't be able to provide any verifiability. But for software, it's still not as good as eliminating non-determinism completely.
E.G. There could be malicious code hidden in the free RAM.
http://stackframe.blogspot.com/2007/10/configuring-applicati...
http://www.replaydebugging.com/2008/08/vmware-workstation-65...
VMware Workstation has such disjointed development spurts that it wouldn't surprise me if the feature had been ripped out at some point. Other useful features such as machine groups have been. :(