What rr does
rr-project.org
rr-project.org
All too many things like this either dive straight into the deep end inundating you with superfluous details when all you want is a primer, or provide so little information as to be nearly useless.
The writers on this did a great job.
I've not read their extended technical report, but I am kind of curious exactly what performance counters AMD is implementing poorly and how that impacts rr.
There is still one remaining annoyance, which is that AMD's NMI latency is super high, which directly tanks rr's reverse execution latency. There's probably some improvements that could be made to the replayer to be more aggressive about optimistic assumptions on NMI latency and retrying if those assumptions are off, but it'd be a fair bit of work. I don't really understand why AMD decided to use this kind of architecture. It also makes profiles much less accurate.
For those that have used it, how useful it is for debugging multithreading heisenbugs? Can I let a process run under rr for days, wait until it crashes to due a heisenbug, and replay the trace without rr having to go through days of recording? i.e. is it possible to fast forward the trace, somehow?
(I nerd sniped myself a bit here, wondering how fast forwarding could be implemented. I think it might be achievable with periodic process memory snapshots and incremental traces.)
I could not reproduce the bug in less than an hour of run time, which meant that analysing the bug in gdb required an hour for it to run forward to the crash point, after which it was possible to skip back and forth.
AIUI you get 'start at an event' basically for free, because 'step backward' is implemented as 'start at the preceding event and then step forward by N', so events are frequent in the trace and the machinery to get to that point without running all the way from the start of the debug session exists anyway. There's some stuff on the website about how this is all implemented, I think.
You'd be better off restarting the recording periodically. Also, rr has a "chaos mode" which randomizes scheduling and often makes threading bugs easier to reproduce. https://robert.ocallahan.org/2016/02/introducing-rr-chaos-mo...
I had a heisenbug which would appear once a week, and that I couldn't trigger on my workstation. Chaos mode did the trick.
I wonder if a) rr could randomize the cpu ticks as well, at least in chaos mode, b) profiled code could somehow hint to rr that a certain instruction would be an "interesting" scheduling point.
You can "fast forward" the trace as you imagine. rr works by recording all non-deterministic input and output to the program so it can start from the core dump and step backwards.
As I understand it anyway; I've never actually used it - the one time I really wanted something like rr was on a Mac.
Not exactly. rr can't magically inflate a core dump into all the open file descriptors and other state accumulated during a process's execution. It needs to run from the beginning.
So starting from the beginning, you can let it run to any arbitrary point. (And there are ways of knowing useful points, eg if you record with -M it will print out event counts with anything written to stdout/stderr, so you can quickly run with -g to start debugging at the point that message was emitted.) But it does still need to run from the beginning. And you're recording a whole process tree, you need to start from the initial process and let it go forward to your requested point in time.
In practice, I usually use it by starting a replay, continuing forward to a crash (or a breakpoint at some line if it didn't crash), and only then starting to pay attention to what's going on. It's a simple, muscle-memory process to get to that point, and if it was a long recording you kind of start it up and wait until it's ready. (Which will take roughly as long as the initial run took to get to the same point. A little slower because of the overhead, a little faster because it doesn't actually have to wait for I/O, averaging out to a mostly unnoticeable amount slower.)
I always have to mention: one of my favorite things about rr is something that doesn't even require all the sophisticated machinery. I often want to debug a single process within a whole process tree, and with most things there aren't --debugger flags (or they're broken). With rr, I can just record the whole tree, then pick out the process I care about after the fact. It's a small thing, but it saves me from my usual hairball of wrapper scripts.
Random example: when debugging a gcc plugin, I record a call to gcc, but the actual compile I care about is done by a forked cc1plus process.
Not that useful, because as it says on the page (under "Limitations"):
> emulates a single-core machine.
For most uses rr is a major win, but for race conditions it sometimes doesn’t help.
Windows has something similar called Time Travel Debugging[1] but in my experience the dump files it creates can be enormous and be a pain to analyze as a result. (It also relies on WinDbg which while being extremely powerful and capable, has a huge learning and usability cliff. I’ve been using it for over a decade and I still need a cheat sheet from time to time. The revamped WinDbg Preview[2] improves the UI a lot, but ultimately it’s still WinDbg.)
[1] https://docs.microsoft.com/en-us/windows-hardware/drivers/de...
[2] https://docs.microsoft.com/en-us/windows-hardware/drivers/de...
> Are reposts ok?
> If a story has not had significant attention in the last year or so, a small number of reposts is ok.
Instant replay: Debugging C and C++ programs with rr - https://news.ycombinator.com/item?id=27034588 - May 2021 (66 comments)
Using time travel to remotely debug faulty DRAM - https://news.ycombinator.com/item?id=24589597 - Sept 2020 (62 comments)
Time Traveling Linux Bug Reporting: Coming in Julia 1.5 - https://news.ycombinator.com/item?id=23069372 - May 2020 (21 comments)
rr: lightweight recording and deterministic debugging - https://news.ycombinator.com/item?id=18388879 - Nov 2018 (52 comments)
Rr 5.0 Released - https://news.ycombinator.com/item?id=15191445 - Sept 2017 (3 comments)
Debugging Leaks with rr - https://news.ycombinator.com/item?id=10573308 - Nov 2015 (4 comments)
Back to the Futu-Rr-e: Deterministic Debugging with Rr - https://news.ycombinator.com/item?id=10492664 - Nov 2015 (9 comments)
Rr 4.0 Debugger Released with Reverse Execution - https://news.ycombinator.com/item?id=10441618 - Oct 2015 (11 comments)
Rr records nondeterministic executions and debugs them deterministically - https://news.ycombinator.com/item?id=8817954 - Dec 2014 (9 comments)
Rr 3.0 Released with x86-64 Support - https://news.ycombinator.com/item?id=8734502 - Dec 2014 (6 comments)
Porting rr to x86-64 - https://news.ycombinator.com/item?id=8543624 - Nov 2014 (9 comments)
> We identify a boundary around state and computation, record all sources of nondeterminism within the boundary and all inputs crossing into the boundary, and reexecute the computation within the boundary by replaying the nondeterminism and inputs. If all inputs and nondeterminism have truly been captured, the state and computation within the boundary during replay will match that during recording.
So for any chunk of time spent entirely in user space doing computation, the replay starts out in the same situation and executes in exactly the same way the original process did, with zero overhead. That's what enables rr to be so low overhead overall; most programs spend the bulk of their time computing stuff and reading/writing memory. The replayed process has no way of knowing that its file descriptors aren't actually open, since anything it does with them will be provided by the recording. Quoting again:
> In particular, user-space memory and register values are preserved exactly, with a few exceptions noted later in the paper. This implies CPU-level control flow is identical between recording and replay, as is memory layout.
1. a `log` command that just records whatever you give it into a plaintext file, together with its "point in time" according to rr. This is useful because when using rr, you tend to move forward and backward in time a lot, and it's hard to keep track of the actual sequence of events and where you are within them. It also creates a checkpoint so you can return to any one of your log points. It also has some niceties like replacing any expression enclosed in curly brackets with the results of executing the gdb expression given, so you can do things like
log starting execution of Init() with v={v}. About to crash.
2. a `label` command that lets you assign names to random hex values. Then in the output of `p expr` or the above `log` with no arguments, which displays the full set of log messages you've recorded, it will replace known hex values with their labels. This is so much nicer than memorizing numeric values and matching them up. (rr) p obj
$1 = (JSObject*) 0x7f606892a200
(rr) label OUTER_OBJ=obj
(rr) p $OUTER_OBJ
(JSObject*) $OUTER_OBJ
(rr) log
701/31299795 [c4] starting processing with obj=(JSObject*) $OUTER_OBJ
983/31299 [c2] starting processing with obj=(JSObject*) 0x7f6068a2a200
2081/7382911 [c3] traversing to (JSObject*) 0x7f6069c2a7e8
3316/199 [c1] crashing while accessing field of object (JSObject*) $OUTER_OBJ
The [c2] markers are the automatically-created checkpoints, numbered in order that you made those log entries in the debugger. It reorders the log messages to show them in execution order rather than debugging order. Pernosco has a very similar feature called the Notebook (where you only have to click on a log entry to view the state at that point in time.)The scripts are also intended for sharing log files and labels between multiple concurrent replays of the same execution, which I find useful to have separate windows each maintaining a different context (point in time, and portion of the execution that I'm examining.) That tends to be the buggier part of the scripts, though. ;-)
If you're working with C or C++ (or Rust? haven't tried it), rr really is a superpower. I rarely bother using straight gdb anymore. It feels crippled.
When you debug an optimized build with debug info in gdb by stepping line by line, it is easy to accidentally step "too far" and completely lose your place. In rr, you can always step back and recover.
Well done!
>rr aspires to be your primary C/C++ debugging tool for Linux, replacing — well, enhancing — gdb. You record a failure once, then debug the recording, deterministically, as many times as you want. The same execution is replayed every time.
>rr also provides efficient reverse execution under gdb. Set breakpoints and data watchpoints and quickly reverse-execute to where they were hit.
Performance overhead of reverse debuggers:
* gdb: >1000x (note: I never tested this one myself; just heard about this overhead in a HN comment a long time ago)
* Microsoft WinDbg TimeTravel Debugger: >40x
* rr: 1.5x
rr is the only one fast enough to be used on a regular basis -- the others are slow enough that they only make sense on particularly nasty bugs (usually memory corruption)
It also supports selective recording so, if this is configured (e.g. selecting certain functions), only a subset of the process execution is actually committed to the trace file, further reducing the overhead.
I'll have to look into selective recording, but I'm not sure how helpful it'll be in my use case (I don't know said library well enough to predict which functions might be causing the bogus values)
GDB does have reverse execution:
https://sourceware.org/gdb/current/onlinedocs/gdb/Reverse-Ex...
"rr currently requires either:
- An Intel CPU with Nehalem (2010) or later microarchitecture.
- Certain AMD Zen or later processors (see https://github.com/rr-debugger/rr/wiki/Zen)"This does also mean that there's https://-project.org, and that https://r-project.org secretly disambiguates into two different projects.