664 karma · joined January 28, 2022
We make software that can record/replay other software (including running it backwards). It's cool, please try it.
Sometimes I share things on Twitter - https://twitter.com/mark_undoio Sometimes I write articles (along with my colleages) about GDB - https://undo.io/resources/gdb-watchpoint/
I'd have said the situation back then was a bit better than that - a Java applet wouldn't have been able to access your filesystem by default, for instance.
Part of the benefit of Sun's Java was that the bytecode itself could be statically verified to only have good behaviour and the plugin would then sandbox what it could access at runtime. The plugin itself would obviously have had bugs - like all software - but it's not obvious to me that was intrinsically worse than having all that code as part of the browser (as we do now).
I'd contrast it with ActiveX and (I think), which was very free about what its applets could do (basically just Windows executable code, I think). Flash I'm less clear on the limitations of.
We have moved on in other ways, of course - browsers are architected to isolate processes more, including use of things like seccomp.
Acorn is a local company for me - a load of my colleagues / contacts through the local tech scene worked there.
It's also a company whose products I originally saw in school when I was about 8 years old - they were the fancy computers, the ones kids like me weren't allowed to mess with.
We're perhaps in a different space (time travel debugging for C/C++/Java/Rust/Go/Python) so we have quite narrow focus versus, say, JetBrains or GitHub who are providing much larger portions of the developer's tooling (everything a modern IDE does / CI / project workflows).
Tooling can be difficult - people sometimes expect it to be free (and, often, open source). They also struggle to justify spending on stuff that makes developers more productive. Almost everybody will say that want that, almost everybody has difficulty quantifying it to take to their budget holders.
But it's still a good business - we make money doing technically fascinating work while changing other developer's day-to-day lives for the better all the time. We have a unique technology that companies hundreds or thousands of times our size our envious of (and pay us for access to). It's not a bad deal!
> No, I was sad because around about the halfway mark he starts going into the title topic: “Can DevTools Get to $1B ARR?” and (the host) Tim Chen agrees and says basically only Hashicorp, GitHub, and now Vercel “made it”, and everyone else has had disappointing, middling outcomes. (I might add GitLab, Sentry, JetBrains, Atlassian, presumably Linear, but yeah it’s a short list). Entire markets were savaged in that conversation - all API tooling — Kong ($2b), Postman (valued $6b, now $3b), kaput, B grade, thanks for playing.
Now, I think this is true in a way: it shows that dev tools businesses (in general) are underappreciated. You're less likely to get change-your-life-rich off them (bearing in mind that most start-ups fail, so the typical case is always "not rich").
But I also think it's a very all-or-nothing take. You don't have to make $1B ARR to make a dent in the problems you want to solve, or to make money or to employ a team of brilliant people. It depends what your win condition is.
And CLion / VS Code for people who prefer an IDE interface.
But a lot of people do really want to stick with their printf debugging.
If your boss won't buy you an Undo you can still use https://rr-project.org/ - or on Windows the built in time travel debug of WinDbg.
If you'd like to try it please get in touch, feedback is always useful.
It's worth noting here that you can also build your binaries and keep debug symbols separately.
You don't need to ship them with the binary (although it will make many scenarios a bit simpler if you do, since you'll always have the right ones available).
Some info that might help: https://www.tweag.io/blog/2023-11-23-debug-fission/ https://undo.io/resources/gdb-watchpoint/reduce-binary-size-...
You don't need to pause anything whilst capturing the bug.
Now, if your code is literally connected to running critical systems in an actual factory then you've probably got additional realtime and safety-critical considerations that might push you towards debugging.
But (for more conventional use cases) time travel debuggers can handle multiple communicating systems without causing timeouts, capture bugs in software that interacts directly with hardware devices, etc. And you don't have to keep rebuilding / rerunning once you've reproduced the bug.
When I got a split layout key it became quite apparent that my technique was "weird" - sometimes I'd notice one hand come wandering over to the other side of the keyboard to find a key it was used to pressing. That became easy to correct once it was so visible!
shudder memories of my early days in kernel-level programming where using a printf used to just "fix" some bugs.
That came down to an uninitialised variable (which calling printf was helpfully initialising by using the stack).
As a result of that early experience, when I'm in the headspace of very low-level bugs, I find it helps to think of printf both as "this will tell me some variable values" and as a source of other clues: if it makes a weird value disappear or change then you might have uninitialised data on your stack, if it makes a flaky behaviour become stable (either disappear or become repeatable) then it's probably a race condition, etc.
A reference to James Mickens' "The Night Watch" feels appropriate here: https://www.usenix.org/system/files/1311_05-08_mickens.pdf
Time travel debugging (https://en.wikipedia.org/wiki/Time_travel_debugging) can help with this because it separates "recording" (i.e. reproducing the bug) from "replaying" (i.e. debugging).
Breakpoints only need to be set in the replay phase, once you've captured a recording of the bug.
In theory, you can make conditional breakpoints very fast using an in-process agent. For GDB (for instance) this gives the ability to evaluate conditional breakpoints within the process itself, rather than switching back to GDB: https://sourceware.org/gdb/current/onlinedocs/gdb.html/In_00...
I've always found the GDB documentation to be a bit vague about how you set up the in-process agent but I found this: (see 20.3.4 "Tracepoints support in gdbserver") https://sourceware.org/gdb/current/onlinedocs/gdb.html/Serve...
When we implemented in-process evaluation of conditional breakpoints in UDB (https://undo.io/products/udb/) it made the software run about 3000x faster than bare GDB with a frequently-hit conditional breakpoint in place. In principle, with a setup that just uses GDB and in-process evaluation you should do even better.
> I know for some people this is often a terrible UX because of the performance of debug builds, so a prerequisite here is fast debug builds.
The reasons debug builds perform badly are kind of mixed, in my experience looking at other people's set ups:
Building without optimisations
It's fairly common to believe that debug builds have to be built with -O0 (no optimisations) but this isn't true (at least, not on the most common platforms out there). There's no need to build something that's too slow to be useful.
You can always add debug info by using -g (on gcc / clang). Use -g3 to get the maximum level. This is independent of optimisation.
You can build with any level of optimisation you want and the debugger will do its best to provide you a sensible interpretation - at higher optimisation levels this can give some unintuitive behaviours but, fundamentally, the debugger will still work.
Gcc provides the "-Og" optimisation level, which attempts to balance reasonably intuitive behaviour under the debugger with decent performance (clang also supports this but, last I checked, it's just an alias to -O1.
Doing a ton of self-checks
People often add a load of self-checking, stress testing behaviours and other things "I might want when looking for a bug" to their code and gate it on the NDEBUG macro.
The logic here is reasonable - you have a build that people use for debugging, so over time that build accumulates loads of behaviours that might help find a bug.
The trouble is, this can give you a build that's too slow / weird in its behaviours to actually be representative. And then it's no use for finding some of your bugs anymore!
I think it would be better here to have a separate "dimension" for self-checking (e.g. have a separate macro you define to activate it), rather than forcing "debug build" to mean so many things.
The time travel debugging available with WinDbg should be able to wind back to the point of corruption - that'd probably have taken a few days off the initial realisation that an async change to the stack was causing the problem.
There'd still be another reasoning step required to understand why that happened - but you would be able to step back in time e.g. to when this buffer was previously used on the stack to see how select () was submitting it to the kernel.
In fact, a data breakpoint / watchpoint could likely have taken you back from the corruption to the previous valid use, which may have been the missing piece.
It seems like the select() was within its rights to have passed a stack allocated buffer to be written asynchronously by the kernel since it, presumably, knew it couldn't encounter any exceptions. But injecting one has broken that assumption.
If the select() implementation had returned normally with an error or was expecting then I'd assume this wouldn't have happened.
I've never satisfied myself that you can't just make a legacy codebase work that way, given enough effort, but I am not fully convinced it's always a good idea.
Last time I tried it you were able to add logging statements "after the fact" (i.e. after reproducing the bug) and see what they would have printed. I believe they also have the ability to act like a conventional debugger.
I think they're changing some aspects of their business model but the core record / replay tech is really cool.
It makes command history not work by default but IIRC "focus cmd" fixes that.
It minimises the mental effort to get to the next potential clue. And programmers are naturally drawn to that because:
1. True focus is a limited resource, so it's usually a good strategy to do the mentally laziest thing at each stage if you're facing a hard problem.
2. It always feels like the next time might be it - the final clue.
But these can lead to a trap when you don't quickly converge on an answer and end up in a cycle of waiting for compilation repeatedly whilst not making progress.
If you have a time travel debugger then you can record concurrency issues without pausing the program then debug the whole history offline, so you get a similar benefit without having to choose what to log up front.
E.g. use Microsoft's WinDbg time travel integration: https://learn.microsoft.com/en-us/windows-hardware/drivers/d...
Or on Linux use rr (https://rr-project.org/) or Undo (https://undo.io - disclaimer: I work on this).
These have the advantage that you only need to repro the bug once (just record it in a loop until the bug happens) then debug at your leisure. So even rare bugs are susceptible.
rr and Undo also both have modes for provoking concurrency bugs (Chaos Mode from rr - https://robert.ocallahan.org/2016/02/introducing-rr-chaos-mo..., Thread Fuzzing from Undo - https://undo.io/resources/thread-fuzzing-wild/)
And you don't need a full debugger setup on the target machine, just the recorder binary.
Gives you the possibility to have a proper debug experience without having to set up debugging that somehow works in a live k8s pod, or connects through special firewall holes or somesuch.
Then they'll do the same thing when you replay.
Non-idempotent system calls are tricky because they interact with the outside world - but that's still OK.
In Time Travel Debug, the process you're debugging is essentially in the Matrix. When it's being recorded everything acts as normal (and it'll see the real results of those non-idempotent calls).
When it's being debugged, any interaction with the outside system is prevented and replaced with the behaviour we saw at record time. It'll still think it's doing the non-idempotent calls, they just won't change (or depend upon) the state of the rest of the system.
You might find our Java product interesting, it adds Time Travel Debug to IntelliJ - https://undo.io/products/java/
Undo captures everything the process does, below the JVM level, so you can reproduce / rewind any problem you record as many times as you want (and copy the recording out of production onto a dev machine to debug, etc etc).
Please get in touch if you'd like a free trial.
A while ago there was a project to port it to GTK3 but I think that went away. I'm glad the mainline project is still going.