HNHacker News
TopNewBestAskShowJobs

mark_undoio

664 karma · joined January 28, 2022

I'm the CTO at Undo - https://undo.io

We make software that can record/replay other software (including running it backwards). It's cool, please try it.

Sometimes I share things on Twitter - https://twitter.com/mark_undoio Sometimes I write articles (along with my colleages) about GDB - https://undo.io/resources/gdb-watchpoint/

submissionscomments
mark_undoio··on Amstrad Emailer, the UK’s first smartphone
I have a real affection for the Em@iler (or, as my friend used to call it, the "Ematiler").

My parents got one around when I went to university. It was a really good fit for them and helped us keep in touch. Main benefits:

* It did its job - a light showed up when you had something to read, otherwise you don't. * Always-on, no boot time. * It fit on the kitchen worktop, whereas PCs took up a lot of space. * The predictable schedule of dialing up, retrieving e-mail and then disconnecting let them know what their costs would be, even if they were a bit higher than the (variable) costs of normal PC internet use. * It had good contacts storage (with dockable electronic address book you could take away), at a time when you would still expect to be managing a list of telephone numbers on a piece of paper.

All for less than a PC. They could have got all the same things out of the PC, the Em@iler just put it in more convenient package.

mark_undoio··on Omniscient Debugging (2007)
That makes a lot of sense - having a deterministic recording makes a big difference to learning how the code actually works.

I also sometimes think that, even when I want a debugger interface to look at the details, I'd also like to see a "big picture" view showing how key values evolve over time.

(something like replay.io does for visualising logging statements is probably in about the right place for both of us, though I'm not a web dev)

mark_undoio··on Omniscient Debugging (2007)
VS Code has a concept called "log points" in the C/C++ debug mode that use a debugger to implement logging.

So you can combine the best of debugging and logging.

At Undo we've done some investigation of how to get this out into the world for time travel debug, so it's interesting to see your idea lines up.

What sort of problem do you think this would be useful for?

mark_undoio··on Omniscient Debugging (2007)
The answer, as always with engineering is a "yes but"

It's pretty appealing. I also believe MS Visual Studio had something that effectively worked like that albeit without changing the coffee manually.

But - if you've got bugs due to memory corruption, race conditions, etc then the logging doesn't necessarily tell you what you need as you don't know when the change actually happened.

The other issue is that once you've started logging everywhere the string formatting is likely going to be a significant slowdown - at which point you could have used a time travel debugger and probably get better performance.

(disclaimer: I work on a time travel debugger)

mark_undoio··on Omniscient Debugging (2007)
Nice quote on one big benefit of such systems:

"there are no non-deterministic problems"

Edit: (disclaimer: I work on one such system)

mark_undoio··on The Hundred-Year Programming Language
>> Java is the most recent popular general-purpose language > This post was written in 2022. Does the author not know about Python? Javascript? Rust? C#? and a bunch of others?

I had to double check this - but Python is several years older than Java. Wikipedia lists 1991 for its first release, vs 1995 for Java.

That said, I felt like Python became really well-known much later than Java (which had massive hype and enthusiasm in the 90s) - so I do actually agree with you listing it here.

mark_undoio··on Weird things I learned while writing an x86 emulator
I had not read it before! Thanks for the link. I don't get too involved with our JIT other than to occasionally peer into some code and go "ooooh"
mark_undoio··on Weird things I learned while writing an x86 emulator
Very impressive you got such good performance out of straight emulation! And WinDbg TTD runs with truly parallel threaded execution within one process, which we do not (yet).

I'm surprised a JIT wasn't worth it but, from what I'm aware of, I can see a few reasons why the trade offs are different.

mark_undoio··on Weird things I learned while writing an x86 emulator
IIRC from our similar code, the LEA (load effective address) instruction is useful for doing adds / subtracts of arbitrary integers without updating the flags.

Flags turn out to be quite the annoyance for the kind of in-process virtualization needed by Time Travel Debug. You need to instrument code with minimal overhead so, on the one hand, you don't want to save/restore flags all the time .... And on the other hand it still all has to work when flags get used.

mark_undoio··on Weird things I learned while writing an x86 emulator
Ha, I was just reading some of the discussion and thinking it sounded quite similar to the JIT we use in our own (Linuxy) Time Travel Debugger.

It's similarly am x86-on-x86 JIT / emulator but (and I'm sure WinDbg's TTD is similar) most of what you want to do there is just code copying for any instructions that don't need special instrumentation.

And you want to run entirely in the cache of JITted code so you're close to native speed, rather than exit to C code and make decisions.

mark_undoio··on Religious and spiritual folklore surrounding programming
That one confounded us for a while. There was no true parallelism going on in the system but the code that needed nop padding was the interrupt handler.

That it needed padding wasn't unreasonable - we had tight timing constraints. It would have absolutely made sense for it to race with another part of the system but that didn't seem to be the nature of the interaction. The interrupt handler could always run when it wanted to.

The eventual deduction: the flash memory had 256 byte pages at hardware level. I inferred that there's some initial access cost to open a page (and probably cache it into SRAM or something).

The build was putting later functions in the file at later addresses, so what you put in other functions changed the alignment of the interrupt handler. If you requested 256 byte alignment for that function then you still needed nop padding but it was completely consistent.

mark_undoio··on Religious and spiritual folklore surrounding programming
Computers are extremely haunted. It's often remarkably they were working at all before you made the change, then you prod and all the goblins come out.

Similar experiences I've had:

* Putting a big array on a stack (early in C programming career) - compiles fine, instant crash at runtime.

* Colleague discovers that a function needs padding with NOPs to achieve precise timing. Number of NOPs needed varies depending on code changes (even NOPs added or removed) in functions higher up the same source file.

* Occasional crashes in low level routines on ARM64 after changing completely unrelated code. The stack appears to contain a struct from elsewhere in memory instead of ... a stack.

(First one was probably just the array was too big for sensible stack management/ growth and a friendlier compiler would have just told me. Second one was to do with the size of memory pages in the flash - there was a delay if you ran over a boundary. Third one was, IIRC, the variable containing the base of a temporary stack occasionally getting splattered to point into other data structures! The actual stack was fine but you couldn't see it anymore)

mark_undoio··on Intro to GCC bootstrap in RISC-V
It looks like the intention is to bootstrap from source using an absolute minimum of precompiled binaries.

For an average user I'd imagine this is mostly just intellectually interesting but it's potentially useful for distros to aggressively minimise the amount of unauditable code involved in their builds.

The GNU Mes project (mentioned in the article) seems to be part of general work in this direction: https://www.gnu.org/software/mes/

Which seems to be GUIX-driven: https://guix.gnu.org/en/blog/2020/guix-further-reduces-boots...

And some funding has been provided to the author to help with this: https://nlnet.nl/project/GNUMes-RISCV/

mark_undoio··on TinyEMU – x86 and RISC-V emulator, small and simple while being complete
My reading is that it emulates either x86 or RISC-V CPU and can emulate various real x86 peripherals but it also includes the abstract VirtIO services for ... both?

I assume the RISC-V emulation is all VirtIO based and not real...

mark_undoio··on Pictures of a working garbage collector
@chubot (author) I'm curious what the debug experience is like. It looks like you're generating C++ code that maps pretty closely onto your statically typed Python code - is it straightforward to say "aha, it's that variable in C++ so it'll be the same name in Python"?

I'm wondering if it's possible to put some directives into the generated C++ code that would map its debug info directly back to source lines in the Python - I can't find any docs to confirm my feeling this should be possible.

Edit: Maybe this? https://gcc.gnu.org/onlinedocs/cpp/Line-Control.html

mark_undoio··on Pictures of a working garbage collector
In our latest release of UDB we added the `last` command: https://docs.undo.io/TrackingValueChanges.html

Which effectively combines a `watch -l`, plus a reverse-continue but also monitors when the underlying memory was allocated / freed.

It's quite nice, basically "git blame" for your variables.

mark_undoio··on Pictures of a working garbage collector
(I work for Undo)

> gdb has had reversible debugging since release 7, in 2009. What does udb offer that it lacks?

GDB's built-in reversible debugging is cool (and it's helped raise awareness) but it doesn't scale well. We build on the same command set and serial protocol that GDB defined - UDB is GDB but with additional Python code hooking it up to our separate record/replay engine.

For UDB, I'd say we offer: 1. Performance & efficiency (orders of magnitude faster at runtime and lower in memory requirements). 2. Recordings can be saved to portable files (share with colleagues, receive from customers, etc). 3. Library API so applications can self-record with control of when to capture and save. 4. Wider support of modern software (proactively tracking modern CPU features, shared memory and device maps, etc). 5. Correctness (in the past we've found the reverse operations in GDB don't have as strong semantics as we'd hoped, though I'd also be happy to be wrong here)

FWIW, rr (https://rr-project.org/) also offers many similar benefits over GDB's built-in system (though not the library API in point 3) but with differences in what CPUs / systems are supported, ability to attach at runtime, etc.

If you're looking for an open source solution, I'd choose rr over GDB's built-in approach.

mark_undoio··on Pictures of a working garbage collector
Good to see GDB's watchpoints (https://sourceware.org/gdb/download/onlinedocs/gdb/Set-Watch...) get a mention. Often called Data Breakpoints in IDEs (which then, confusingly, often also use "Watch" for a different concept - argh).

Watchpoints, when they're suitable, are an incredibly powerful way to query your program's runtime behaviour (When does this happen? Why does this value end up here?) instead of stepping through and printing things / logging stuff.

They also fast, provided you watch values suitable for the CPU's logic to handle them directly. They need to be small-ish and located in memory for hardware support to handle it, otherwise GDB needs to single step.

(the `-l` flag is also important here and even less well known - it tells GDB you really care about the memory address the expression is stored in, not the expression itself)

mark_undoio··on C Posix-compliant argument parsing in 42 LoC, inspired by Duff's device
Because there's really a switch statement (hidden behind a macro) that will jump to labels within that block.

The fact that the if condition is false means it won't just run the whole block straight through but you can still jump to a label in it. A goto statement would also allow you to jump into an otherwise-unreachable block.

mark_undoio··on What’s Left in the Apple Silicon Transition
Hardware level stuff is a bit lower than my area but...

The two tricky things I'd anticipate are:

What does it look like to hardware? Different CPUs speak different protocols to maintain a cache coherent view of memory. So they couldn't easily share memory in a completely natural way - but GPUs manage to share memory between host and GPU, so it can be done (just not necessarily as conveniently to software).

What does it look like to software? Lots of weird options here. Most convenient for software compatibility is it the Intel is also somehow running MacOS (or appears to be). I guess you could do that - it'd be like having a two node cluster in one machine. But that's not as convenient as being able to distribute computation between compatible CPUs.

None of that is a complete technical blocker to a company that controls their stack, I'd just guess it would make it too inconvenient / expensive for them to do.

That said - as an intern, years ago, I once had access to a piece of prototype hardware that had an x86 chip in one motherboard CPU socket and a non-x86 chip in the other (with various adaptation hardware to allow them to play nice). I think that was more a proof of concept than anything else, not necessarily viable without other hardware changes. Interesting nonetheless.

mark_undoio··on Why is Rosetta 2 fast?
Something that fascinates me about this kind of A -> A translation (which I associate with the original HP Dynamo project on HPPA CPUs) is that it was able to effectively yield the performance effect of one or two increased levels of -O optimization flag.

Right now it's fairly common in software development to have a debug build and a release build with potentially different optimisation levels. So that's two builds to manage - if we could build with lower optimisation and still effectively run at higher levels then that's a whole load of build/test simplification.

Moreover, debugging optimised binaries is fiddly due to information that's discarded. Having the original, unoptimised, version available at all times would give back the fidelity when required (e.g. debugging problems in the field).

Java effectively lives in this world already as it can use high optimisation and then fall back to interpreted mode when debugging is needed. I wish we could have this for C/C++ and other native languages.

mark_undoio··on Negative 2000 Lines of Code (1982)
I like to call this "net negative development"! Valuable contributions to an established codebase can involve a tiny amount of code but a massive amount of understanding, so it's easy for a developer to clean up more in old code than they add in new code.

There must be a point where this doesn't hold as well - new feature projects or, especially, newer projects themselves - where adding code is more correlated with value.

mark_undoio··on Taste vs. Skills
I agree with you. Skill helps you appreciate what is good taste.

Both come with experience and a willingness to learn.

I'm not keen on the idea of taste as simply some innate quality, rather than something learned / taught. Thinking of it as some X factor makes it easier to write off people who just disagree or lack experience.

That said, I've totally seen code with poor taste and I've written it myself, so I believe in the underlying concept.

mark_undoio··on Taste vs. Skills
This also reminds me there are different dimensions to taste.

A developer can have both excellent taste and skill in code and lack one or both in its user interface. For instance, I'd personally say git seems to be amazingly successful and robust but that the UX could be much more consistent without sacrificing any power.

I fully expect there are other dimensions to this too for other overlapping domains in software engineering. Something like E.g:

- testing - does the test code work? (skill), is it making the most of opportunities for coverage and self-documentation (taste? Or a different kind of skill?) - API design - can it do what's required? (skill), is it a sensible ABI too? (skill), can others understand it? (taste), is boilerplate minimal? (taste) - documentation - is it correct and complete (skill) and understandable? (taste) - code review - can you spot bugs and inconsistencies? (skill), are you giving feedback that allows incremental improvements without overwhelming? (taste)

I fully believe both skill and taste can be learned in all these areas but different thought processes will help in each.

mark_undoio··on Old tech is haunted
Periodically, my diagnosis on test failures has been "maybe a ghost did it".

Test failures that are flaky / sporadic often signal something you really need to look into - either you're not in control of your test's corner cases or your product's. Either way you can't ignore it.

But, sometimes, with the whole "running hundreds of thousands of tests of complex, interacting software on cloud infrastructure that has virtualisation, custom kernels, whatever ..." the answer is just "woooooo spooky". When it comes to customer boxes, it's even weirder ;-)

=== When PTRACE_SINGLESTEP got haunted ===

A few years ago, we saw a particular issue in which the PTRACE_SINGLESTEP operation (which should, you know, step by a single instruction!) would sometimes step two instructions - but it basically only happened in testing.

Eventually we figured out that:

* For a particular older enterprise distro on x86.

* If you tried to single-step a thread in one process.

* And a thread in another process hit a watchpoint at the same time.

* Then your thread would step by two instructions.

Spooky action at a distance. This might also have required you to be running on a virtualized system - it didn't matter to us at that point, since clearly we need to work around issues people may see on cloud infrastructure.

Of course, unless you tested debugger workloads extensively you'd probably never have two processes under debug simultaneously in order to notice this.

=== When transparent hugepages got haunted ===

When RHEL backported support for the Linux kernel's THP optimisation, quite a few years ago, to one of their enterprise kernels a bug also got backported that I'd call "extremely haunted".

THP worked fine, maps would automatically get huge pages where appropriate and would get split up again if necessary.

EXCEPT ... if a THP-ed memory location was:

* Write-protected because of COW memory sharing (e.g. after fork)

* And the page was split by a write

* And that write had come from a PTRACE_POKE operation, not an in-process write.

Then something crazy would go wrong in the system. That process, IIRC, would hang and become unkillable.

Some sort of goblin got into the memory management subsystem at that point and there was no getting it out again - gradually the whole system would start to lock up until it became unusable - eventually requiring the box to be reset. This happened even from an unprivileged process.

Again, this is something most workloads wouldn't see but - once we'd understood the underlying issue - we had to find ways to prevent it happening. You just can't trigger a kernel bug like that, even if you know it's not your fault. If you're always there when it happens, nobody is sympathetic to you!

mark_undoio··on What’s wrong with medieval pigs in videogames
I believe it's the origin of "bootstrapping" as a verb, in which it immediately becomes unironic technical jargon.

i.e. as in getting a start-up off the ground or, when shortened further, "booting" a computer.

mark_undoio··on Product Market Fit
I guess you can control the markets you're focusing in - within reason. E.g. picking a particular sector to sell into. How much choice you get depends on the product you're building, so I suppose it's a bit circular.
mark_undoio··on Why adults still dream about school
> When am sufficiently stressed at work - a recurring dream I have is that I am in the exam room ready to take a test, but did not study or I have mixed up the dates and came prepared for the wrong subject.

Because my high school Biology course (A Levels in the UK) was divided into modules assessed throughout two years I came to the last exam pretty sure that I didn't have to score many marks to get the grade I was aiming for.

I ended up in a weird situation where my instinct was to prepare really well but, in practice, I should skimp on revision for Biology and concentrate on subjects that weighted the end-of-year exams higher.

The strategy worked out - I hit what I was aiming for in Biology and in the other subjects. But that whole very tactical approach so went against my character that, to this day, anxiety dreams manifest about not having done the work for that specific exam. Often they're further exaggerated (e.g. I've simply forgotten to go to any lessons all year, etc) but it focuses around that event.

I find it weird that my subconscious prefers that over arguably even more stressful (and sometimes less successful!) exams at university. Maybe I was just more equipped to deal with it by then.

mark_undoio··on Podcasting is just radio now
> And, advertising everywhere. Most podcasts aren't the handsfree experience they used to be, now that I have to seek past ads every 15 minutes.

In some comedy podcasts, I can find the ad reads to be entertaining enough to be worth a listen. In others my heart sinks every time I hear the same pre-recorded ad (that annoyed me the first time) popping up again.

Just how bad is it in general? Feels like there's an opportunity for good content in the ad breaks, even on non-silly podcasts - but that requires effort to generate something interesting / useful / fun that's not the same every time.

OTOH, for some products it seems like spamming out the same content to build recognition does seem to be a rational (or at least popular!) strategy, which is a shame.

mark_undoio··on io_uring_spawn: Launching New Processes with io_uring
I think vfork is intended to avoid the CoW overhead? (at the cost of having really bizarre other semantics that make it tricky to use)

Since it pauses the execution of the parent process (or at least the parent thread?) it gives a window for the child to use the current stack frame to set up an exec() call.

When I first learned about vfork() (early 2000s) I had the impression it had become a bit pointless since CoW sharing had made fork() really fast. But with modern-sized address spaces it is (has become?) horribly slow again.

← PreviousPage 6 of 8Next →