HNHacker News
TopNewBestAskShowJobs

timmisiak

70 karma · joined October 7, 2017

I used to work on WinDbg. I write blog posts occasionally at https://timdbg.com
submissionscomments
timmisiak··on Things I learned while writing an x86 emulator (2023)
I suspect the biggest issue is that courses like to talk about how instructions are encoded, and that can be difficult with x86 considering how complex the encoding scheme is. Personally, I don't think x86 is all that bad as long as you look at a small useful subset of instructions and ignore legacy and encoding.
timmisiak··on Things I learned while writing an x86 emulator (2023)
Glad you like it. I used m10c, with a few tweaks: https://github.com/vaga/hugo-theme-m10c
timmisiak··on Things I learned while writing an x86 emulator (2023)
I think both are useful, but designing a modern CPU from the gate level is out of reach for most folks, and I think there's a big gap between the sorts of CPUs we designed in college and the sort that run real code. I think creating an emulator of a modern CPU is a somewhat more accessible challenge, while still being very educational even if you only get something partially working.
timmisiak··on Writing a debugger from scratch: Breakpoints
Honestly it's a great exercise for learning how low level stuff works in general. Happy to answer any questions you have!
timmisiak··on Writing a debugger from scratch: Breakpoints
Yes, software breakpoints are difficult to get correct (the main reason why I started with hardware breakpoints). It gets more complicated with kernel debugging, where a single step (trap flag) could get pre-empted by an interrupt handler. And you can't always single-step a CPU and leave all other CPUs frozen.
timmisiak··on Writing a debugger from scratch: Breakpoints
Absolutely, rust sitter is fantastic. I haven't used any other parsers in Rust so I don't have much of a comparison point, but it's probably hard to get much more clear and concise, which I think really helps.
timmisiak··on Writing a debugger from scratch: Breakpoints
I think my record when I was on the WinDbg team was 5 debuggers deep.

I honestly think one of the best parts of writing a debugger is being your own recursive customer. I think that's something you only get to do for a few things. Debuggers, languages/compilers, and operating systems. And probably a few others.

timmisiak··on Omniscient Debugging (2007)
I worked on the one that ships with WinDbg: https://learn.microsoft.com/en-us/windows-hardware/drivers/d...

Literally every person who uses it for debugging a hard problem absolutely raves about it. Debugging something like a stack corruption or a heap corruption is trivial. And every other class of bugs are also incredibly easy, as long as it doesn't rely on very precise timing where the tracing changes the behavior. So why don't more people use it? Why haven't more people heard about it?

I'm not entirely sure, but I do have a guess. Most bugs are shallow. We tend to think about the really hard, really deep bugs, but the vast majority of devs are working on bugs where it's easier to just add some logging to figure out a logic error.

timmisiak··on Weird things I learned while writing an x86 emulator
Apologies. I'm just using a Hugo template and haven't spent a lot of time figuring out how to customize it. You're right though, the contrast isn't great.
timmisiak··on Weird things I learned while writing an x86 emulator
Yes, that's exactly what I did. Most useful thing I did was setting up a unit test framework to test a single instruction, and then generating a huge number of variations of those instructions.
timmisiak··on Weird things I learned while writing an x86 emulator
I talked a bit about the Intel SDM in the last post I wrote (linked at the top). The SDM is great for getting the specific details, but isn't always approachable for high level overview things.
timmisiak··on Weird things I learned while writing an x86 emulator
TTD (the WinDbg one) works very well for complex multithreaded apps. (The caveat being the performance hit you get from emulation). It's one of the big advantages it has over rr.

The main use case I saw for TTD was debugging complex memory corruption issues. Certain types of issues like stack corruption became trivial to debug under TTD. It was also very useful for capturing a repro. If a customer complained about something and I couldn't immediately reproduce it or get a crash dump, I'd ask them to record a TTD trace. More than 75% of the time I'd say it was enough to root cause the bug, without spending tons of time figuring out the repro steps.

timmisiak··on Weird things I learned while writing an x86 emulator
Yes, saving and restoring flags is very expensive. I thought about talking about that in the article but figured that was too much of a detour.

Darek Mihocka wrote a really interesting article about how to optimize flag calculations in an x86 emulator:

http://emulators.com/docs/nx11_flags.htm

Although looking at your username I suspect you may have read this one before...

timmisiak··on Weird things I learned while writing an x86 emulator
My point is the smaller encoding is what gives you the performance benefit. When done across an entire function/module, you can get a measurable increase in hit rate for the instruction cache. Not that the instruction itself is faster. Apologies if that wasn't clear.
timmisiak··on Weird things I learned while writing an x86 emulator
Yes, that made it 100x easier, because I was able to write unit tests that literally did a single step over the instruction and compare the register context to the emulated register context. The single step approach didn't work for all instructions (and didn't work for undefined flags) but was extremely effective and let me brute force a whole lot of testing.
timmisiak··on Weird things I learned while writing an x86 emulator
That's really interesting. I had no idea there were structurally valid instructions that are longer than 15 bytes.
timmisiak··on Weird things I learned while writing an x86 emulator
Funny enough, we tried going down the road of doing a JIT, but for our use cases pure emulation was often fast enough and it wasn't worth the extra complexity to do JIT. Part of that is probably due to the architecture of how TTD works, but part of it is how efficient we were able to get with the emulation mode. Later, a different emulator was created for running x86 code on ARM64, and that one uses an extremely efficient JIT. (I didn't get to work on the x86-on-arm64 emulator though. Would have been fun)

Working on the JIT in TTD was a lot of fun though, and it's a shame it didn't make sense to keep it around.

timmisiak··on The faker's guide to reading x86 assembly language
That's true. It's far too broad to learn everything without having a specific goal in mind. Learning the parts of asm that are useful for writing asm code is very different than learning enough asm to understand what APIs are being called from malware.
timmisiak··on The faker's guide to reading x86 assembly language
Covering how to practically use it is far too much for a single post (and I don't think I really claimed to teach assembly... I just want to give people the tools they need to start learning assembly). This was mainly to get people to not be scared and start reading asm, which is really the only way you can learn this stuff. In my opinion, writing a hello world in asm is not useful for most people. But learning enough to understand what part of a complex C++ expression caused a program to crash is much easier, and more widely useful. My plan for future posts is to talk about common compiler-generated code patterns to help people recognize what's causing a crash even if they don't understand every line of asm.
timmisiak··on The faker's guide to reading x86 assembly language
The target audience here were folks that have very little experience with asm, but you're completely right that a lot of complexity gets glossed over. That's not even to mention cases where an instruction behaves differently depending on the code segment attribute or privilege level it's running in.
timmisiak··on The faker's guide to reading x86 assembly language
That's actually something I want to tackle in the next post I'm writing. I need good examples though of stuff like that, so I think I need to spend some time with godbolt...
timmisiak··on The faker's guide to reading x86 assembly language
I'm planning to write a future post on "common compiler generated code" to recognize those patterns. The intention of this post was to encourage folks who think asm is scary to give it a try because it's not as bad as they think. You really need to spend a bunch of time with godbolt or a debugger with a side-by-side source and asm view to really build an intuition for these things. I want to try to give folks the tools they need to start building that intuition, because you can read dozens of articles about asm and still have no idea what you're doing.
timmisiak··on Time Travel Debugging
Oh OK, that makes sense then. Yeah, that's been fixed in current internal builds and will go out on the store soon.
timmisiak··on Thoughts on Microsoft's Time-Travel Debugger
Yup, you got it. Not just jumping back and forward, but seeking to points in time when memory changed. For instance... have some memory that got corrupted? With normal debugging you're probably out of luck unless you get some hints as to where it got corrupted. With TTD, you just set a memory access breakpoint and "run in reverse" until the memory is modified. It pinpoints the exact point in time where the memory got corrupted.
timmisiak··on Thoughts on Microsoft's Time-Travel Debugger
FYI, Microsoft TTD is supported on AMD CPUs.
timmisiak··on Time Travel Debugging
Right now you can attach TTD to a process after initialization (or any point in time prior to the repro), but you can't trace starting from a breakpoint. We'd love to add that functionality though, so I'm sure you'll see it in a future version of the tool.

Glad you like the new WinDbg! It's been very polarizing (which you can see just reading the comments on this post!), so it's good to hear the positive feedback sometimes :)

We've got a ton of plans to make debugging even faster and more effective in WinDbg, so hopefully we'll win more people over as we make our tools easier to use and more powerful.

timmisiak··on Time Travel Debugging
What duplicate lines are you talking about?

The debugger isn't written from scratch, just the UI. All of the underlying functionality is essentially the same, just in a more usable shell, and if you collapse the ribbon and retheme the UI to look like the 90s, you could almost squint and think it was the old WinDbg. The change is clearly very polarizing, but we're nearly at parity with what you could do in the old WinDbg UI, and we've already been able to give people features that we could have never dreamed of supporting in the old UI (not for lack of trying). The JavaScript support is just one example.

(Also, the cut-off title bar was an issue in the Fluent.Ribbon component we use, and I think it's fixed in an updated version that we're taking soon)

timmisiak··on Time Travel Debugging
The underlying engine is the same, so all of the functionality is fundamentally the same.
timmisiak··on Time Travel Debugging
I believe all of those projects are based on single-core debugging. Restricting the trace to a single core is one way of introducing determinism in the process as a way of allowing replay later. Microsoft time travel debugging is a multi-core technology, which allows you to capture interactions between threads in multi-threaded processes.

Edit: Since I'm still "posting too fast", let me respond here to clarify. Both rr and undodbg support multiple threads, but do not record them simultaneously. They restrict execution to a single core. That's what I mean by recording interactions between threads. For instance, if you have shared memory between processes, that won't work with rr/undodb generally (although I know at least rr can work around this by recording both processes).

I'm not saying one is intrinsically better than the other since there are tradeoffs with both approaches. There is a high constant overhead for recording all cores, but it does have the advantage of scaling well to large numbers of cores.

timmisiak··on Time Travel Debugging
The overhead varies depending on how cpu/IO bound the application is. IO isn't really affected, so IO bound applications tend to not see a big slowdown. In theory, you could see a very large slowdown in the worst case, but in the average case for a "medium sized" process the slowdown would be noticeable but not affect the usability. This technology isn't based on tracepoints, it's based on in-process cpu emulation. The emulation overhead is on the order of 10-20x in many cases, whereas tracepoint overhead would be on the order of 1000x I believe (maybe worse).

If there is something specific you dislike about the visuals of WinDbg Preview, let us know through the feedback hub or emailing windbgfb@microsoft.com. We realize that folks that have been using WinDbg for 20 years are likely to not be interested in a new UI, but we face 20 years of legacy code every time we want to add a new feature to the UI. As an example, the new WinDbg UI has a javascript window for writing scripts that extend and automate the debugger. It took us approximately 8x less time to implement in WinDbg Preview than what we estimated it would cost in the legacy UI (and a much more junior dev was able to do it as well). We want to innovate without disrupting folks that have effective workflows in WinDbg, so we really want to hear feedback on the new UI. If there are specific things that we can change to make you more efficient in the new UI, please let us know.

Page 1 of 2Next →