How breakpoints are set
majantali.net
majantali.net
This makes setting a breakpoint really easy, as all you have to do is replace a single byte (and restore a single byte) where you want to place your breakpoint. INT 3 being only one byte is also important when you're setting a breakpoint instead of a another single byte instruction - your newly set breakpoint won't override the consecutive instruction, which might be jumped to somewhere else in the code.
It's kind of the other way around. The reason it has a single byte opcode is because Intel wanted INT3 to be for break points, so they designated 0xCC for it. In fact, 0xCD 0x03 works, but just isn't used.
Lets say you have a lot of single byte opcodes:
40 INC EAX
43 INC EBX
41 INC ECX
C3 RET
And you want to set a breakpoint on INC EAX.
If you replace "40" with "CD03" - you'll overwrite INC EBX as well.
That can cause your program to crash if there are control flows that end up jumping to INC EBX without going through INC EAX first.That's the main reason why 0xCD, 0x03 isn't used.
One-byte opcodes won't save you when the code is jumping into the middle of instructions. The instruction you want to breakpoint might be in the middle of some other instruction that will run.
That is one way to look at it, but I find it a bit too limiting (debuggers can attach to an existing process as well) and too confusing (requires knowing what fork does/is, same for execl - and are those even used when attaching to an existing process?) and because of the latter functions used obviously coming from a linux background (nothing wrong with that, on the contrary, but I can imagine windows people or beginners still having no clue whatsoever about a debugger after reading this - though it's likely not the target group).
While a lot of tech is rapidly moving and constantly changing - this is the type of fundamental knowledge that will probably prove valuable for the rest of your career.
See this explanation of how it works in SpiderMonkey, for instance: http://rfrn.org/~shu/2014/11/20/speeding-up-debugger.html
Of course, one needs to take care of the respective secure access. :)
With an interpreter you can just probably ask the interpreter to stop at a certain instruction. For the purposes of your debugger the interpreter is the CPU in that respect. Except that you don't necessarily need to rewrite memory (although this probably exists too where the IL has a special breakpoint opcode).
With a JIT compiler you could do the same as in the article, but it complicates things because the code you want to debug may not have been jitted yet, or, with some JIT compilers, may not be jitted at all, ever (e.g. a small method that runs exactly once). You could also do the same as above, with a breakpoint opcode, or asking the runtime to break at a particular statement. Both cases require the JIT to play along and do the right thing. For code that isn't jitted you'd have to fall back to the interpreter anyway, though, so in some cases JIT compilation is simply disabled in the debugger (e.g. Java, if I remember correctly), which has the unfortunate side effect that not only you lose optimizations, which is normal for debug code, but you also get a hefty performance hit because you're now running interpreted.
http://eli.thegreenplace.net/2011/01/23/how-debuggers-work-p... http://eli.thegreenplace.net/2011/01/27/how-debuggers-work-p... http://eli.thegreenplace.net/2011/02/07/how-debuggers-work-p...
The fact that the code was encrypted meant the debugger couldn't disassemble it in any meaningful way, and also made it impossible to set a breakpoint (since the breakpoint would just end up being "decrypted" into some other opcode that would inevitably crash). The debugger also couldn't step through the code, because taking over the single step interrupt would prevent the decrypter from running, so you'd just be stepping through garbage.
The way I worked around this was by writing a debugger that could hook the single step interrupt in such a way that it still forwarded the interrupt onto the previous hook. I still couldn't set breakpoints, but I could step through the code, watching it decode itself as it proceeded.
For example, this is how the trace exception on the M68k works - the program counter of the next instruction to be executed can be read off the stack by the exception handler. The 6502 doesn't have built-in software single-stepping but the same effect was sometimes achieved by tying a short timer to the NMI - and when the NMI is asserted, the interrupted program counter is pushed onto the stack.
Breakpoints are also easy to spot. Just look for 0xCC in the memory. But of course there are also multiple ways to work around those checks (such as using harder to detect hardware breakpoints instead of software breakpoints).
Another trick to detect breakpoints as an anti-debugging measure was to compute a key over the block of memory that was to be protected. Any int3s in there would give a different output. A countermeasure was to use "memory breakpoints" which work via the x86 debug registers: https://en.wikipedia.org/wiki/X86_debug_register
http://pferrie.host22.com/papers/antidebug.pdf
For a fresh compilation (which includes the previous paper) you can check this article:
http://antukh.com/blog/2015/01/19/malware-techniques-cheat-s...
IsDebuggerPresent() is the most straightforward way: https://msdn.microsoft.com/en-us/library/windows/desktop/ms6...
In the specific case of x86 and int3 bps, you can abuse variable-length instruction encoding to jump into the middle of another instruction. Then code execution can differ depending on the presence of 0xcc.
Fun times then and now :)
i got my start in programming by trying to break various copy protection schemes in MS-DOS games. i used a software debugger that would set int 3 breakpoints, as the article mentions.
BUT, for the debugger to work, it has to rewrite the interrupt vector table so that int 3 instructions will cause the debugger's own code to be executed. really wily code would use the entry for int 3 in the interrupt vector table as a scratch variable to perform a bunch of its own calculations, perhaps thousands of them so you can't find and modify them all, which would ultimately wind up with an address to jump to, to continue execution of the program. this means that, when the debugger rewrites the int 3 vector, the program's internal calculations would be thwarted, effectively stopping the program at that point.
this was really difficult to get past. there were other people doing this kind of thing who seemed to find ways around this technique, but i never did.
If this kind of thing interests you check out 4AM's Apple II cracks. He's removing the copy protection on old software so the techniques aren't directly applicable to current machines, but there's a font of creativity in the code he's reverse engineering. https://twitter.com/a2_4am
However, it didn't explain how the debugger can stop again at the breakpoint after the last step? The interrupt command has been replaced with the original command, so the process won't stop again..
(You can engineer a deadlock in gdb due to this, e.g., on x64, by stepping over a SYSCALL instruction that reads from a pipe that's about to be filled by another thread. But you're unlikely to experience this in practice, as system calls are wrapped by a glibc function, and you'll probably be stepping over that rather than the instruction directly.)
The breakpoint registers are accessed via JTAG/SWD using your j-link/FET/whatever.
Quite often, when you're debugging embedded systems, you run out of "hardware" breakpoints and have to resort to software-style breakpoints described in the article.
The situation is kinda OK when you don't often change breakpoints and your CPU has an instruction register writable by the debug probe via JTAG/SWD/etc. Upon stepping or continuing from a breakpoint, the debugger will write the actual instruction at the breakpoint into the instruction register and tell the CPU "run again, but don't load the instruction from memory as I have already loaded it into your instruction register.".
Another option is to emulate the effects of the instruction in the debugger and write the results back into registers/RAM/I/O. This is not always possible.
If you don't have the options above, your flash will wear down quickly, as stepping away or continuing from a breakpoint entails writing the actual instruction back, stepping, then writing the breakpoint again.
> And these software breakpoints are written into flash.
The GP is referring to the eight FPB hardware breakpoints, which do not need to be written to flash. Most Cortex-M have an additional four hardware breakpoints from the DWT. You can get a lot done with 12 breakpoints before you need to start writing things to flash.
More generally, as long as the debug interface allows you to run with interrupts disabled and has at least one HWBP, you can single-step without writing SWBP to memory.
1. Have a NOR flash part rated for 10,000 or fewer cycles. 2. Have no facility for remapping bad sectors. 3. Have no wear-leveling mechanism.
All of this is reasonable for a product that will only be flashed once during manufacturing (and there are a lot of those products) or a product that will receive a firmware update a single-digit number of times in its lifetime.
In contrast to an SSD, where:
1. NAND flash is used, with 100,000 to 1,000,000 write cycles 2. The drive can transparently remap bad sectors, so flash can start to fail before anybody notices. 3. The drive performs automatic wear leveling - if you try to write a single sector a million times, the drive will do something closer to writing a million sectors once.
Damn, where can you get that kind of flash chips with even 100k cycles endurance? I'd like to place a large order. 64 Gbit chip, please.
1000-3000 P/E cycles is typical endurance for MLC NAND flash, not 100-1000k. SLC chips would fare better, but larger ones are too expensive for typical applications.
One question though: why is NAND flash so much more resilient?
For AVR specifically, have a look at e.g. https://github.com/raimue/avarice/blob/master/src/jtagbp.cc (just random find on github, likely not the official repository).
This is, by the way, also possible on a typical PC (as in "personal computer" not program counter), also x86/64 has debug registers.
http://www.codeproject.com/Articles/28071/Toggle-hardware-da...
...and very likely also on ARM, m68k, MIPS, Sparc, PPC, ... I just didn't look it up.
On Windows this can be done with SEH and Linux has its own thing too.
2. ptrace() modifies the currently debugged program.
3. if the program currently being debugged is sufficiently small, it might end up in the processor's instruction cache.
4. instruction cache invalidation, at least on some processors like MC68000 family, causes crashes.
5. since ptrace() effectively performs the equivalent of self-modifying code, how is instruction cache invalidation avoided?
https://sourceware.org/git/gitweb.cgi?p=binutils-gdb.git;a=b...