How does gdb work?
jvns.ca
jvns.ca
Having written at least 4 complete DWARF readers and writers (GCC's location list support, GDB's expression evaluator, the thing that became google breakpad's debuginfo reader etc), it's really not that bad.
In fact, compared to pretty much any other debug format, it's wonderful. All of the forms are consistent, and outside of the index tables, and a few places where backwards compatibility was needed (ie the world moved from 32 bit to 64 bit, but before that, DWARF supported 64 bit sized debug info on 32 bit processors), the encoding is sane.
libdwarf, on the other hand, is ... not so much. I love david a, and (AFAIK) he's been working on DWARF since SGI was at 1600 amphitheatre parkway, and keeping libdwarf up to date. It's one of those open source projects nobody ever realizes has been around 20 years and that someone has kept it working great (see ftp://ftp.sgi.com/sgi/dev/davea/objectinfo.html)
However, libdwarf is just not a pleasant interface to work with, IMHO.
It's also memory intensive.
If you just want a reader, the thing that made it into breakpad is probably a good reference (others may have better ones, i've thankfully been out of the debug info game for years):
https://chromium.googlesource.com/breakpad/breakpad/+/master...
It's mostly a callback interface, and has a function info reader meant as a demonstration It should work without trouble outside of breakpad (when it was contributed, I made it portable to be able to just compile standalone. it doesn't look like much has changed).
it does not support DWARF4/5, but nobody should need to care.
It also has no expression evaluator but i have a bunch of them if someone needs them :)
Just something that writes DIE's?
Or do you need line info, accelerator tables, etc?
I've actually considered looking more at the LLVM dwarf writing components. Do you have any experience with that?
http://jvns.ca/blog/2016/03/16/tcpdump-is-amazing/ quick introduction to TCP dump
http://jvns.ca/blog/2014/09/27/how-does-sqlite-work-part-1-p... diving into SQLite and sharing the findings with the readers
http://jvns.ca/blog/2014/08/12/what-happens-if-you-write-a-t... implementing TCP in Python
Most of them have valuable HN discussions as well.
http://3.bp.blogspot.com/-J5bsfRdkOdk/UsHqCho2huI/AAAAAAAACV...
and structures:
https://msdnshared.blob.core.windows.net/media/TNBlogsFS/Blo...
* "set print pretty on" , this makes structs readable
* "x/20b &var" to show a hex dump. But no ASCII, and you have to give it the count (e.g. 20 bytes..)
edit-001 : minor formatting updates
We just saw three ranges in ///maps:
5598a9605000-5598a9886000 r-xp 00000000 [...]
5598a9a86000-5598a9a8b000 r--p 00281000 [...]
5598a9a8b000-5598a9a8d000 rw-p 00286000 [...]
what are these? The r??p look like permissions, read/write/execute/something, but how do we know which one to look for the variable in?(And what are the numbers afterwards? 00281000 on the second line is the length of the range on the first line, but then there's a gap of 00200000 between the end of the first range and the start of the second. 00286000 on the third line is again 00200000 less than the distance from the start of the first range and the start of the second.)
In this case, each of those mappings corresponds to one of the sections in the binary. The permissions indicate that the first one is executable code, the second is read-only data, and the third is writable (copy-on-write). The number after the permissions is an offset into the underlying file.
I suspect the article is glossing over some details e.g. how gdb figures out which mapping corresponds to which section, but it gets the basic idea across.
Something like http://www.cs.columbia.edu/~junfeng/09sp-w4118/lectures/int3...
Although unfortunately debug registers are limited to just 4 simultaneous breakpoints in the same time.
http://jvns.ca/blog/2016/06/12/a-weird-system-call-process-v...
I love the JVM's easily trace-ability, though that involves safepoints, so that's not completely out of process either.
Java stems from Sun. Sun made Solaris. And Solaris had an ideology about transparency and discoverability of running systems. They have some outright amazing tools for introspection into a binary that is running. This probably creates an environment in which the same would happen for the JVM. In fact, the `jstack` name seem very familiar to DTrace (and mdb(1)) users.
Days since I last solved a problem on the JVM through a heapDump: 2.
If it follows standard System Linkage, its easy to point gdb or any other system debugger or profiler to debug and profile the application.
Some runtimes have a mix of System and Private linkage, ie C functions will follow System Linkage but JIT'ed code frames might follow private linkage. This makes for difficult stack-walking by system native debuggers and profilers. You'd have to teach GDB via an extension how to walk the non-standard frames.
So yea, long story short, it depends on the linkage convention the implementers of the language runtime decided to follow.
* "Python" in general might mean you're on Linux/Windows/whatever, and it might mean CPython, PyPy, or some other runtime. But any out-of-process instrumentation is gonna have to be pretty platform/runtime specific.
* Even if we restrict ourselves to, say, CPython on Linux, the interpreter's internals aren't super friendly to this sort of inspection from the outside. You have to rely on and also work around implementation details.
Example: to get a Python call stack, you want to look at `PyThreadState_Current` (basically the same idea as `ruby_current_thread` in that excellent linked post of Julia's, I think). But this happens to be null whenever the GIL is released, e.g. when doing network I/O, and then you're kind of out of luck. So you'll already have trouble usefully profiling a single-threaded I/O-intensive program.
* Oh and you pretty much need debug symbols in your CPython binary (I think? Tell me if this isn't true!). Most production CPython builds don't have them. So you have to get the right binary, and rebuild any application dependencies with C extensions. Not hard but annoying.
There is potential though! With some work, we definitely could have a better story for out-of-process Python profiling a la Linux perf.
Also, I've seen several comments here of someone making what I presume are honest mistakes in using the wrong pronoun and being publicly shamed for it.
It doesn't to me. Maybe it varies by region.
For example, what would you refer to me as? "Nadya wrote on her blog" or "Nadya wrote on his blog"? (hint: You'd need to dig rather deep into my post history to find the "proper" one.)
@ruraljuror's example
Further context removes that ambiguity. Although you could rephrase it as "Julia wrote on their own blog" or even remove the pronoun altogether: "Julia wrote on Julia's blog" (which is only ambiguous if there is another Julia).
And Nadya is hard—I would go for gender neutral on that one. :-) (It's not that I dislike singular ‘they’, but sometimes it can make sentences awkward.)
I haven't considered the situation you bring up. I am for the singular they when it is being used to describe and indeterminate person. Are you suggesting we use singular they for a particular person until we get explicit confirmation about their pronoun preference?
I wonder if you weren't downvoted for using the subjunctive mood in your comment ("if that were the intention here") which seems to imply that you believe it was not the intention and was instead a "mistake".
In my opinion this starts to feels very self self-promotional.
As someone who enjoys variety and diversity I think it's a valid question.
The domain jvns.ca gets submitted by a variety of people. (Lots from ingve, but that user submits a lot of other stuff too.)
So it's probably that people post it, and other people upvote it.
Do you think it's something that is interesting?
It's fine for you not to find every post on HN interesting or discussion-worthy. There are other people who are interested and having a good discussion.
Many (maybe "more interesting") articles about deep details of something don't get much attention because you need a lot of previous knowledge to understand what is going on, that hurdle almost never happens with her articles. Introductory articles often get more attention that way.
And sometimes I notice that apparently I didn't know as much as I thought about the details...
This article isn't stuff I didn't already know in some form, but now it's made me want to write a debugger! :)
This just reinforces my suspicion that that this blogger's friends and coworkers are the reason a single blog constantly ends up on the front page of HN regardless of merit.
You really think her blog posts have undeserved merit? :-)
"I suspect vote rigging" accusations aren't nice. You should send them to mods at the email address rather than post them to the thread.
There are lots of books we could all read instead, but she uses a great mix of explanations and code and has an approachable reading style.
She also uses tools that we might want to use, so it's less abstract.