Magic-trace – High-resolution traces of what a process is doing
github.com
github.com
magic-trace was submitted before, our first announcement was this blog post: https://blog.janestreet.com/magic-trace/.
Since then, we've worked hard at making magic-trace more accessible to outside users. We've heard stories of people thinking this was cool before but being unable to even find a download link.
I'm posting this here because we just released "version 1.0", which is the first version that we think is sufficiently user-friendly for it to be worth your time to experiment with.
And uhh... sorry in advance if you run into hiccups despite our best efforts. Going from dozens of internal users to anyone and everyone is bound to discover new corners we haven't considered yet. Let us know if you run into trouble!
echo '*.ml linguist-language=OCaml' >> .gitattributes
echo '*.mli linguist-language=OCaml' >> .gitattributes
[1] https://github.com/github/linguist/blob/master/docs/override...Any plans to support Arm in the future? Thanks!
We do try to support scripted languages with JITs that can emit info about what symbol is located where [1]. Notably, this more or less works for Node.js. It'll work somewhat for Python in that you'll see the Python interpreter frames (probably uninteresting), but you will see any ffi calls (e.g., numpy) with proper stacks.
[1]: https://github.com/torvalds/linux/blob/master/tools/perf/Doc...
PIX is great! It gets regular updates and an active and responsive Discord channel.
Yup, PIX is THE tool for game developers. Direct3D team also has a very responsive discord channel :)
I came across the incr_dom library [1], which efficiently calculates the diff on the projected DOM based on a diff of the model data, sorta like React but more... mathematically grounded (?). incr_dom was then reformulated in Bonsai [2] which refactors and generalizes the idea to work with more than DOM. There was a recent Signals and Threads podcast about it a few months ago [3].
[0]: https://news.ycombinator.com/item?id=31100023
[1]: https://opensource.janestreet.com/incr_dom/
https://man7.org/linux/man-pages/man1/perf-intel-pt.1.html
> Intel Processor Trace is an extension of Intel Architecture that collects information about software execution such as control flow, execution modes and timings and formats it into highly compressed binary packets. […] The main distinguishing feature of Intel Processor Trace is that the decoder can determine the exact flow of software execution.
I worked on a very similar (if not identical lol) project at a job once upon a time and the biggest problem I had (and one that I never really solved well) was recovering call stacks from trace data. I effectively ended up using DWARF and just simulating execution and keeping a call stack in the decoder. This mostly worked fine for small and simple programs, but I ran into SO MUCH trouble because I found that (at least on my generation of cores) IPT actually overflows and drops packets very frequently if you have too many calls/returns too quickly. This is largely not an issue for C code but once you start getting into more dynamic languages with fancy features, IPT cannot keep up. Once packets get dropped, the entire call stack for the entire rest of the thread is ruined since you have no idea who called/returned in the dropped packets.
One option that we had but didn't really chase down due to time was maybe combining IPT with low frequency stack traces so that we can both just reset every so often and, if needed, work backwards/apply heuristics in order to arrive at that next callstack.
How did y'all manage this? Your call stacks look totally correct and I'm very impressed :)
- I imagine the extra memory bandwidth of newer parts doesn't hurt. The example traces were taken on server-class Ice Lake machines. They just don't overflow for our typical workloads.
- We found the specific IPT configuration matters a lot. Turning off return compression is more liable to result in overflows. We allow varying this in magic-trace via the `-timing-resolution` parameter, more detail available in the wiki. We don't typically see overflows under the default configuration even on Broadwell server-class parts.
- Clark spent a week on an Intel NUC (mobile Tiger Lake part) toiling away on decode error recovery. For the most part, the data lost are uninteresting branches, and you only need one of the call in / return out of a frame to survive the decode error to be able to construct a frame for it.
We also considered the periodic stack sampling approach for error recovery, but ended up not implementing it since the decode error recovery we implemented ended up being robust enough in practice.
We ended up having more trouble with runtimes that mess with the stack pointer directly. (The kernel does this for the retpoline Spectre mitigation! But perf is smart and rewrites that part of the instruction stream into a jump for us.) There's code in magic-trace to special-case OCaml exceptions, for instance, and it's likely similar code is necessary for some other runtimes too (we have an open issue for Go's coroutine switching).
https://github.com/janestreet/perfetto/
It is mentioned in the documentation but for anyone quickly skimming and expecting a SaaS pricing model underneath (I think many do now), that isn't obvious. From my initial scroll through this wasn't obvious and it was significantly less attractive. Looks very interesting!
I don't like that this comment is hidden on the bottom of the page, as my first impression of the page was that the work of creating the high frequency trace was done by Jane Street (I don't like it in the other direction either, when a big company ,,rebrands'' what a person/small company does).
perfetto is "just" the profiler UI.
I don't think I can say I support their company's primary mission, but their commitment to improving the world of software through various means (language influence, academic publication, open-source software releases, etc) is admirable and well worth respecting.
I don't think that tracks. They like OCaml, and they are pretty adamant that it is a good tool for the job. Maybe you disagree, but you should not project your opinions on them.
I'm not sure what you mean by this. What is the situation?
> Being stuck on a dead tech is even worse.
Why do you think OCaml is a "dead tech"? Can you justify that? Or is it just based on the notion that if most people don't use it, it must be a bad tool?
Cost is more than just monetary. There are significant indirect costs as well.
Can you elaborate on these costs? Do you have knowledge of JS's internal needs and resources to suggest a better alternative?
They are not just hapless consumers of a dead language; they actively maintain it and invest in it because it works well for them. The language itself gives them the kind of guarantees they want in their work, and their work on the language and surrounding tooling (among other things) helps them to acquire high-skill talent. I don't know how you can claim that they would transition to another language if they could without having some pretty firm data to back that claim up. Otherwise, I think you're just projecting your own feelings about OCaml onto them.
Nowhere did I mention that Ocaml was a dead language or a dying language. I simply stated that there are insurmountable switching costs which incentivizes them to contribute to the larger Ocaml community.
If all they wanted was to hire people, they... would. You don't have to sponsor a conference to attend or hire from that conference. And you especially also don't need to sponsor additional workshops, or carbon neutrality initiatives, or anything else.
It's genuinely silly to suggest that they spend all this money on things just to hire people. There are so many more effective uses of their money if that is the only goal.
Because they believe in the value of science?
They don't need to publish to compete with top-tier public companies. I don't know of any other trading firm that contributes to open-source development, or to a language infrastructure, or to academic advancement in the way and to the extent that Jane Street do. Most companies keep everything proprietary and highly secret.
But JS chooses to publish. And it's not like that's an easy task that you can just do for fun on a whim; it takes a long time to put together a good paper. They also regularly collaborate with people in academia on long-term projects and evaluations.
I understand the perspective you're suggesting, but I genuinely believe it is wrong, and I also believe that you do not have any evidence to back it up. I think their public contributions speak for themselves, but I've also met some of their more academically inclined engineers (including their CTO), and they come from an academic background and seem to genuinely believe in academic publishing as a goal in itself. It's not totally crazy that there exists one such company out there. (There are actually a couple, but not terribly many, and the others are not relevant in the present discussion.)
https://signalsandthreads.com/
> Listen in on Jane Street’s Ron Minsky as he has conversations with engineers working on everything from clock synchronization to reliable multicast, build systems to reconfigurable hardware. Get a peek at how Jane Street approaches problems, and how those ideas relate to tech more broadly.
Do you have a basis for that claim?
JS have developed tons of libraries and tools for OCaml development, and new developers and quants that they hire go through an OCaml bootcamp to come up to speed. They put lots of work into the OCaml compiler, and in blog posts about that work they talk about why this is useful for trading. Maybe I'm missing something crucial, but I think it's more likely that you just don't know what you're talking about.
There is, of course, something to be said about the by-products of their work. Jane Street is far from an evil company, and I would not be entirely morally opposed to working for or with them. They do a lot of good in academic research in areas I care about. I just wish that that was their primary purpose instead of direct money-making, if that makes sense.
- We use cash to buy circuit boards, screens, enclosures, etc, write software, and sell mobile phones.
- We use cash to rent a building, order pallets of inventory, and sell that inventory locally to walk-in customers.
- We use cash to buy shares, hold onto them for a bit, and sell those same shares and make money off the spread.
I'm not making any kind of comment at all about the value of market makers, just... those three businesses feel like they're different models.
Not exploiting teenagers in some 3rd world country?
Not gambling with your pension?
Not manipulating some physical commodity like oil?
What’s the problem here?
> Not gambling with your pension?
> Not manipulating some physical commodity like oil?
There is no reason to believe #1 and #3 aren't true, and I should very much suspect they are. #2 is not possible as far as I can tell, I agree there.
> - We use cash to buy shares, hold onto them for a bit, and sell those same shares and make money off the spread.
Those two sound like pretty much the same thing.
In the case of a trading shop, the stuff they do is playing the market liquidity, collecting interest, arbitrage, etc...
Sure there are some evil ones, but other businesses have those too.
Jane Street is a proprietary trading firm. They have no external customers to whom they would provide goods or services. Their primary purpose is to invest the company's own money.
cat /sys/devices/cpu/caps/pmu_name
to find if they're invited to the party grep intel_pt /proc/cpuinfo
should do the trick.Intel PT has a bunch of rough edges that we've tried to paper over in magic-trace, but the gritty caveats are documented in the wiki.
Glad to see "overhead" mentioned and quantified. I'd put the 2-10% at the top though, as that's heavy handed for some environments (can trigger a production fail-over).
I see magic-trace has implemented what some call "flame charts" (time on the x-axis) and not "flame graphs" (alphabet on the x-axis). The best tools do both (e.g., TraceCompass). Please do both! Will make seeing the big picture easy (flame graphs) and then zooming into time-based patterns easy too (flame charts).
Good point about overhead. I've moved the 2%-10% number front and center, and wrote up a bit more detail about where that comes from in a new wiki page: https://github.com/janestreet/magic-trace/wiki/Overhead
We'll think about adding flame graphs. We unfortunately have little experience writing responsive web UIs, the excellent Perfetto developers did all of the heavy lifting on that front. But who knows, maybe an enterprising Open Source Contributor could help us out. I see Matt Godbolt was asking questions in their discord the other day...
The Perfetto UI already supports flamegraphs btw (we use it for memory profiling and CPU stack sampling). We've never bothered to implement it for userspace slices because we've never had high frequency data there to make that a worthwhile view of the data.
Contributions for this upstream are very welcome :)
A quick way to check if PMU access is enabled is this:
dmesg | grep "Performance Events"
Edit: Oh, unfortunately VMware Fusion 12 (on Mac OSX Intel) does not expose the performance counters to the VM anymore, as it's using the OSX Hypervisor Framework instead of its own kernel module.There's also the regular performance traces you can capture with wpr and friends. I don't think these provide function-level traces, and I also don't think it's possible to do that (but I could be wrong). You just get sampled callstacks, which may or may not be enough for your needs.
In my experience on Windows you need to instrument applications to get function-level tracing.
> 2. perf can do this too, but that's not how most people use it. In fact, if you peek under the hood you'll see that magic-trace uses perf to drive Intel PT.
I think this (the first sentence quoted) is a bit misleading. The main feature is not really a "key difference from perf" if the main feature is implemented using perf. From a brief read, it looks like the real key difference is a friendlier and more interactive UI (both when capturing and viewing the trace).
Regardless, I think it looks neat, and will try to take it for a spin sometime soon.
We shouldn't need to scroll through pages of text to figure that out.
[0] https://www.intel.com/content/www/us/en/io/data-direct-i-o-t...
There was a great presentation from 2017 about some of Optiver's low latency techniques[1]. I had assumed they released it because the had obviated all of them by switching to FPGAs, but I don't know. Either way, he suggested that if you ever needed to ping main memory for anything, you already lost. So, I wouldn't have thought DDIO plays into their thinking much.
You can select a trigger symbol for magic-trace to snapshot upon the next call of. This can be whatever you want, and you can imagine writing code like
if (something_really_wonky_happened) { take_magic_trace(); }
and asking magic-trace to take a snapshot of the past only when `take_magic_trace` is called.Is it possible to combine PMU or sampling with IPT to get multiple profiling dimensions in the same run? Not just what sequence of instructions were executed but where in time-and-code the branch mispredictions, cache misses, etc. occurred?
In fact, if you look carefully at the demo gifs in the README, that trace had 5 decode errors! Nonetheless, it was extremely usable.
Snapshot sizes are configurable--you can go back as far as you like. However, the trace viewer tends to crash when the trace files reach the hundreds of MB and you'll need to do some work to set up a trace processor outside of your browser for the UI to connect to. The UI will offer up some docs if you actually run into this.
I'm so glad you asked us about PMU events, we've been thinking a lot about those. These are available in traces of the efficiency cores of Alder Lake CPUs, but nothing else. When we get our hands on a server class part with PMU tracing we'll add support ASAP. We conjecture that it will be absurdly useful to see cache events on a timeline next to call stacks.
magic-trace uses perf. If you want, you can think of it as a mere "alternative frontend" for the Intel PT decoding offered by perf.
First no-kidding application I've seen in that language.
If yes, then yes that is much much higher overhead than processor trace.
1) A program that inserted a macro invocation with a GUID, at every single new scope. "{", except not switch, struct, or class, etc.
2) I made the macro instantiate an object on the stack, passing in the GUID. In the constructor and in the destructor, it called a singleton with thread-local storage to a file pointer, where it would append a few bytes indicating whether it was entering or exciting scope, what the GUID was, and what the time was. In this way, each thread was writing to its own file.
3) I made a program which walked my source, looking for file name, line number for each of my macro invocations, and the GUID. If I were more sophisticated, I would have tried to get the function name out of it, too.
4) I made a program which would turn one of the thread's files into a Visual Studio recognized output. Basically "filename(linenumber): [content, such as the time]". Then I set that up as a tool in Visual Studio, and when I would run it, it would output in a Visual Studio window. The reason for that was then I could hit (I think) F4 and Shift-F4 to step forward and backward through the output, and each time it would jump to the source code at that location.
So then I had a forward-and-backward time travelling debug script. I think I also started manually passing in function parameters into the macros, which would format (a: "a") on the debug line, too.
We had automated testing of our whole integrated application. I wanted to record my output from each automated test. Then when I was checking in new code, I could see which new GUIDs were never touched by any of our integration tests, to tell me how much coverage we had. And I could tell the testers which automated tests were most likely to exercise my code changes.
I liked that the GUIDs would have been stable, even as code moved. (Unlike file name, line number, or even class and function name.)
And yes, seeing this in the code wasn't great:
{ TIMER("5c7c062f-84a3-40d0-b7cd-77bd9db59f3e");
// real code
}I wanted to teach Visual Studio how to basically ignore those, and if I copied code and pasted it, have it generate new GUIDs when I pasted.
But I could imagine using the output to generate the fire charts, and other debugging tools, like in the article.
And it all compiled to 0, in Release mode.
The payoff of this felt large, and the cost felt small. But the biggest pain was that humans would see these macro invocations, and need to maintain them.
So I chickened out and didn't force my coworkers to see all of this.