HNHacker News
TopNewBestAskShowJobs

tbr1

141 karma · joined March 26, 2020

submissionscomments
tbr1··on Magic-trace – High-resolution traces of what a process is doing
I think this answer has several parts:

- I imagine the extra memory bandwidth of newer parts doesn't hurt. The example traces were taken on server-class Ice Lake machines. They just don't overflow for our typical workloads.

- We found the specific IPT configuration matters a lot. Turning off return compression is more liable to result in overflows. We allow varying this in magic-trace via the `-timing-resolution` parameter, more detail available in the wiki. We don't typically see overflows under the default configuration even on Broadwell server-class parts.

- Clark spent a week on an Intel NUC (mobile Tiger Lake part) toiling away on decode error recovery. For the most part, the data lost are uninteresting branches, and you only need one of the call in / return out of a frame to survive the decode error to be able to construct a frame for it.

We also considered the periodic stack sampling approach for error recovery, but ended up not implementing it since the decode error recovery we implemented ended up being robust enough in practice.

We ended up having more trouble with runtimes that mess with the stack pointer directly. (The kernel does this for the retpoline Spectre mitigation! But perf is smart and rewrites that part of the instruction stream into a jump for us.) There's code in magic-trace to special-case OCaml exceptions, for instance, and it's likely similar code is necessary for some other runtimes too (we have an open issue for Go's coroutine switching).

tbr1··on Magic-trace – High-resolution traces of what a process is doing
DDIO operates mostly transparently to software, with the I/O controller feeding DMAs into a slice of L3. Hardware can opt out by setting PCIe TLP header hints, and you have some system-wide configurability via MSRs, but it's not something a userspace application can take into its own hands.
tbr1··on Magic-trace – High-resolution traces of what a process is doing
Absolutely, check out https://github.com/janestreet/magic-trace#privacy-policy and https://github.com/janestreet/magic-trace/wiki/Setting-up-a-.... With a bit of extra configuration, magic-trace can host its own UI locally. You just need to build the UI from source, and point magic-trace to it (via an environment variable).
tbr1··on Magic-trace – High-resolution traces of what a process is doing
Yes, in fact this is how we've been narrowing down performance problems in it and its dependencies :)

- https://github.com/let-def/owee/issues/23

- https://github.com/janestreet/magic-trace/issues/93

tbr1··on Magic-trace – High-resolution traces of what a process is doing
Absolutely! This is one of the main features of magic-trace, and in fact a primary use-case.

You can select a trigger symbol for magic-trace to snapshot upon the next call of. This can be whatever you want, and you can imagine writing code like

  if (something_really_wonky_happened) { take_magic_trace(); }
and asking magic-trace to take a snapshot of the past only when `take_magic_trace` is called.
tbr1··on Magic-trace – High-resolution traces of what a process is doing
It works best on compiled programs.

We do try to support scripted languages with JITs that can emit info about what symbol is located where [1]. Notably, this more or less works for Node.js. It'll work somewhat for Python in that you'll see the Python interpreter frames (probably uninteresting), but you will see any ffi calls (e.g., numpy) with proper stacks.

[1]: https://github.com/torvalds/linux/blob/master/tools/perf/Doc...

tbr1··on Magic-trace – High-resolution traces of what a process is doing
It's worth noting that aside from the overhead, function call / returns are not quite enough to reconstruct the callstack: tailcalls are just regular branch instructions.
tbr1··on Magic-trace – High-resolution traces of what a process is doing
We don't have plans to add ARM support largely because we have no in-house expertise with ARM. That said, ARM has CoreSight which sounds like it could support something like magic-trace in some form, and we'd definitely be open to community contributions for CoreSight support in magic-trace.
tbr1··on Magic-trace – High-resolution traces of what a process is doing
You may also be interested in this wiki page: <https://github.com/janestreet/magic-trace/wiki/Supported-pla...>

Intel PT has a bunch of rough edges that we've tried to paper over in magic-trace, but the gritty caveats are documented in the wiki.

tbr1··on Magic-trace – High-resolution traces of what a process is doing
It's all OCaml, GitHub is just misclassifying it as SML :)
tbr1··on Magic-trace – High-resolution traces of what a process is doing
We have a bit more color on compatibility in general up on <https://github.com/janestreet/magic-trace/wiki/How-could-mag...> for those interested.
tbr1··on Magic-trace – High-resolution traces of what a process is doing
One of the maintainers here -- it should work on Broadwell if you're not super keen on the tens-of-nanoseconds timing precision and are okay with microsecond precision (i.e. only want accurate callstacks),

  grep intel_pt /proc/cpuinfo
should do the trick.
tbr1··on Zero-copy network transmission with io_uring (2021)
Small correction: DDIO is not limited to Intel NICs, it's a mostly-transparent-to-the-hardware mechanism by which DMAs are snooped and some fraction of the NIC-local (in the case of multi socket systems) L3 cache is filled with incoming data.
tbr1··on Looking Glass: Run a Windows VM on Linux in a window with native performance
It's quite easy to get it set up once you have a VM with GPU passthrough running (for which there are plenty of guides available online) -- just a double-click installation of a service on the Windows side, compiling a cmake project on the Linux side, and (optionally, for some extra performance) compiling a Linux kernel module.

After that it kind of just works, and continues working. I use it to play games and run Office apps, and have not had it break on me in a ~year of use. (Disclaimer: I occasionally contribute to the project now, but remember being impressed at how easy it was to get going when I first tried it out. Getting the VM working at all was the hardest part of the endeavor, but only took a few hours.)

tbr1··on Looking Glass: Run a Windows VM on Linux in a window with native performance
Yes, QXL can do that and has built-in support in e.g. virt-manager. It's not going to be fast, but it does work, and is how installation of Looking Glass can be bootstrapped without a physical monitor attached.
tbr1··on How a Bad Random Number Generator Froze Sway (2020)
(Sway dev mentioned in the article here.)

To be clear, this loop doesn't exist in Sway proper: it lives in json-c, a third-party library used by a lot more than just Sway. A lockup bug was reported to a lot of projects using json-c at the same time.

This really wasn't a Sway/Wayland thing, we just happened to be the first to diagnose and contribute a fix upstream.

tbr1··on I'm tired of this anti-Wayland horseshit
He also wrote "the book" on the subject: https://wayland-book.com/
tbr1··on I'm tired of this anti-Wayland horseshit
"Wayland forwarding": https://gitlab.freedesktop.org/mstoeckl/waypipe

VNC: https://github.com/any1/wayvnc

tbr1··on The X.Org Server Is Abandonware?
Mir is a Wayland compositor. It may not have started off as one, but is today.
tbr1··on The X.Org Server Is Abandonware?
Xorg has had a meson build system since 2017: https://gitlab.freedesktop.org/xorg/xserver/-/commit/1549e30...

Are you referring to something else?

tbr1··on The X.Org Server Is Abandonware?
To make the comparison closer to apples-to-apples, the Wayland analog to the Xorg server would be something like GNOME's mutter compositor, which had its first Wayland support out in 2013[1] -- 7 years ago. And the rate of progress has only sped up since then -- take a Wayland compositor from a year ago and compare it to the same one today, and things tend to be much more polished.

[1]: judging by https://wiki.gnome.org/Initiatives/Wayland

tbr1··on The X.Org Server Is Abandonware?
The v4l2 trick "works", but usually the application will use a lossy video codec optimized for faces, not screens. wf-recorder to v4l2 as a poor-man's screenshare to Discord ends in a blurry mess. YMMV with Teams.
tbr1··on The X.Org Server Is Abandonware?
Developers willing/able to work on X11/Wayland plumbing are a (very) finite number. A single developer can have a large impact. With more of them switching their primary efforts to Wayland, the past few years have seen large improvements in the landscape, and this trend is bound to continue.

Granted, there are rough edges, and I wouldn't claim any Wayland compositor is as polished as an X11 one -- but we're not that far off, and for many people the benefits of running a Wayland session today outweigh the cons.

tbr1··on The X.Org Server Is Abandonware?
This comment is mostly correct (as a daily Wayland user), with a few exceptions.

> Wayland does almost nothing besides render buffer handling. Input? Applications job.

Applications don't do more work to handle input on Wayland as opposed to e.g. X11. It's still event-based, and the compositor feeds input events to applications that can process them as normal. Keyboard, mouse and touch input are part of the core Wayland protocol, and tablet input is part of an extension that all major compositors fully support.

> Window decorations? Compositors job.

Kind of, it's the job of the application (client-side decorations) or compositor (server-side decorations). The compositor can choose which to use. CSDs give more custom look-and-feels to applications that have them (think Firefox or Chrome); SSDs provide consistent looks across all apps. GNOME only supports CSDs, but is an exception in that regard.

> Clipboard? Maybe compositor or toolkit.

Both: https://emersion.fr/blog/2020/wayland-clipboard-drag-and-dro...

Clipboards are (implementation-complexity-wise) scary in X11 as well.

tbr1··on The X.Org Server Is Abandonware?
> ChromeOS runs on Wayland

Kind of. The Chrome browser itself doesn't run on Wayland, it runs on a custom compositor (I believe Aura?). Sommelier[1] is a Wayland compositor used for Linux apps on CrOS (Crostini), but Chrome doesn't use it. There is an ongoing effort (Lacros) to make the Chrome browser itself run under Wayland on CrOS, but it's not public outside of development builds (and not yet on par with the "native" version).

[1]: https://chromium.googlesource.com/chromiumos/platform2/+/HEA...

tbr1··on Jane Street and the OCaml Compiler (2018) [video]
You can write zero-alloc OCaml code, since the standard compiler is very predictable in when it will allocate closures, etc. The Flambda compiler makes things even nicer in removing allocations for some common idioms.

[1]: https://twitter.com/yminsky/status/947064713237684224