30 karma · joined February 15, 2023
Neoverse N1 is hardly a dog. It’s a decent enough core in a server processor that has 80 of them.
Altra is not intended to compete core for core with any laptop/desktop/workstation processor.
In section 4.3 the author incorrectly describes Fermi’s comments to Dyson as a criticism of Dyson’s work on QED. At the time of this meeting, Dyson had already given up on his QED program and was working on the strong interaction. Fermi’s comments that Dyson had neither a “clear physical picture” or a “precise and self-consistent mathematical formalism” were regarding some (pre-quark) pion theory Dyson had been working at at Cornell starting in 1951, not QED.
Dyson describes this meeting in https://www.webofstories.com/play/freeman.dyson/94 but the required context is in the prior segment https://www.webofstories.com/play/freeman.dyson/93
Also the author describes Oppenheimer’s initial resistance to Dyson’s work, but fails to mention that Oppenheimer was eventually convinced that Dyson’s work was valuable and gave Dyson a lifetime appointment at the IAS.
Compare and swap still typically requires the cache line to bounce around between cores, so if that’s the primary cost of locks for you it seems like compare and swap doesn’t really fix the problem of frequent synchronization between cores.
https://developer.nvidia.com/blog/accelerating-standard-c-wi...
> Why not implement your approach and mail it to LKML :-)
because this would still be an in-kernel dwarf unwinder and I would expect an instant reject, and because I am lazy and/or don’t care enough about this problem or linux to work on it. Even if people could be persuaded, I don’t have the interest or temperance to debate this with LKML.
This is also common for samples in leaf functions.
compiler & tool chain folks tend to think (quite justifiably imo) that this and similar stuff is fine because dwarf allows reconstructing everything perfectly. The problem is just that the user experience of dwarf-based unwinding is poor, because the only implemented method in Linux is sampling the contents of the stack and doing the unwind in post processing.
But i agree the impact of preserving frame pointers is generally quite small and doesn’t often actually need mitigation - on amd64 there’s not much impact from losing 1 more of 16 arch registers.
AFAIK frame pointers work fine for unwinding on aarch64. And on aarch64 the gcc default is not to omit frame pointers and IIRC when the default was switched at some point it was treated as a bug and reverted (not sure if required in the ABI or just strongly preferred by the community). So IME generally unwinding with frame pointers on aarch64 works more often than on amd64 since you don’t have to recompile the world.
No. The only reason it works like this is because the upstream Linux kernel has thus far rejected in-kernel dwarf unwinders, but copying the stack is simpler and available / implemented.