965 karma · joined April 4, 2013
https://github.com/tpn
https://pyparallel.org/
https://speakerdeck.com/trent/
https://twitter.com/trentnelson
https://reddit.com/u/trentnelson
That was a nightmare to triage back in the late 90s when I did it. Thankfully Ctrl-Q (I think it’s Q) “resumes”, so, easy fix if you know what you’ve done.
There’s never been a POSIX equivalent to this. It requires sophisticated kernel support and the exact same parity can’t be achieved in user space alone.
So if you had mmap heavy read/write workloads… you could do some pretty cool stuff.
Tried to make the article as readable as possible on mobile, tablet, and desktop. Mobile necessitated a smaller font size for the code to obviate the need for horizontal scrolling.
Light/dark mode is supported, and the images are even cognizant of the selected mode!
I am doing a talk at PyData Seattle this year (Nov 7-9) focused on this topic, so any feedback regarding additional areas of interest would be appreciated.
Hovering was so humbling! You’d be stable for a few seconds and then oops now we’re suddenly crabbing backwards whilst rolling laterally whilst exacerbating everything with pilot-induced oscillations in every conceivable axis of movement.
Having to constantly enter three inputs whenever the external environment changes (ie wind, gust), or any time any one of the three inputs change… it absolutely requires some new neural pathways to be forged!
I flew with Patty Wagstaff many years later and even she admitted hovering was so hard, to the point it looked like she wasn’t going to be able proceed with her rotor license (before it all clicked).
(The CPU algos for doing this work great on CPUs but are woeful on GPUs.)
Nor do they have anything like IOCP, where the kernel is aware of the number of threads servicing a completion port, and can make sure you only have as many threads running as there are underlying cores, avoiding context switches. If you write your programs to leverage these facilities (which are very unique to NT), you can max perform your hardware very nicely.
And yeah, IOCP has implicit awareness of concurrency, and can schedule optimal threads to service a port automatically. There hasn’t been a way to do that on UNIX until io_uring.
If you have timestamped module load/unload info with base address + range, plus context switch times that allow you to figure out which specific thread & address space was running at any given CPU node ID + point in time, you can always answer that question. (Assuming the debug infrastructure is robust enough to map any given IP to one specific function, which it should be able to do, even if the optimizer has hoisted out cold paths into separate, non-contiguous areas.)
I realize this isn't very helpful to you on Linux (if it's any consolation I'm on Linux these days too), but, sometimes it's interesting to know how other platforms handle it.
(There's no other out-of-the-box solution to this right? i.e. having symbol info live somewhere else other than the .so/exe, that can be loaded on demand when debugging? Like .pdbs basically.)
I wrote a plugin for a past employer to visualize our internal product event hierarchy performance as if it were a normal C/C++ call stack, it was pretty cool. ETW and WPA are phenomenal tools. I miss them both dearly when on Linux.
200 9.1GB SCSI disks for 1.8TB!
And still only 4GB RAM on that SQL Server 6.5 box they used for TPC-C. Wild.
And yeah, $770k for that server.
Edit: I guess whilst I'm at it...
Data General AV8600: https://www.tpc.org/results/individual_results/dg/dg.8600.es...
HP NetServer LXr Pro8: https://www.1000bit.it/ad/bro/hp/netserverlxrpor8.pdf
I really do like NT internals though.
# C++ stdlib
skip -gfi /home/trent/mambaforge/envs/td/x86_64-conda-linux-gnu/include/c++/10.3.0/\*
skip -gfi /home/trent/mambaforge/envs/td/x86_64-conda-linux-gnu/include/c++/10.3.0/*/*
skip -gfi /home/trent/mambaforge/envs/td/x86_64-conda-linux-gnu/include/c++/12.3.0/\*
skip -gfi /home/trent/mambaforge/envs/td/x86_64-conda-linux-gnu/include/c++/12.3.0/*/*
# tl::expected
skip -gfi /home/trent/.cache/cpm/expected/5acc53468c550d1f25ce819a675b60bfa0bbc69d/include/tl/\*
It's not perfect, but at least I don't have to step through annoying things like std::unique_ptr<Foo>.get() a million times whilst debugging.It's geared to generating perfect hash tables with the fastest possible lookup/index times (for 32-bit keys), for key sets in the <=100,000 range. (It scales well up to millions of keys, but the solving time takes a lot longer.)