HNHacker News
TopNewBestAskShowJobs

trentnelson

965 karma · joined April 4, 2013

https://trent.me/

https://github.com/tpn

https://pyparallel.org/

https://speakerdeck.com/trent/

https://twitter.com/trentnelson

https://reddit.com/u/trentnelson

submissionscomments
trentnelson··on Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?
May need to be fail-closed.
trentnelson··on Linux 7.3 improves performance when running out of vRAM
Possibly an unpopular opinion on this forum, but virtual memory has always been one of the strengths of the NT kernel. (And integration between the cache manager, virtual memory manager, file system (NTFS), overlapped I/O, completion ports, and threading. It’s a very robust and performant foundation when leveraged correctly.)
trentnelson··on A Practical Guide to SSH Tunnels: Local and Remote Port Forwarding
Not ssh related but I regularly suspend my terminal with Ctrl-S by accident, usually when going for Ctrl-C/V.

That was a nightmare to triage back in the late 90s when I did it. Thankfully Ctrl-Q (I think it’s Q) “resumes”, so, easy fix if you know what you’ve done.

trentnelson··on Wine 11 rewrites how Linux runs Windows games at kernel with massive speed gains
WaitForMultipleObjects is fascinating behind the scenes. A single thread can wait on up to 64 independent events, which is done by plumbing the KTHREAD data structure with literally 64 slots for dispatcher header stuff, plus all the supporting Ke/dispatcher logic in the kernel.

There’s never been a POSIX equivalent to this. It requires sophisticated kernel support and the exact same parity can’t be achieved in user space alone.

trentnelson··on What makes Intel Optane stand out (2023)
When the PDIMMs were used with an appropriate file system + kernel, it was pretty cool. NTFS + DAX + kernel support yielded a file system where mmap’ing didn’t page fault. No page faults because the file content is already there, instantly.

So if you had mmap heavy read/write workloads… you could do some pretty cool stuff.

trentnelson··on Fontcrafter: Turn Your Handwriting into a Real Font
Well now I’m curious how they did it in the 90s. Some poor schmo doing pixel by pixel font creation?
trentnelson··on Prism
If you’ve got an existing paragraph written that you just know could be rephrased more eloquently, and can describe the type of rephrasing/restructuring you want… LLMs absolutely slap at that.
trentnelson··on OpenBSD-current now runs as guest under Apple Hypervisor
I mean to be fair, WSL1 and WSL2 are extremely successful engineering efforts by Microsoft. I can’t imagine having to go back to the Cygwin days.
trentnelson··on PyTorch and Python Free-Threading
I finished this article in February this year, just before joining NVIDIA. It didn't get officially published then for... reasons. Posting now despite some of the information being a little out of date as I still think the content might be useful to others.

Tried to make the article as readable as possible on mobile, tablet, and desktop. Mobile necessitated a smaller font size for the code to obviate the need for horizontal scrolling.

Light/dark mode is supported, and the images are even cognizant of the selected mode!

I am doing a talk at PyData Seattle this year (Nov 7-9) focused on this topic, so any feedback regarding additional areas of interest would be appreciated.

trentnelson··on The Dawn of Nvidia's Technology
Oh man, the Abit motherboards! That takes me back. How much did this cost and at what time? Presume very late 90s.
trentnelson··on Windows NT for GameCube/Wii
I remember my first job in 2000, straight out of 1.5 years of college, getting to play directly with Digital UNIX and Alpha processors! The Alpha 21264 was a beast at the time.
trentnelson··on Qwen2.5-1M: Deploy your own Qwen with context length up to 1M tokens
Based on an earlier comment, I think the person you're replying to is the author of aider.
trentnelson··on Neuroplasticity in F16 fighter jet pilots
It’s insane how hard hovering is. I had about 35 hours of fixed wing time, and treated myself to a helicopter lesson for my birthday.

Hovering was so humbling! You’d be stable for a few seconds and then oops now we’re suddenly crabbing backwards whilst rolling laterally whilst exacerbating everything with pilot-induced oscillations in every conceivable axis of movement.

Having to constantly enter three inputs whenever the external environment changes (ie wind, gust), or any time any one of the three inputs change… it absolutely requires some new neural pathways to be forged!

I flew with Patty Wagstaff many years later and even she admitted hovering was so hard, to the point it looked like she wasn’t going to be able proceed with her rotor license (before it all clicked).

trentnelson··on Initial CUDA Performance Lessons
Had any exposure to r=2 hypergraph implementations on the GPU? Ideally with an efficient way to determine if the graph is acyclic?

(The CPU algos for doing this work great on CPUs but are woeful on GPUs.)

trentnelson··on Windows NT vs. Unix: A design comparison
None of the UNIXes have the notion of WriteFile with an OVERLAPPED structure, that’s the key to NT’s asynchronous I/O.

Nor do they have anything like IOCP, where the kernel is aware of the number of threads servicing a completion port, and can make sure you only have as many threads running as there are underlying cores, avoiding context switches. If you write your programs to leverage these facilities (which are very unique to NT), you can max perform your hardware very nicely.

trentnelson··on Windows NT vs. Unix: A design comparison
Yeah I’d definitely include RegisteredIO and IoRing. When I was interviewing at Microsoft a few years back, I was actually interviewed by the chap that wrote RegisteredIO! Thought that was neat.
trentnelson··on Windows NT vs. Unix: A design comparison
I should do an updated version of that deck with io_uring and sans the PyParallel element. I still think it’s a good resource for depicting the differences in I/O between NT & UNIX.

And yeah, IOCP has implicit awareness of concurrency, and can schedule optimal threads to service a port automatically. There hasn’t been a way to do that on UNIX until io_uring.

trentnelson··on How to build highly-debuggable C++ binaries
FWIW, on Windows, the ETW event instrumentation that captures dispatch (i.e. thread scheduling) and loader info (I think it's literally the DISPATCH+LOADER flags to xperf) solves this problem, which, inherently is: at any arbitrary point in time, given an IP/PC, what module/function am I in?

If you have timestamped module load/unload info with base address + range, plus context switch times that allow you to figure out which specific thread & address space was running at any given CPU node ID + point in time, you can always answer that question. (Assuming the debug infrastructure is robust enough to map any given IP to one specific function, which it should be able to do, even if the optimizer has hoisted out cold paths into separate, non-contiguous areas.)

I realize this isn't very helpful to you on Linux (if it's any consolation I'm on Linux these days too), but, sometimes it's interesting to know how other platforms handle it.

trentnelson··on How to build highly-debuggable C++ binaries
Interesting... I've been lamenting the absence of .pdbs on Linux. It sounds like this would allow dissasociating symbol info from the build artifact itself?

(There's no other out-of-the-box solution to this right? i.e. having symbol info live somewhere else other than the .so/exe, that can be loaded on demand when debugging? Like .pdbs basically.)

trentnelson··on How to build highly-debuggable C++ binaries
I like the idea of hacking the crap out of `compile_commands.json` and subverting it for your evil machinations outside of the normal build process. Such a hideously pragmatic tip.
trentnelson··on How to build highly-debuggable C++ binaries
That's neat. The modern equivalent to that these days, on Windows, is to leverage ETW and Windows Performance Analyzer. Potentially with a custom plugin that can visualize your specific perf data as a first-class WPA citizen (i.e. indistinguishable from any other perf data being analyzed, which means you can group/query/filter etc. just like anything else).

I wrote a plugin for a past employer to visualize our internal product event hierarchy performance as if it were a normal C/C++ call stack, it was pretty cool. ETW and WPA are phenomenal tools. I miss them both dearly when on Linux.

trentnelson··on Modifying the OG Xbox to have 256M of RAM [video]
Had fun googling those system names. NX801: https://www.tpc.org/results/individual_results/axil/axil.nx8...

200 9.1GB SCSI disks for 1.8TB!

And still only 4GB RAM on that SQL Server 6.5 box they used for TPC-C. Wild.

And yeah, $770k for that server.

Edit: I guess whilst I'm at it...

Data General AV8600: https://www.tpc.org/results/individual_results/dg/dg.8600.es...

HP NetServer LXr Pro8: https://www.1000bit.it/ad/bro/hp/netserverlxrpor8.pdf

trentnelson··on ExectOS – brand new operating system which derives from NT architecture
I should probably do an updated talk/article/deck on io_uring.

I really do like NT internals though.

trentnelson··on lsix: Like "ls", but for images
What happens when you press page up or down in tmux? Have you configured it to scroll page up & down? Or do you do that via Ctrl-B ] or whatever the magic incantation was?
trentnelson··on RISC vs. CISC by John R. Mashey (1995)
Yeah I was thinking how different those days seemed. And how hard it would be to run into this sort of content these days if you’re a generally curious youngster.
trentnelson··on GCC 14.1
And train control systems.
trentnelson··on C++ Insights – See your source code with the eyes of a compiler
This has made stepping through C++ in gdb slightly less painful for me: use gdb's skip command. E.g., in my ~/.gdbinit:

    # C++ stdlib
    skip -gfi /home/trent/mambaforge/envs/td/x86_64-conda-linux-gnu/include/c++/10.3.0/\*
    skip -gfi /home/trent/mambaforge/envs/td/x86_64-conda-linux-gnu/include/c++/10.3.0/*/*
    skip -gfi /home/trent/mambaforge/envs/td/x86_64-conda-linux-gnu/include/c++/12.3.0/\*
    skip -gfi /home/trent/mambaforge/envs/td/x86_64-conda-linux-gnu/include/c++/12.3.0/*/*

    # tl::expected
    skip -gfi /home/trent/.cache/cpm/expected/5acc53468c550d1f25ce819a675b60bfa0bbc69d/include/tl/\*

It's not perfect, but at least I don't have to step through annoying things like std::unique_ptr<Foo>.get() a million times whilst debugging.
trentnelson··on Scientists find optimal space-time balance for hash tables
What’s its fastest index function look like in assembly? My MultiplyShiftRX clocks in at like 5 cycles on x64 and 3 cycles on my M1. Mine is optimized for offline table generation so construction speed isn’t really relevant for its primary use.
trentnelson··on Scientists find optimal space-time balance for hash tables
Hey, if you're looking for a real-world pragmatic and performant implementation of a theoretically-cool algorithm, my https://github.com/tpn/perfecthash project might fit the bill.

It's geared to generating perfect hash tables with the fastest possible lookup/index times (for 32-bit keys), for key sets in the <=100,000 range. (It scales well up to millions of keys, but the solving time takes a lot longer.)

trentnelson··on Meta AI releases Code Llama 70B
Can that be split across multiple GPUs? i.e. what if I have 4xV100-DGXS-32GBs?
Page 1 of 10Next →