777 karma · joined September 23, 2014
The paper is: Brandis, Marc M., and Hanspeter Mössenböck. "Single-pass generation of static single-assignment form for structured languages." (https://bernsteinbear.com/assets/img/brandis-single-pass.pdf). It was quite understandable to me. For a deeper dive with proper SSA construction with dominance frontiers I could not find time to dig deeper, many other papers on SSA require focused CS work on them, not practically feasible for a side project. Also, single-pass is a requirement for very fast compilation to bytecode and LSP feedback.
I tried to read TS and Pyright source code, they share the same style of immense files and nested local functions, that was quite a steep wall to understand actual inner workings in detail. Maybe TS implementation in Go will be easier to read, it's on my later TODO list. It's tempting to use AI for help, but I'm quite experienced already with undoing AI work when it takes a wrong direction and I do not notice early.
It's too low level, I cannot even tell that I understand everything, only the concept, so I will just share some links. It works only per NAPI epoll context, and one cannot easily control NAPI ID, but if an entire machine is dedicated for a proxy one can do a simple trick of assinging sockets by NAPI ID to dedicated pollers.
In my use case, it was not a proxy, but N socket polling on a machine that then processes received data. It does not look feasible for such case, maybe round-robin polling of NAPI contexts from a single thread may work. What I would really want to have one day from the kernel is that I can easily tell it: trust me, I will poll this single socket eventually, never ever use IRQ path for it.
Previous HN discussion of the kernel feature: https://news.ycombinator.com/item?id=43749271 Nice presentation by the Fastly contributor, with nice diagrams making the big picture much easier to understand: https://netdevconf.info/0x18/docs/netdev-0x18-paper10-talk-s... LWN articles: https://lwn.net/Articles/1008399/, https://lwn.net/Articles/997491/, https://lwn.net/Articles/959462/ Kernel docs: https://docs.kernel.org/networking/napi.html#irq-mitigation
The proposed way just normalizes tracking.
https://en.wikipedia.org/wiki/The_Garden_of_Earthly_Delights...
But my use case is never 24/7, I hibernate it overnight and every time I leave for longer than going to a grocery shop, and I have several Proxmox boxes with proper OSes for hosting stuff. Windows + WSL is my dev/media/web/files/OneDrive machine, a compact silent SFF box that is powerful enough for 90+% of my daily tasks. Lately I try Linux Desktop on Fedora/Ubuntu with every major version, however RDP server and secure boot that I can trust to work and not break myself - these things remain unsatisfactory.
I think the Pro version is enough for reasonable experience, most of the terrible stories originate from the Home version, which should be avoided like the plague.
The only thing I still enjoy is that any data smaller than 1M rows is sliced and diced almost without thinking. I am sometimes really grateful that MS did not break the shortcuts, while almost breaking the product overall. The muscle memory works perfectly.
I'm going to take the metro now and thinking how long do we have until the entire transit network goes down because of a similar incident.
I assume both cases have the file cached in RAM already fully, with a tiny size of 100MB. But the file read based version actually copies the data into a given buffer, which involves cache misses to get data from RAM to L1 for copying. The mmap version just returns the slice and it's discarded immediately, the actual data is not touched at all. Each record is 2 cache lines and with random indices is not prefetched. For the CPU AMD Ryzen 7 9800X3D mentioned in the repo, just reading 100 bytes from RAM to L1 should take ~100 nanos.
The benchmark compares actually getting data vs getting data location. Single digit nanos is the scale of good hash tables lookups with data in CPU caches, not actual IO. For fairness, both should use/touch the data, eg copy it.
After I lost 8 months of photos with a phone ~10 years ago, being sure it was all backed to Google Photos, I would rather trust Microsoft, than risk losing data, and now backup to both clouds. The paid Office+OneDrive is great value.
It just works. Yes, defaults are annoying, but could be changed. I recently enabled a blocked-by-default outgoing firewall, and I have much more questions to JetBrains Rider trying to ignore my system DNS setting and so to bypass Pi-Hole multiple times per minute, than to Microsoft.
internal static class ArrayExtensions
{
[MethodImpl(MethodImplOptions.AggressiveInlining)]
public static ref T RefAtUnsafe<T>(this T[] array, nint index)
{
#if DEBUG
return ref array[index];
#else
Debug.Assert((uint)index < array.Length, "RefAtUnsafe: (uint)index < array.Length");
return ref Unsafe.Add(ref MemoryMarshal.GetArrayDataReference(array), (nuint)index);
#endif
}
}
then your example turns into: public static void AddBatch(int[] a, int[] b, int count)
{
// Storing a reference is often more expensive that re-taking it in a loop, requires benchmarking
for (nint i = 0; i < (uint)count; i++)
a.RefAtUnsafe(i) += b.RefAtUnsafe(i);
}
The JITted assembly: https://sharplab.io/#v2:EYLgxg9gTgpgtADwGwBYA0AXEBDAzgWwB8AB...I'm convinced C# is so much better for high perf code, because yes it can do everything (including easy-to-use x-arch SIMD), but it lets one not bother about things that do not matter and use safe code. It's so pragmatic.
See also the top comments from a recent thread, I totally agree. https://news.ycombinator.com/item?id=45253012
BTW, do not use [MethodImpl(MethodImplOptions.AggressiveOptimization)], it disables TieredPGO, which is a huge thing for latest .NET versions.