Most of the world manages to run their binaries on a mainline kernel so I'd love to know what's so special in their binaries.
Most of the world manages to run their binaries on a mainline kernel so I'd love to know what's so special in their binaries.
Most of the world doesn’t need to worry about their kernel - this is a good thing. But at FAANG scale, you inevitably need to make changes and optimizations.
Eg: “Twitter has a kernel team!?” - https://danluu.com/in-house/
Discussion: https://news.ycombinator.com/item?id=28691676
Not sure if the system call was accepted or if it still exists as Google specific kernel code. I can't seem to find the original article either...
Jump to the 15-minute mark for the three new system calls, switchto_wait, switchto_resume, switchto_switch.
The cgroups code/API that made it into the upstream kernel was the product of a couple of years of internal experimentation at Google into ways to do kernel-level resource isolation, and was pretty different from the approach that was used internally in production on a rather older kernel version. (A bit like Borg vs Kubernetes). And it still took a couple more years for cgroups to replace the original internal mechanisms, since the internal way worked OK and upgrading the kernel across so many machines was risky.
Both of these are completely feasible savings and make custom binaries and a custom kernel completely with it.
https://github.com/abseil/abseil-cpp/blob/1ae9b71c474628d60e...
https://github.com/torvalds/linux/commit/b6a2fea39318e43fee8...
> Those patches implement various internal APIs (e.g. for Google Fibers), provide hardware support, add performance optimizations, and contain other "tweaks that are needed to run binaries that we use at Google".
I'd assume if you're making allowances for custom-made processors at the kernel level you wouldn't also want that leaking into your user land binaries as well.
Also, for some features you need both kernel and userland cooperation. Just think of eg fuse or io_uring or mmap that are in the public kernel. You can surely imagine that Google might brew up similar features.
(I used to work at Google. But I didn't have any special insider knowledge about their kernel stuff. Which is good, so I can't violate any lingering NDAs here with my speculation..)
https://github.com/abseil/abseil-cpp/blob/master/absl/base/i...
They also have for years had major patches to TCP which Linux refuses to adopt. Some details at:
https://www.ietf.org/proceedings/97/slides/slides-97-tcpm-tc...
These won't make binaries "not run" as such but they are necessary for the correct and efficient operation of large-scale distributed systems.
QUIC, Snap/Pony, and user-space thread scheduling are strictly superior to what happens inside the kernel and their development makes the delta between prodkernel and upstream kernel less relevant over time. They also, collectively, make Linux itself less relevant.