Moving Google toward the mainline
lwn.net
lwn.net
Since we've moved to tracking -current, we rebase every few weeks. We catch any newly introduced bugs from upstream almost immediately, since there are only a few weeks of changes, not a few years, to sort through. Its far easier to contribute code upstream. In fact, we do most development upstream and wait until the next rebase to pick our changes up (or cherry pick them if it is something urgent).
All in all, life is much, much, much better on the bleeding edge.
I have one service that is low traffic that I use to test out new base images and core api changes. It’s easy to generate synthetic traffic to see how it behaves, as a lot of its traffic comes from the build process of more complex services.
Same feeling here, but on a personal level (meaning, using a desktop computer for professional work). Started using Ubuntu many years ago, ran updates maybe once every month. Every now and then something breaks, but always unsure of what/how/when, so digging through things that changed always took long time.
Started using Arch Linux and few years ago and also started doing updates at least once per day. Now it's way easier to discover what particular change broke something, as the releases are way smaller.
I attribute this more to how often I do the updates now VS before, and less to which OS you use. Same can be achieved with Debian Testing et al.
Rebasing a Linux Kernel Branch that’s Two Years Old and 9k Commits Deep
…but you would have had thousands of us running to the bathroom to revisit our last meal.
What a Herculean effort by these two engineers. When your role is to track down fellow employees from completely different teams and crack the, er, carrot to get them to update their patch from two years ago, you can get very tired very quickly.
I’m sure the hardware driver modules are fairly self contained but things like the userspace QoS patch could be pretty hairy. The kind of thing where, for upstream, you’d spend 90% of your time landing a framework just to even be able to implement your feature with the final 10% patch.
Oof. Bravo.
A possible better editorialized title would be
Creating a parallel kernel so we can continually rebase commits instead of every two years
Herculean effort indeed.
I feel this was a fair description of CGroups.
I hate to be a pessimist, and I may be simply ignorant of some realities, but it FEELS like this is setting an impossibly high goal. I think there will turn out to be something significant in 9000 variances, which cannot simply be uplifted into mainstream, and which turn out to be hard to maintain as a continual rebase into the emerging mainstream.
I agree that 40 sounds low, but probably is a good indication it can work. 1% is within view, if they maintained pace so even on a 1% if the rebase cost can be handled, they can converge but that assumes there isn't a showbuster in the 9000.
There is a system for using git* in conjunction with the monorepo that some developers use, but that system has nothing to do with the prodkernel.
* there is also a similar system for using mecurial with the monorepo that some folks love.
The monorepo (google3) and its review tools were used for most other projects.
google1 (I can't remember the actual directory name, maybe //google/) used Makefiles and Makefile generators. IIRC there is a small amount of still-live code here.
google2 failed fairly quickly.
google3 is blaze/bazel and is where almost all development is today.
All three trees are actually in the same source control "depot".
If you read enough of their papers this is actually detailed there but spread across different papers.
Considering that the kernel OOM killer tends to be way too late in doing its thing I don't see how this is inelegant, maybe there's a reason you can't just have the kernel kill processes earlier in the face of memory pressure.
What usually happens is that in near-OOM conditions, the kernel starts reclaiming memory pages backed by file (sometimes called "trashing"). This operation manages to keep some extra memory available, but it makes the system almost unresponsive because it's constantly copying memory back and forth from the disk. It may take anywhere between minutes to hours before the system finally OOMs and the OOM killer is invoked.
This problem has been there forever but has been made worse by the improved speeds of modern storage technologies: with slower disk I/O, the OOM condition was reached sooner.
There are several solutions:
- Buy more RAM: if your system routinely goes nearly OOM something is not right.
- Add a (small) swap. It doesn't have to be a partition: nowadays most filesystems support swap file. Just create an empty file and mark it as swap.
- Limit the amount of thrashing or protect some pages from being reclaimed. This has been proposed by Google first and several other people since then, but AFAIK it has never been implemented in the mainline kernel.
Regarding the latter solution, there is a patchset called le9-patch[1] that is included in some alternative Linux kernels and it should be relatively safe to use.
All hardware being the same (RAM, SSD, CPU) as OOM is reached, Linux will freeze, whereas Windows continues to run smoothly. All OSes try to reclaim memory pages, just Linux seems to hang the user space while doing so.
As someone who has dual-booted Windows and Linux for a decade, I can 100% attest to this glaring problem.
This could be fixed. Android doesn't behave like that, and it's not too far from the mainline kernel nowadays.
I'm not a kernel developer or anything like that, I've just spent some time investigating why this issue happens and has been happening for more than 10 years now.
> the user experiences the PC freeze during certain conditions, however you choose to name them
I'm not trying to defent the Linux kernel, I just described how it works. In particular it's not true that the OOM killer "takes too long" or doesn't work: it's just not invoked at all. If you invoke it manually (enable the magic SysRq with`sysctl kernel.sysrq=1` and press `alt-sysrq-f`) it does its job and solves the OOM instantly.
So, if you don't want to deal lockups and don't like an OOM userspace daemon (I don't), these are the possible solutions.
> I would take one crashed program over power cycling the entire PC any day of the week.
On a laptop or desktop PC, you don't need to power cycle in a near-OOM: use the magic SysRq key.
Thanks for the tip! If my Linux were ever to start locking up regularly, I will apply it.
But right now (so I don't have to give up the use to which I currently put my SysRq key) I would prefer some method for determining after I forcefully powered down the computer, then powered up again, whether the lockup or slow-down that motivated the force-power-down was caused by a near-OOM condition.
Do you happen to have a tip for that?
My experiences with swap on Linux have been similarly bad. If even brief memory pressure forces the kernel to move things to swap, the only way to revert that in any reasonable timeframe is to unmount the swap partition or to restart the machine.
Meanwhile using Windows with a swap file of twice the size of physical RAM runs smooth as butter. I have a 200GB swap file right now and no problems.
I've long wondered this too. How does Windows handle memory pressure differently?
It does thrash much more gracefully than Linux, though. In fact the "your computer is low on memory" prompt actually can show up even when severely thrashing, something utmost impossible in Linux (even starting something like zenity may take hours..).
That's what resource control via cgroups is about. Fedora desktop folks (both GNOME and KDE) are working on ensuring minimum resources are available for the desktop experience, via cgroups, which then applies CPU, memory, and IO isolation when needed to achieve that. Also, systemd-oomd is enabled by default. The resource control picture isn't completely in place yet, but things are much improved.
Putting desktop apps into individual cgroups is one of the more counter-productive ideas that has cropped up lately.
> ... allows maximizing hardware utilization without sacrificing workload health or risking major disruptions such as OOM kills.
1. Compiling the kernel with custom knobs 2. Write the custom code as Linux modules
It is when you get into kernel internals that rebasing continuously becomes a challenge. Some parts of the core Linux kernel don't change often, but I imagine many other parts see significant churn.
A little more philosophicallh, if all these customization points were available as modules, the process of updating modules to work with new versions of the kernel would be exactly as much of a mess.
As for the scheduler stuff, the main change is a SwitchTo set of syscalls that allow threads to bypass the kernel’s scheduler and just continue execution as a different thread. https://lkml.org/lkml/2020/7/22/1202
https://www.youtube.com/watch?v=KXuZi9aeGTw They explain it around 15:01 - google added it's own syscall switchto_thread - that puts the current thread to sleep and is switching to the argument thread id. (and some other calls too). That one helps with cutting down latency in inter thread calls for m:n threading. The real effort is to make latency for individual application requests predictable, while keeping it low.
There seems to be a Phoronix article with some more info and a link to a preparatory patchset that's public: https://www.phoronix.com/scan.php?page=news_item&px=Google-F...
Google borrowed for proprietary reasons; meanwhile the people who maintain and improve the kernel and give away the fruits for free have produced something that google wants so they'll merge, again for FAANG proprietary reasons. I'm not losing sleep over it, hell, I paid for it already with my privacy.
Maybe it was so in the past, but I doubt that any significant contributions nowadays come as a free (as in beer) work.
https://lwn.net/Articles/867540/
Any source for this claim?
Yes, and they say exactly that. What's your point?
Mark was a senior front end Dev who was bad at complex logic. He sat near me and I took pity on him, helped him out to help us out.
Steve was smarter and more experienced but had attention to detail problems and did not understand git. He should have been one of my lieutenants but I could not trust him and often had to help him do branch surgery.
Steve did not like this arrangement but it couldn’t be helped. He got cranky about it and went after Mark. I can’t prove he was coming at me sideways, but it was curious.
One day Steve claims Mark broke our shit. Here’s this line with his name on it. It even looked like a Markism. But the thing was, I reviewed that code. I remember being happy Mark got it right on the first try. This was not the code I merged but git says it is. Da fuq?
So I start bisecting and looking at branches and sure enough, Steve screwed a three way merge, again, and blame treated the change as if Mark wrote it.
Thanks, you stupid git.
If you have a patch that saves significant resources, improves important performance metrics or unblocks hardware that does those things you don't care if upstream will take the patches or wait for them to decide if they will. You simply start using the patches ASAP and reap the benefits, then you continue trying to upstream the patches to reduce the maintenance burden.
I wonder how some of these tools will change due to progress on io_uring and similar idioms. Batching system calls to amortize the cost.
To be honest I don't think it is a major loss. IIUC most features are either only interesting internally, or are submitted upstream to reduce maintenance cost. Google's "secret sauce" isn't how great their kernel is, so I am not aware of anything that is held back for competitive advantage other than maybe some drivers for custom hardware. So while it would be nice to see all of these patches collected into a "live" snapshot of their kernel it probably wouldn't be that much more helpful (if at all) to a third-party.
Most of the world manages to run their binaries on a mainline kernel so I'd love to know what's so special in their binaries.
> Those patches implement various internal APIs (e.g. for Google Fibers), provide hardware support, add performance optimizations, and contain other "tweaks that are needed to run binaries that we use at Google".
I'd assume if you're making allowances for custom-made processors at the kernel level you wouldn't also want that leaking into your user land binaries as well.
Also, for some features you need both kernel and userland cooperation. Just think of eg fuse or io_uring or mmap that are in the public kernel. You can surely imagine that Google might brew up similar features.
(I used to work at Google. But I didn't have any special insider knowledge about their kernel stuff. Which is good, so I can't violate any lingering NDAs here with my speculation..)
https://github.com/abseil/abseil-cpp/blob/master/absl/base/i...
They also have for years had major patches to TCP which Linux refuses to adopt. Some details at:
https://www.ietf.org/proceedings/97/slides/slides-97-tcpm-tc...
These won't make binaries "not run" as such but they are necessary for the correct and efficient operation of large-scale distributed systems.
QUIC, Snap/Pony, and user-space thread scheduling are strictly superior to what happens inside the kernel and their development makes the delta between prodkernel and upstream kernel less relevant over time. They also, collectively, make Linux itself less relevant.
Most of the world doesn’t need to worry about their kernel - this is a good thing. But at FAANG scale, you inevitably need to make changes and optimizations.
Eg: “Twitter has a kernel team!?” - https://danluu.com/in-house/
Discussion: https://news.ycombinator.com/item?id=28691676
Not sure if the system call was accepted or if it still exists as Google specific kernel code. I can't seem to find the original article either...
Jump to the 15-minute mark for the three new system calls, switchto_wait, switchto_resume, switchto_switch.
https://github.com/abseil/abseil-cpp/blob/1ae9b71c474628d60e...
https://github.com/torvalds/linux/commit/b6a2fea39318e43fee8...
The cgroups code/API that made it into the upstream kernel was the product of a couple of years of internal experimentation at Google into ways to do kernel-level resource isolation, and was pretty different from the approach that was used internally in production on a rather older kernel version. (A bit like Borg vs Kubernetes). And it still took a couple more years for cgroups to replace the original internal mechanisms, since the internal way worked OK and upgrading the kernel across so many machines was risky.
Both of these are completely feasible savings and make custom binaries and a custom kernel completely with it.