2,934 karma · joined March 25, 2009
Computers are cool
$ sudo bpftool prog dump xlated id 50
int restrict_filesystems(unsigned long long \* ctx):
; int BPF_PROG(restrict_filesystems, struct file *file, int ret)
0: (79) r3 = *(u64 *)(r1 +0)
1: (79) r0 = *(u64 *)(r1 +8)
2: (b7) r1 = 0
; uint32_t *value, *magic_map, zero = 0, *is_allow;
3: (63) *(u32 *)(r10 -20) = r1
; int BPF_PROG(restrict_filesystems, struct file \*file, int ret)
4: (bf) r1 = r0
5: (67) r1 <<= 32
6: (77) r1 >>= 32
; if (ret != 0)
7: (55) if r1 != 0x0 goto pc+59
8: (b7) r1 = 32
9: (0f) r3 += r1
10: (bf) r6 = r10
https://github.com/systemd/systemd/blob/c76691d708ac7fe13b7c...Unfortunately, most of the programs loaded by systemd are more-or-less hand-generated (the ingress/egress programs specifically) and do not include this information.
It's a surprisingly small group of folks who work in this space upstream, but I know that they're aware of this as an opportunity to improve things :)
TCP gives you three things: 1. Reasonable performance - This is hard but not impossible to replicate 2. Reliability - This is very hard to replicate because networking edge cases are very hard to isolate 3. Fairness - this one is roughly impossible, because the "fairness" is an artifact of the experimentation and tweaking of Congestion Control Algorithms.
To elaborate on fairness, dynamic traffic control of all flows within a DC while maintaining high utilization is roughly impossible. You can get really close to this by picking your battles wisely (i.e. solid demand control for data warehouse workloads), but you'll always end up counting on individual flows to react appropriately to loss. They need to back off enough to make room for others without tanking their own throughput.
The people who design and implement these algorithms are definitely geniuses, but even they rely on TONS of empirical evidence to narrow parameters to what's appropriate. Of the Kernel Networking people I've worked with, Lawrence Brakmo had the most sophisticated network testing harness I've seen. Even then, you don't really know if it works (and can't finish tuning it) until you run it in production.
Running novel congestion control algorithms in production at a sufficient scale to figure out whether or not they're working appropriately is a great way to kill your network, so we end up conducting the equivalent of CCA drug testing to roll it out slowly and safely.
The end result of all of this is that it's really hard to solve the "arbitrary connections sharing arbitrary network topologies with high utilization" problem quickly enough for it ever to look like a breakthrough rather than just steady progress.
It's also worth noting that it's usually easiest to prove performance, so you'll see a lot of excitement about performance benchmarks from people who don't yet know what they're about to learn about networking. We were very much in this camp at Facebook when we were all-in on memcache-over-udp, and we later abandoned it completely.
This is ridiculous.
Hyperscalars see an immediate ROI from efficiency/reliability improvements and actively invest in TCP alternatives all of the time. It's just really hard.
Networking companies see an ability to differentiate their products from their peers and work on this kind of thing as well. I did a 3 second google for "QUIC acceleration Mellanox" and got a hit on Nvidia's blog right away.
You just can't trivially replace something with an investment totally 50 years of clock time and thousands of years of engineer time. It will either take a long time or a massive shift in needs/technology. FWIW, I wouldn't be surprised if the high-performance RDMA networks being put together for AI workloads were the thing that grew into the "next" thing.
As for the risk of selling something with proprietary scheduler bits, the upshot is that this is well handled, because most of the interfaces are GPL-only. IANAL, but this [1] is probably an interesting read.
[0] btw, more often than not the reason that these things don't go upstream is because upstream says "no." Even if such changes are worth a lot of money to the business, it's not enough that the competitive advantage outweighs the maintenance cost of a forked scheduler
[1] https://lore.kernel.org/netdev/20210916032104.35822-1-alexei...
Main Memory => zswap (compressed memory) => swap
In this case, the pages may be logically allocated or not -- the assurance is that the data will be the value you expect it to be when it becomes resident.
Should those pages be uninitialized, the "Swapped" state is really just "Remember that this thing was all zeros."
We could do computing your way, but it'd be phenomenally more expensive. I know this because every thing we introduce to the hierarchy in practice makes computing phenomenally less expensive.
my takeaway was essentially that you get sweet perf wins from semantics that are hard to replicate with a type system that's also making really strong guarantees without making the code SUPER gross.
I manage a linux team so this battle was lost for me long ago, but having managed a site-oriented team before, it was a lot easier to build relationships for me and the team members.
Yes, it’s hard to do research if you start at “let’s design a kernel from scratch,” but you’d never need to do that. You can just hack Linux or a bsd, or even use something like bpf to extend it.
The thing that annoyed me about 410 is that when I switched to Linux I realized that pusha/popa didn’t matter at all. It was a good course for writing reentrant C and learning the very basics of hardware, but the really hard stuff is in the weird dynamics of memory management on NUMA systems and work conserving io, which you can’t get anywhere near if you are starting from scratch.
I think we're going to continue to see more of this shared-memory message passing style for two reasons: 1. Hardware performance (NICs, SSDs, etc) is out accelerating CPU performance. You can't keep up if you're hitting a ton of context switches 2. Context switches are crazy expensive, and (at least temporarily) getting more expensive due to Spectre/Meltdown mitigations.
These considerations pique my interest more so than micro/exo/whatever kernel architecture considerations do.
In terms of stack openness, I think the biggest changes are what's happening with bpf. While it's always been possible to go hack the scheduler however you want, it's not really been feasible -- you're likely to break the thing, and carrying patches around is a giant pain. With bpf hooks, you can manipulate kernel behaviors in a very fine-grained fashion, which has already created a bunch of academic interest and really changed the way we build our low-level systems software (containers, networking, etc).
The biggest thing here is that visibility into internals and strong ABI guarantees are inherently at conflict, and it'll take some time to figure out the more nuanced view.
The interesting problem is that it’s still kind of a developer experience mess. We can get pretty far with libbpf skeleton, libbpf-rs, etc, but I think we’re still waiting on the “killer” framework for this (or some other kind of language support).
[1] I work with Pavel, Jens, and Alexei
I “weaseled” my way into CMU via athletic admissions (I was an actual athlete, not a Lori Laughlin style one), but did very well at CMU once I got there. People who aced the SATs did not do as well. Fwiw I still did okay, 32 ACT score, but there were 35/36’s around.
IOW, prediction of academic success is hard; career success harder. These standardized tests don’t add much.
Managers who only ever “interface” or act as a “shit umbrella” end up screwing the team over in the long run.
Your surgeon metaphor only stands up in a situation where 100+ Surgeons are operating on the same person.
Second, even if you do the first, the whole thing gets screwed up again when you start trying to introduce vendor code into the mix. Generally, no one upstream gives a crap that you have super compelling business reasons to compromise on code quality (or even trivial things like how code is committed: tarballs vs good git hygiene), and vendors sometimes compromise a lot.
So it’s not surprising that sometimes groups choose to do the expedient thing to get something to market instead of doing things “the right way.” In a lot of respects, the original Android did this with Linux.
Competition is good.
[0] https://github.com/facebook/folly/blob/16d6394130b0961f6d688...
That said, looking at the (apparently) leading implementation, capsicum
> Capsicum also introduces capability mode, which disables (with ECAPMODE) all syscalls that access any kind of global namespace; this is mostly (but not completely) implemented in userspace as a seccomp-bpf filter.
So I do feel that bpf ultimately enables building the kinds of abstractions that people want.
For the last few years, I managed the Container Runtime group at Facebook. My experience has been:
1. `if (has_capability(..., X)) { ... }` gets put into code pretty haphazardly in a way that's not necessarily super well structured. Once it's there, it's ABI, and you're screwed if you want to iterate on it. That's why cap_sys_admin is /almost/ root.
2. If you wanted to do the right thing from the jump (e.g. for bpf itself), you'd have to add a new capability. This is a heavy lift for something that might not actually get any traction. It requires changing a bunch of common tools, and you likely end up breaking a bunch of applications.
3. Debugging capability failures is a pain in the ass. We ended up building and deploying capability tracing infrastructure just to figure out what people are actually using.
4. For gradual roll outs of enforcement/changes, you need the flexibility to warn first, enforce second. We did large scale monitoring of all such changes to make sure we didn't break the workloads.
5. Even if you nail all of the above, the ability to make finer-than-capability-grained decisions (i.e. binding to port 20 or 80 is okay but not port 22) is really valuable.
I'm all for kernel abstractions that just work and solve all problems for all people, but I think the overwhelming trend has been towards kernel interfaces that provide a lot of flexibility and then more opinionated libraries/tools that kind of let us have our cake and eat it to (io_uring => liburing, bpf => libbpf, btrfs => btrfstools).
You also point to a lot of problems that are container-independent and lay them at the feet of docker, which is unfair.
Upgrading the OS is always hard unless you have some awesome, declarative config and you managed to depend on zero of the features that have changed. It doesn't matter if you're in a container or not, switching from iptables in Centos 7 to nftables in Centos 8 is going to introduce some pain.
And somehow we get mad at people for not knowing how to install things, but the complexity of installing them is itself a problem. More steps means more inconsistency, which means it's more likely that "it works on my machine, but breaks on yours."