HNHacker News
TopNewBestAskShowJobs

alexgartrell

2,934 karma · joined March 25, 2009

SWE @ Thinking Machines Previously Operating Systems, etc @ Meta

Computers are cool

submissionscomments
alexgartrell··on CMU CS Academy: a free online computer science curriculum by Carnegie Mellon
I think people dramatically overestimate the value of CMU social connections. The value of a CMU Comp Sci degree is that you've proven (to yourself and everyone else) that you can get one. The program is essentially one "weeder" course after another to identify the people who don't belong (21-127 Concepts of Mathematics, 15-251 Great Theoretical Ideas of Computer Science, 15-213 Introduction to Compute Systems, 15-410 Operating Systems, 15-451 Design and Analysis of Algorithms). Every course is the "hard" version of ever other course (e.g. "Welcome to Operating Systems, you'll be writing a kernel that runs on an X86_64 processor. Your hint is 'pusha/popa.' Go."). While it was horrifying the whole way through, I think that it built a resiliency in me that wouldn't exist otherwise. It's incredibly hard to replicate that through online classes.
alexgartrell··on Google exec fired after female boss groped him at drunken bash, suit says
There are reasons beyond compensation to like a job (fulfillment, relationships, status, etc). If you want to do Google-scale anything, at best you've got 4 other companies to choose from. So you're going to feel hurt if you lose that opportunity for reasons you feel are beyond your control. Especially if you've put 16 years into it and you feel the loyalty is due back.
alexgartrell··on eBPF – Adding functionality to OS at runtime
FWIW, programs compiled through a modern toolchain can ship their own debug data. For example, the restrict_filesystems program loaded by systemd

    $ sudo bpftool prog dump xlated id 50
    int restrict_filesystems(unsigned long long \* ctx):
    ; int BPF_PROG(restrict_filesystems, struct file *file, int ret)
       0: (79) r3 = *(u64 *)(r1 +0)  
       1: (79) r0 = *(u64 *)(r1 +8)
       2: (b7) r1 = 0 
    ; uint32_t *value, *magic_map, zero = 0, *is_allow;
       3: (63) *(u32 *)(r10 -20) = r1
    ; int BPF_PROG(restrict_filesystems, struct file \*file, int ret)
       4: (bf) r1 = r0                                                             
       5: (67) r1 <<= 32             
       6: (77) r1 >>= 32
    ; if (ret != 0)           
       7: (55) if r1 != 0x0 goto pc+59                                             
       8: (b7) r1 = 32               
       9: (0f) r3 += r1
      10: (bf) r6 = r10        
https://github.com/systemd/systemd/blob/c76691d708ac7fe13b7c...

Unfortunately, most of the programs loaded by systemd are more-or-less hand-generated (the ingress/egress programs specifically) and do not include this information.

It's a surprisingly small group of folks who work in this space upstream, but I know that they're aware of this as an opportunity to improve things :)

alexgartrell··on eBPF – Adding functionality to OS at runtime
eBPF is also a lot more efficient than ptrace
alexgartrell··on We need a replacement for TCP in the datacenter [pdf]
In this case we're talking about within the Datacenter, and you could conceivably update every network device and system to talk the new thing if you wanted. This is more achievable at a hyperscalar, where there tends to be < 3 distinct protocols, proxies, etc.

TCP gives you three things: 1. Reasonable performance - This is hard but not impossible to replicate 2. Reliability - This is very hard to replicate because networking edge cases are very hard to isolate 3. Fairness - this one is roughly impossible, because the "fairness" is an artifact of the experimentation and tweaking of Congestion Control Algorithms.

To elaborate on fairness, dynamic traffic control of all flows within a DC while maintaining high utilization is roughly impossible. You can get really close to this by picking your battles wisely (i.e. solid demand control for data warehouse workloads), but you'll always end up counting on individual flows to react appropriately to loss. They need to back off enough to make room for others without tanking their own throughput.

The people who design and implement these algorithms are definitely geniuses, but even they rely on TONS of empirical evidence to narrow parameters to what's appropriate. Of the Kernel Networking people I've worked with, Lawrence Brakmo had the most sophisticated network testing harness I've seen. Even then, you don't really know if it works (and can't finish tuning it) until you run it in production.

Running novel congestion control algorithms in production at a sufficient scale to figure out whether or not they're working appropriately is a great way to kill your network, so we end up conducting the equivalent of CCA drug testing to roll it out slowly and safely.

The end result of all of this is that it's really hard to solve the "arbitrary connections sharing arbitrary network topologies with high utilization" problem quickly enough for it ever to look like a breakthrough rather than just steady progress.

It's also worth noting that it's usually easiest to prove performance, so you'll see a lot of excitement about performance benchmarks from people who don't yet know what they're about to learn about networking. We were very much in this camp at Facebook when we were all-in on memcache-over-udp, and we later abandoned it completely.

alexgartrell··on We need a replacement for TCP in the datacenter [pdf]
> corporations don't see any "immediate shareholder value", so they sit around happy as pigs in shit with the status quo.

This is ridiculous.

Hyperscalars see an immediate ROI from efficiency/reliability improvements and actively invest in TCP alternatives all of the time. It's just really hard.

Networking companies see an ability to differentiate their products from their peers and work on this kind of thing as well. I did a 3 second google for "QUIC acceleration Mellanox" and got a hit on Nvidia's blog right away.

You just can't trivially replace something with an investment totally 50 years of clock time and thousands of years of engineer time. It will either take a long time or a massive shift in needs/technology. FWIW, I wouldn't be surprised if the high-performance RDMA networks being put together for AI workloads were the thing that grew into the "next" thing.

alexgartrell··on GhOSt: Fast and Flexible User-Space Delegation of Linux Scheduling
Well you can already do whatever proprietary stuff you want with your own kernels if you're not putting them on devices that you sell to people. In practice, this means that Amazon, Google, etc likely have pretty fancy, totally unshared scheduler changes that are worth a significant amount of money. [0]

As for the risk of selling something with proprietary scheduler bits, the upshot is that this is well handled, because most of the interfaces are GPL-only. IANAL, but this [1] is probably an interesting read.

[0] btw, more often than not the reason that these things don't go upstream is because upstream says "no." Even if such changes are worth a lot of money to the business, it's not enough that the competitive advantage outweighs the maintenance cost of a forked scheduler

[1] https://lore.kernel.org/netdev/20210916032104.35822-1-alexei...

alexgartrell··on In C, how do you know if the dynamic allocation succeeded?
Allocated-but-unavailable is a totally reasonable part of the memory hierarchy.

Main Memory => zswap (compressed memory) => swap

In this case, the pages may be logically allocated or not -- the assurance is that the data will be the value you expect it to be when it becomes resident.

Should those pages be uninitialized, the "Swapped" state is really just "Remember that this thing was all zeros."

We could do computing your way, but it'd be phenomenally more expensive. I know this because every thing we introduce to the hierarchy in practice makes computing phenomenally less expensive.

alexgartrell··on Simple Linux kernel memory corruption bug can lead to complete system compromise
I shared this perspective, but luckily my job is awesome and (in a routine 1:1!) Paul told me why it's less straightforward than I thought: https://paulmck.livejournal.com/62436.html

my takeaway was essentially that you get sweet perf wins from semantics that are hard to replicate with a type system that's also making really strong guarantees without making the code SUPER gross.

alexgartrell··on The great executive-employee disconnect
Don't you think it's a lot more plausible that extroversion correlates with becoming a manager in the same way that it correlates with wanting to see people in the office? Or that it's simply easier to do the manager job in person than over the internet?

I manage a linux team so this battle was lost for me long ago, but having managed a site-oriented team before, it was a lot easier to build relationships for me and the team members.

alexgartrell··on It's Time for Operating Systems to Rediscover Hardware
Admittedly, I got a B in 410 and only managed to TA 213, but I do work in the space now.

Yes, it’s hard to do research if you start at “let’s design a kernel from scratch,” but you’d never need to do that. You can just hack Linux or a bsd, or even use something like bpf to extend it.

The thing that annoyed me about 410 is that when I switched to Linux I realized that pusha/popa didn’t matter at all. It was a good course for writing reentrant C and learning the very basics of hardware, but the really hard stuff is in the weird dynamics of memory management on NUMA systems and work conserving io, which you can’t get anywhere near if you are starting from scratch.

alexgartrell··on Europe's all-time heat record set in Sicily at nearly 120 degrees
I don't think it's particularly ambiguous since anything other than Fahrenheit would imply that everyone is dead.
alexgartrell··on Descriptorless Files for Io_uring
Big disclaimer: I support these guys but I don't tell them what to do, so these are just my dumb, manager opinions :)

I think we're going to continue to see more of this shared-memory message passing style for two reasons: 1. Hardware performance (NICs, SSDs, etc) is out accelerating CPU performance. You can't keep up if you're hitting a ton of context switches 2. Context switches are crazy expensive, and (at least temporarily) getting more expensive due to Spectre/Meltdown mitigations.

These considerations pique my interest more so than micro/exo/whatever kernel architecture considerations do.

In terms of stack openness, I think the biggest changes are what's happening with bpf. While it's always been possible to go hack the scheduler however you want, it's not really been feasible -- you're likely to break the thing, and carrying patches around is a giant pain. With bpf hooks, you can manipulate kernel behaviors in a very fine-grained fashion, which has already created a bunch of academic interest and really changed the way we build our low-level systems software (containers, networking, etc).

The biggest thing here is that visibility into internals and strong ABI guarantees are inherently at conflict, and it'll take some time to figure out the more nuanced view.

alexgartrell··on Descriptorless Files for Io_uring
We’re working on it :) [1]

The interesting problem is that it’s still kind of a developer experience mess. We can get pretty far with libbpf skeleton, libbpf-rs, etc, but I think we’re still waiting on the “killer” framework for this (or some other kind of language support).

[1] I work with Pavel, Jens, and Alexei

alexgartrell··on Give me /events, not webhooks
I think the Stripe API stuff you did was fine, but you really did your best work as a concepts of mathematics TA.
alexgartrell··on MIT suspending SAT/ACT requirement for next application cycle
The fundamental issue is that crushing the SAT/ACT is more of a reflection of “mom and dad got me good tutoring or prep” than it is intellectual merit.

I “weaseled” my way into CMU via athletic admissions (I was an actual athlete, not a Lori Laughlin style one), but did very well at CMU once I got there. People who aced the SATs did not do as well. Fwiw I still did okay, 32 ACT score, but there were 35/36’s around.

IOW, prediction of academic success is hard; career success harder. These standardized tests don’t add much.

alexgartrell··on Oil companies buying up EV charging networks: Shell acquires Ubitricity
I'm not an electrical engineer, but from watching this video [0], there seem to be about a million and a half different mechanisms to prevent sparks from chargers.

[0] https://www.youtube.com/watch?v=RMxB7zA-e4Y

alexgartrell··on Leading someone with more years of experience than my age
I think the reality is that a manager’s job is alignment. Sometimes that is resetting the expectations of the organization, and sometimes that is redirecting the team or team members. Realigning without pissing people off is what leadership is.

Managers who only ever “interface” or act as a “shit umbrella” end up screwing the team over in the long run.

Your surgeon metaphor only stands up in a situation where 100+ Surgeons are operating on the same person.

alexgartrell··on TerarkDB, ByteDance's RocksDB replacement
This already happens, because most (?) people aren’t running vanilla kernels. Many (most?) distros compile their kernels with config options and patches that “make sense to them.” In the most egregious cases, you end up with things like bpf being intentionally broken by default.
alexgartrell··on TerarkDB, ByteDance's RocksDB replacement
First, it’s reeaallllyyyy expensive to invest enough in an open source project that you have a reasonable chance of steering it.

Second, even if you do the first, the whole thing gets screwed up again when you start trying to introduce vendor code into the mix. Generally, no one upstream gives a crap that you have super compelling business reasons to compromise on code quality (or even trivial things like how code is committed: tarballs vs good git hygiene), and vendors sometimes compromise a lot.

So it’s not surprising that sometimes groups choose to do the expedient thing to get something to market instead of doing things “the right way.” In a lot of respects, the original Android did this with Linux.

Competition is good.

alexgartrell··on U.S. Treasury breached by hackers backed by foreign government – sources
I am a huge open source fanboy, but there's nothing magical about open source that makes it secure against nation states.
alexgartrell··on About 150 U.S. Cadillac Dealers to Exit Brand, Rather Than Sell Electric Cars
It’s common for dealerships to hold several franchises too. So they wouldn’t even necessarily be struggling, just reacting to Cadillac performing worse than Toyota, for example
alexgartrell··on How io_uring and eBPF Will Revolutionize Programming in Linux
I wouldn't be doing my job if I failed to mention that both Alexei (eBPF) and Jens (io_uring, block) work at Facebook. Beyond them, we've got a bunch of folks working on the primitives as well as low-level userspace libraries [0] that enable us to use all of this stuff in production, so, by the time you're seeing it, we've demonstrated that it works well for all of Facebook's load balancers, container systems, etc.

[0] https://github.com/facebook/folly/blob/16d6394130b0961f6d688...

alexgartrell··on eBPF – The Future of Networking and Security
I was not aware of object capabilities -- TIL.

That said, looking at the (apparently) leading implementation, capsicum

> Capsicum also introduces capability mode, which disables (with ECAPMODE) all syscalls that access any kind of global namespace; this is mostly (but not completely) implemented in userspace as a seccomp-bpf filter.

So I do feel that bpf ultimately enables building the kinds of abstractions that people want.

alexgartrell··on eBPF – The Future of Networking and Security
Can you provide more context on why you feel that's true (or even possible)?

For the last few years, I managed the Container Runtime group at Facebook. My experience has been:

1. `if (has_capability(..., X)) { ... }` gets put into code pretty haphazardly in a way that's not necessarily super well structured. Once it's there, it's ABI, and you're screwed if you want to iterate on it. That's why cap_sys_admin is /almost/ root.

2. If you wanted to do the right thing from the jump (e.g. for bpf itself), you'd have to add a new capability. This is a heavy lift for something that might not actually get any traction. It requires changing a bunch of common tools, and you likely end up breaking a bunch of applications.

3. Debugging capability failures is a pain in the ass. We ended up building and deploying capability tracing infrastructure just to figure out what people are actually using.

4. For gradual roll outs of enforcement/changes, you need the flexibility to warn first, enforce second. We did large scale monitoring of all such changes to make sure we didn't break the workloads.

5. Even if you nail all of the above, the ability to make finer-than-capability-grained decisions (i.e. binding to port 20 or 80 is okay but not port 22) is really valuable.

I'm all for kernel abstractions that just work and solve all problems for all people, but I think the overwhelming trend has been towards kernel interfaces that provide a lot of flexibility and then more opinionated libraries/tools that kind of let us have our cake and eat it to (io_uring => liburing, bpf => libbpf, btrfs => btrfstools).

alexgartrell··on Leaders focus on doing the right thing, managers focus on doing things right
It’s less about having enough information and more about having leaders who have established reasonable constraints and incentives. Organizations get into trouble when doing the right thing runs counter to (intentional or accidental) incentive structures. The point of the tweet is that leaders find a way to do the right thing anyway.
alexgartrell··on Foundations of Software Engineering
I took and TAed this class. I ended up dipping my toes into the open source waters with some Firefox contributions and completing the rest of the minor program including a set of contributions to Chromium. One thing that was heavily emphasized in that curriculum was how open source can be a huge asset to companies. I really took this to heart, and it has played a significant role in shaping open source strategy for containers and Linux stuff at Facebook.
alexgartrell··on Unlocking eBPF power
I work with some of the maintainers of bpf, bpftrace, etc. What can we do to make this stuff more accessible?
alexgartrell··on Programming Inside a Container
Maybe you're talking about non-native containers (i.e. not Linux), but there's no technical merit to the idea that a container by itself could introduce 15x latency on a Linux host for something like a web request, unless something like network namespaces, tc, etc was being used very improperly.

You also point to a lot of problems that are container-independent and lay them at the feet of docker, which is unfair.

Upgrading the OS is always hard unless you have some awesome, declarative config and you managed to depend on zero of the features that have changed. It doesn't matter if you're in a container or not, switching from iptables in Centos 7 to nftables in Centos 8 is going to introduce some pain.

And somehow we get mad at people for not knowing how to install things, but the complexity of installing them is itself a problem. More steps means more inconsistency, which means it's more likely that "it works on my machine, but breaks on yours."

alexgartrell··on Five Years of Btrfs
Yeah there are a remarkable set of container runtime tasks (package downloads, rootfs creation and management, etc) that are way easier with btrfs. It wasn’t always smooth sailing but luckily Chris, Josef, Omar and others are awesome and now (and for the last while) we are asking for features rather than fixes.
← PreviousPage 2 of 18Next →