HNHacker News
TopNewBestAskShowJobs

jsolson

2,421 karma · joined February 21, 2007

Google principal engineer based out of Seattle. I am the technical lead for delivering the next generations of AI GPU Supercomputers to Google Cloud. Previously, I worked on the hypervisor that powers Google Compute Engine.

jonolson at google dot com

Opinions are my own, not those of my employer.

submissionscomments
jsolson··on Why SQLite Does Not Use Git
How do you build up the incremental changes? By staging them into your working directory from the one giant commit?
jsolson··on Tesla is heading for a cash crunch
It literally bought me a boat this week.

I picked up some at the money May calls Monday. I sold them today, and (mostly coincidentally) bought a sailboat. The options, after taxes, covered the boat and this year's moorage.

Could've gone the other way of course, but after the beating they'd taken? Surely they'd do something to correct between Monday and earnings.

They said, I might amend your statement to be emotional volatility is your friend. I mostly got lucky on the response to Tesla coming closer to meeting their production goals (entirely independently of the underlying financials). It elicited a predictable emotional response.

jsolson··on Stephen Hawking has died
Yes, but given today's technology (and my lack of faith in a suitable deity), I do not have confidence in continuity after corporeal death.
jsolson··on Stephen Hawking has died
Interesting.

I would like to attribute my desire to live "forever", at least to some extent, to cosmic perspective.

One day the universe might make me bored, but today is not that day, and I suspect tomorrow isn't either.

I don't ever expect that day to come, but if I thought it was near I'd be very afraid indeed.

jsolson··on Mountain View approves 10,000 homes by Google's North Bayshore project (2017)
If by mixed do you just mean residential over retail?

Is the Bay Area actually opposed to that?

jsolson··on SpaceX Rocket Survives Experimental High-Thrust Landing at Sea
> And once you're throwing it out, might as well do some experiments.

I understand what you're saying here, but the extent to which the aerospace industry just does not work this way is staggering.

This is a fairly unique quality of SpaceX.

jsolson··on Google thinks the sun is 15.81 light years from earth
Having interacted with a lot of computers over the years (including Google's for the last several of my employment), the machines at Google come as close to thinking as I've ever encountered. Two or three times a day they come up with answers that I cannot begin to comprehend how they arrived at. Sometimes they are brilliant, sometimes they are dumb, and sometimes they are merely mad.

They give every coworker I've ever had a run for their money on catching my fuckups.

jsolson··on Google thinks the sun is 15.81 light years from earth
Wait, what about the whole sum... ohhhhh, right, clouds. Clouds. Always clouds.

No summer ever.

Clouds always.

jsolson··on The mysterious case of the Linux Page Table Isolation patches
Have you run u to latency issues with Postgres?

(I work on GCE)

jsolson··on The Christmas crypto correction. What really happened
You can trade BTC/LTC and BTC/ETH as well, though, and all of the cryptocurrencies can move into/out of them at will. There are steps they could take to ensure liquidity in USD, but they're relatively expensive.

I, personally, prefer to keep my dollars in dollars and in an FDIC insured account.

jsolson··on The Christmas crypto correction. What really happened
I transferred in excess of $50k USD out without incident.

Could I have done the same with $1M USD? I have no idea.

The bar is considerably higher than $1k, though, and I'd wager is considerably higher than most people have to comfortably move around.

jsolson··on The Christmas crypto correction. What really happened
The bit about not being able to cash out over $1k is flat out false.

The order book history from GDAX (I maintain a personal level 3 copy) also doesn't suggest market pressure at $12k beyond a lot of opportunists (myself included).

jsolson··on Max Howell's take on getting rejected by Google
I mean, I suppose I'd be happier with Racket?

Bonus points for building a #lang for expressing trees and operations on them.

(note: this probably wouldn't actually get you bonus points, and might result in my wildly gesticulating to get your attention to try to drag you back to the actual interview)

jsolson··on A difficult experience with Autopilot in the mountains of Canada
This isn't that terrifying guardrail-free road, though -- I drive this in a Model S ~15-20 times in each direction every winter on my way between Seattle and Whistler.

Autopilot is fantastic for that trip between UW and turning onto CA1 where BC99 intersects it. From there up to Lions Bay and from the far side of Lions Bay[0] to Squamish I keep the car in self drive after a few too many attempted suicides by Eddie[1]. It's not that there aren't guard rails, it's just that it really winds about as it hugs the cliff face. For most of that trip Autopilot generally doesn't want to engage -- it knows better when it can't predict the lane boundary far enough into the future. Past Squamish it really depends on conditions. My trip up this Saturday morning was downright boring and autopilot handled the whole thing. On other trips at night the road has essentially been one solid white surface where even humans avoid each other by treating it as one very wide lane in each direction and maintaining lots of clearance.

[0]: Lions Bay has a 60kph speed limit; autopilot can generally handle it fine, regardless of conditions, although the wide variance in prevailing traffic speed makes this among the most dangerous parts of the trip. I often wonder how the accident and injury statistics for Lions Bay compare to the rest of Sea to Sky.

[1]: My Tesla is named Edison's Lament. Yes, you name your car. Yes, I enjoy this feature.

jsolson··on Why Google stores billions of lines of code in a single repository (2016)
No.

Unsubmitted changes at Google usually come in one of two flavors, short-lived (abandoned or submitted within a few days) or perpetual. The latter flavor is often for "I think we might want this". It's not uncommon for those to be completely rewritten if they're actually needed. There's usually a preference for submitting useful things (with tests!) and flag gating them to cut down on bitrot.

I have seen exceptions -- I reported a bug in a fiddly bit of epoll-related code and an engineer on my team had a multi-year-old fix -- he hadn't submitted it because he wasn't confident he'd found an actual bug. The final changelist number was more than double the original CL number (unsubmitted changes get re-numbered to fit in sequence when they're submitted -- the original number redirects to the final submitted version in our tooling).

jsolson··on Why Google stores billions of lines of code in a single repository (2016)
Mercurial (with lots of extensions) sits on top of Piper at Google. It doesn't replace it.
jsolson··on Why Google stores billions of lines of code in a single repository (2016)
While technically true due to some features of tooling, that is really only masking off part of the repo under a READ-ONLY directory.

Builds can (and usually do) depend on things that aren't part of your local checkout.

I'd say CitC is a much more accurate representation of the way Piper and blaze "expect" things to work.

jsolson··on Andromeda 2.1 reduces GCP’s intra-zone latency by 40%
c5 wasn't available to me when I made that comment, or at least c5 numbers weren't — we have them now, although we're observing ~10 µs worse than your one-off in ours.

It's certainly a nice improvement over what we see on the c4s. Is that using a placement group to ensure proximity (I believe our tests do, but I'd have to double check)? Our benchmarking philosophy is generally to aim for "default" numbers for GCP and "best" numbers for others -- keeps us honest about our "fresh out of the box" behavior.

Also, if we should be seeing better on earlier instance types, I'd love to know what we're potentially doing wrong.

jsolson··on Microsoft and GitHub team up to take Git virtual file system to macOS, Linux
The ability to have something like CitC (or this post's git virtual filesystem) is certainly one big advantage -- no need to clone new packages, they're right there in your "local" source tree. Bazel (blaze) is another, particularly when coupled with working at HEAD.

My experience with farms of git repos is that the lack of atomic operations over many tiny repos leads to things like version sets and having to periodically merge dependencies. I've worked on teams where that was inevitably neglected during hectic periods resulting in painful merges of large numbers of changes. That problem simply doesn't exist with working at HEAD and high quality presubmit test automation/admission control. The single repo also allows for single code reviews spanning multiple packages which makes it MUCH simpler to re-arrange code (Bazel again helps here since a "package" is any directory with a BUILD file). Package creation is lighter weight for the same reason, and has fewer consequences for poor name choices since rearrangement is easy and well supported by automated tools.

Sharing one build system where a build command implicitly spans many packages also results in efficient caching of build artifacts and massively distributed builds (think a distributable and cacheable build action per executable command rather than a brazil-build per package). Each unit test result can be cached and only dependent tests re-run as you tweak an in-progress change. This is fantastic for a local workflow (flaky tests can be tackled with --runs_per_test=1000, which with a distributed build system is often only marginally slower than a single test run). Also, you can query all affected tests for a given change with a single "local" bazel query command. The list goes on from here -- I keep thinking of new things to add (finer grained dependencies, finer grained visibility controls, etc.).

It's not that you can't build most of this for distributed repos, but I'd argue it's harder and some things (like ease of code reorg) are nearly impossible.

Subjectively, having worked with both approaches at scale, Google's seems to result in much better code and repo hygiene.

jsolson··on Is software development really a dead-end job after age 35-40?
I cannot agree with this enough.

I'm a pretty capable "generalist software engineer", but what keeps me employed is expertise in virtualization and high-performance virtual networking. Prior to that, it was expertise in distributed systems and a knack for debugging distributed failures. The history here is relevant: later in the thread Thomas points out that you can switch domains. My time spent in distributed systems with multi-millisecond quorum periods is more or less directly applicable to debugging synchronization issues at CPU clock speeds. The tools used to observe the issue are different, but the reasoning process for untangling the set of plausible partial orderings is the same.

Specialize in something valuable. Continuously evaluate the next most valuable skill to acquire based on where you are and where you want to be.

jsolson··on Microsoft and GitHub team up to take Git virtual file system to macOS, Linux
Piper: https://cacm.acm.org/magazines/2016/7/204032-why-google-stor...

Unfortunately, it's not available externally.

Even working at Google, my jaw still dropped at this:

    Google's codebase is shared by more than 25,000 Google software
    developers from dozens of offices in countries around the world.
    On a typical workday, they commit 16,000 changes to the codebase,
    and another 24,000 changes are committed by automated systems.
    Each day the repository serves billions of file read requests,
    with approximately 800,000 queries per second during peak traffic
    and an average of approximately 500,000 queries per second each
    workday. Most of this traffic originates from Google's
    distributed build-and-test systems.
Mostly I felt terribly unproductive — my changelist generation rate is way below average.
jsolson··on Tesla Roadster
That was my initial reaction, but it leads to a question: do you stop doing R&D (and shed talent as a result) because production in another part of the company is blocked by some (hopefully transient) supply chain issue?
jsolson··on FreeBSD/EC2 on C5 instances
Indeed! I didn't leave too much to the imagination with my replies on the original post[0], though.

Honestly, I'm mostly curious about how much of "KVM" you're running that's stock, how much is modified, and how much of the userland is running on the far side of PCIe rather than in host ring3 (particularly given "C5 instances are built using a new light-weight hypervisor, which provides practically all of the compute and memory resources to customers’ instances.").

[0]: Especially this one: https://news.ycombinator.com/item?id=15641391

jsolson··on FreeBSD/EC2 on C5 instances
Yes, sorry, I fully expect that's what they're doing -- I was really writing two replies there --

1) You don't have to pass PCIe through to hardware to get performance that makes it look like you have.

2) We know they're offloading though, so here's what I think it looks like.

jsolson··on FreeBSD/EC2 on C5 instances
I do, hence my speculation about putting the NVMe in firmware instead of hardware :)

(Colin's speculation in a peer reply is also reasonable -- personally I've seen enough errata in "off the shelf" IP to shudder at the idea of anything in silicon that doesn't have to be, but fundamentally I'm a SWE, so that would be my take, wouldn't it)

jsolson··on FreeBSD/EC2 on C5 instances
It's fun to speculate about how other clouds do things :)

> There's absolutely no way that they would get the performance I'm seeing from an emulated disk. We're talking to real hardware, exposed via PCI passthrough.

There's a wide spectrum between "emulated" and "real hardware, exposed via PCI passthrough". Passing through to PCI hardware, in and of itself, gains you very little in terms of absolute guest-visible performance versus eliding all VMEXITs via other means, but it has other important characteristics that I suspect AWS very much wants in the c5 family.

> I would assume it's something like "NVME interface hardware" + "ARM CPU which implements the EBS protocol" + "25 GbE PHY", but that guess is based solely on "that's how I would design it".

I would expect something along these lines, although I'd be a little surprised if they bothered putting the NVMe bits down in silicon. My personal guess would be a good silicon DMA engine and PCIe interface married to sufficient general-purpose processing (ARM SoC, FPGA, etc.) to keep pace with the NVMe Command/Completion queues.

Regardless of how they've implemented it, the end-to-end result seems to hang together very nicely. Kudos to the team at AWS.

(note: I work on Google Compute Engine's hypervisor; my speculation about AWS really is speculation about how they'd do this — the two companies have very different engineering approaches, so I may be entirely wrong trying to project onto theirs — like I said at the top, it's fun to try :)

jsolson··on Andromeda 2.1 reduces GCP’s intra-zone latency by 40%
:)
jsolson··on Andromeda 2.1 reduces GCP’s intra-zone latency by 40%
edit: It's late here and I think I misread this originally :)

The differential between public IPs and internal IPs is tied into the path packets take after leaving the host. The path out of the guest is identical for both, but using VM public IPs (rather than internal) can result in passing through additional hops versus being routed straight to the target VM. Common firewall configurations can also impact perf here.

Original comment:

With respect to guest CPU, the approach used by Andromeda 2.1 eliminates VM exits both on transmit and for interrupt delivery (where supported by Intel). In that regard it's essentially identical to PCIe passthrough. There are customers running DPDK to further reduce variance (and eliminate the cost of interrupt handling entirely).

The choice to not pass through host hardware comes down to a few factors, but high on the list are supporting live migration and NIC vendor flexibility.

(I worked on this effort; see other comments for specifics)

jsolson··on Andromeda 2.1 reduces GCP’s intra-zone latency by 40%
It's similar, although distinct. By building on a common foundation of Google networking dataplane bits, Jake's team (and peer teams) get easier integration with Google's other networking infrastructure for features like DoS protection, encryption, etc. The core bits underlying Andromeda 2.1 are related to those used for Espresso (https://www.blog.google/topics/google-cloud/making-google-cl... — HN discussion: https://news.ycombinator.com/item?id=14037830).
jsolson··on Andromeda 2.1 reduces GCP’s intra-zone latency by 40%
Replied to a sibling, but I believe our latency is currently coming in under theirs. Their largest/newest VMs advertise higher peak bandwidth than we do. The latency difference would certainly be most directly visible in microbenchmarks, although HPC applications and those relying on in-memory databases are also likely to see practical benefit.
← PreviousPage 8 of 24Next →