2,421 karma · joined February 21, 2007
jonolson at google dot com
Opinions are my own, not those of my employer.
I picked up some at the money May calls Monday. I sold them today, and (mostly coincidentally) bought a sailboat. The options, after taxes, covered the boat and this year's moorage.
Could've gone the other way of course, but after the beating they'd taken? Surely they'd do something to correct between Monday and earnings.
They said, I might amend your statement to be emotional volatility is your friend. I mostly got lucky on the response to Tesla coming closer to meeting their production goals (entirely independently of the underlying financials). It elicited a predictable emotional response.
I would like to attribute my desire to live "forever", at least to some extent, to cosmic perspective.
One day the universe might make me bored, but today is not that day, and I suspect tomorrow isn't either.
I don't ever expect that day to come, but if I thought it was near I'd be very afraid indeed.
Is the Bay Area actually opposed to that?
I understand what you're saying here, but the extent to which the aerospace industry just does not work this way is staggering.
This is a fairly unique quality of SpaceX.
They give every coworker I've ever had a run for their money on catching my fuckups.
No summer ever.
Clouds always.
(I work on GCE)
I, personally, prefer to keep my dollars in dollars and in an FDIC insured account.
Could I have done the same with $1M USD? I have no idea.
The bar is considerably higher than $1k, though, and I'd wager is considerably higher than most people have to comfortably move around.
The order book history from GDAX (I maintain a personal level 3 copy) also doesn't suggest market pressure at $12k beyond a lot of opportunists (myself included).
Bonus points for building a #lang for expressing trees and operations on them.
(note: this probably wouldn't actually get you bonus points, and might result in my wildly gesticulating to get your attention to try to drag you back to the actual interview)
Autopilot is fantastic for that trip between UW and turning onto CA1 where BC99 intersects it. From there up to Lions Bay and from the far side of Lions Bay[0] to Squamish I keep the car in self drive after a few too many attempted suicides by Eddie[1]. It's not that there aren't guard rails, it's just that it really winds about as it hugs the cliff face. For most of that trip Autopilot generally doesn't want to engage -- it knows better when it can't predict the lane boundary far enough into the future. Past Squamish it really depends on conditions. My trip up this Saturday morning was downright boring and autopilot handled the whole thing. On other trips at night the road has essentially been one solid white surface where even humans avoid each other by treating it as one very wide lane in each direction and maintaining lots of clearance.
[0]: Lions Bay has a 60kph speed limit; autopilot can generally handle it fine, regardless of conditions, although the wide variance in prevailing traffic speed makes this among the most dangerous parts of the trip. I often wonder how the accident and injury statistics for Lions Bay compare to the rest of Sea to Sky.
[1]: My Tesla is named Edison's Lament. Yes, you name your car. Yes, I enjoy this feature.
Unsubmitted changes at Google usually come in one of two flavors, short-lived (abandoned or submitted within a few days) or perpetual. The latter flavor is often for "I think we might want this". It's not uncommon for those to be completely rewritten if they're actually needed. There's usually a preference for submitting useful things (with tests!) and flag gating them to cut down on bitrot.
I have seen exceptions -- I reported a bug in a fiddly bit of epoll-related code and an engineer on my team had a multi-year-old fix -- he hadn't submitted it because he wasn't confident he'd found an actual bug. The final changelist number was more than double the original CL number (unsubmitted changes get re-numbered to fit in sequence when they're submitted -- the original number redirects to the final submitted version in our tooling).
Builds can (and usually do) depend on things that aren't part of your local checkout.
I'd say CitC is a much more accurate representation of the way Piper and blaze "expect" things to work.
It's certainly a nice improvement over what we see on the c4s. Is that using a placement group to ensure proximity (I believe our tests do, but I'd have to double check)? Our benchmarking philosophy is generally to aim for "default" numbers for GCP and "best" numbers for others -- keeps us honest about our "fresh out of the box" behavior.
Also, if we should be seeing better on earlier instance types, I'd love to know what we're potentially doing wrong.
My experience with farms of git repos is that the lack of atomic operations over many tiny repos leads to things like version sets and having to periodically merge dependencies. I've worked on teams where that was inevitably neglected during hectic periods resulting in painful merges of large numbers of changes. That problem simply doesn't exist with working at HEAD and high quality presubmit test automation/admission control. The single repo also allows for single code reviews spanning multiple packages which makes it MUCH simpler to re-arrange code (Bazel again helps here since a "package" is any directory with a BUILD file). Package creation is lighter weight for the same reason, and has fewer consequences for poor name choices since rearrangement is easy and well supported by automated tools.
Sharing one build system where a build command implicitly spans many packages also results in efficient caching of build artifacts and massively distributed builds (think a distributable and cacheable build action per executable command rather than a brazil-build per package). Each unit test result can be cached and only dependent tests re-run as you tweak an in-progress change. This is fantastic for a local workflow (flaky tests can be tackled with --runs_per_test=1000, which with a distributed build system is often only marginally slower than a single test run). Also, you can query all affected tests for a given change with a single "local" bazel query command. The list goes on from here -- I keep thinking of new things to add (finer grained dependencies, finer grained visibility controls, etc.).
It's not that you can't build most of this for distributed repos, but I'd argue it's harder and some things (like ease of code reorg) are nearly impossible.
Subjectively, having worked with both approaches at scale, Google's seems to result in much better code and repo hygiene.
I'm a pretty capable "generalist software engineer", but what keeps me employed is expertise in virtualization and high-performance virtual networking. Prior to that, it was expertise in distributed systems and a knack for debugging distributed failures. The history here is relevant: later in the thread Thomas points out that you can switch domains. My time spent in distributed systems with multi-millisecond quorum periods is more or less directly applicable to debugging synchronization issues at CPU clock speeds. The tools used to observe the issue are different, but the reasoning process for untangling the set of plausible partial orderings is the same.
Specialize in something valuable. Continuously evaluate the next most valuable skill to acquire based on where you are and where you want to be.
Unfortunately, it's not available externally.
Even working at Google, my jaw still dropped at this:
Google's codebase is shared by more than 25,000 Google software
developers from dozens of offices in countries around the world.
On a typical workday, they commit 16,000 changes to the codebase,
and another 24,000 changes are committed by automated systems.
Each day the repository serves billions of file read requests,
with approximately 800,000 queries per second during peak traffic
and an average of approximately 500,000 queries per second each
workday. Most of this traffic originates from Google's
distributed build-and-test systems.
Mostly I felt terribly unproductive — my changelist generation rate is way below average.Honestly, I'm mostly curious about how much of "KVM" you're running that's stock, how much is modified, and how much of the userland is running on the far side of PCIe rather than in host ring3 (particularly given "C5 instances are built using a new light-weight hypervisor, which provides practically all of the compute and memory resources to customers’ instances.").
[0]: Especially this one: https://news.ycombinator.com/item?id=15641391
1) You don't have to pass PCIe through to hardware to get performance that makes it look like you have.
2) We know they're offloading though, so here's what I think it looks like.
(Colin's speculation in a peer reply is also reasonable -- personally I've seen enough errata in "off the shelf" IP to shudder at the idea of anything in silicon that doesn't have to be, but fundamentally I'm a SWE, so that would be my take, wouldn't it)
> There's absolutely no way that they would get the performance I'm seeing from an emulated disk. We're talking to real hardware, exposed via PCI passthrough.
There's a wide spectrum between "emulated" and "real hardware, exposed via PCI passthrough". Passing through to PCI hardware, in and of itself, gains you very little in terms of absolute guest-visible performance versus eliding all VMEXITs via other means, but it has other important characteristics that I suspect AWS very much wants in the c5 family.
> I would assume it's something like "NVME interface hardware" + "ARM CPU which implements the EBS protocol" + "25 GbE PHY", but that guess is based solely on "that's how I would design it".
I would expect something along these lines, although I'd be a little surprised if they bothered putting the NVMe bits down in silicon. My personal guess would be a good silicon DMA engine and PCIe interface married to sufficient general-purpose processing (ARM SoC, FPGA, etc.) to keep pace with the NVMe Command/Completion queues.
Regardless of how they've implemented it, the end-to-end result seems to hang together very nicely. Kudos to the team at AWS.
(note: I work on Google Compute Engine's hypervisor; my speculation about AWS really is speculation about how they'd do this — the two companies have very different engineering approaches, so I may be entirely wrong trying to project onto theirs — like I said at the top, it's fun to try :)
The differential between public IPs and internal IPs is tied into the path packets take after leaving the host. The path out of the guest is identical for both, but using VM public IPs (rather than internal) can result in passing through additional hops versus being routed straight to the target VM. Common firewall configurations can also impact perf here.
Original comment:
With respect to guest CPU, the approach used by Andromeda 2.1 eliminates VM exits both on transmit and for interrupt delivery (where supported by Intel). In that regard it's essentially identical to PCIe passthrough. There are customers running DPDK to further reduce variance (and eliminate the cost of interrupt handling entirely).
The choice to not pass through host hardware comes down to a few factors, but high on the list are supporting live migration and NIC vendor flexibility.
(I worked on this effort; see other comments for specifics)