HNHacker News
TopNewBestAskShowJobs

jsolson

2,421 karma · joined February 21, 2007

Google principal engineer based out of Seattle. I am the technical lead for delivering the next generations of AI GPU Supercomputers to Google Cloud. Previously, I worked on the hypervisor that powers Google Compute Engine.

jonolson at google dot com

Opinions are my own, not those of my employer.

submissionscomments
jsolson··on AWS Ground Station
I am assuming that anyone who managed to accumulate 100 PB of data (or close enough to make this viable) is still accumulating data at a solid clip, or likely has a lot _more_ than just 100 PB of data
jsolson··on AWS Ground Station
I said Snowmobile doesn't make sense. Snowball makes total sense. The economics are entirely different.

Antarctica was the only place that came to mind, but I wasn't sure about the costs of getting a shipping container in / out.

jsolson··on AWS Ground Station
I work on Google Cloud and have an obvious bias here, but having thought about this for a couple year now, this does not seem like a good or cost-effective way to move that much data.

In particular if you have ~100PB of data to move and you are in a location that can be reached by a 45-foot long shipping container on a truck with access to the ~350KW of power it takes to run Snowmobile, you're clearly not somewhere _completely_ inaccessible. Given that 100PB of data will cost you $400k/mo in storage costs on AWS Glacier (before the discounts that you'll obviously negotiate for), even relatively remote locations become "accessible" for 100 Gbps+ fiber or microwave connectivity, and _very_ remote locations for tens of Gbps. The Snowmobile itself is only 1 Tb/s, so if you think you're going to save time this work, consider how long it takes to move it to you and then move it to an AWS facility versus the time it takes to fill it.

I don't get where this is the right call for _any_ customers, even as a one-off. I'd love it if someone from AWS could tell me where my math is off on this.

jsolson··on Why Is Facebook Not in the Cloud Business?
Thanks for the feedback here -- as I mentioned in another comment I'm generally fond of our UIs, and it's nice to know that I'm not _entirely_ crazy.
jsolson··on Why Is Facebook Not in the Cloud Business?
Thanks for the comment -- I think (although I honestly don't know for certain) that the goal is for the names to convey something about what the service does. In particular, so that in the context of the catalog the general purposes of the various services are at least _somewhat_ obvious.

The docs are always a work in progress -- anything you can recall trying to do with GCP that was particularly inscrutable (can't fix everything in one go, but one fix at a time is better than nothing :)

jsolson··on Why Is Facebook Not in the Cloud Business?
Thanks for the feedback -- I'd be curious to understand what about the Azure UI you find most striking. For obvious reasons, I don't have much experience with it.

I'm generally fond of our console, particularly the tools that show both the API and CLI invocations for most UI activities. On the other hand, I work on core virtualization, so the things I want to express tend to be very simple ("make me a giant VM", "make me a gianter VM", etc. :)

jsolson··on Why Is Facebook Not in the Cloud Business?
> a lot of developers hate it

Any particular complaints?

(I work on Compute Engine)

jsolson··on Bottlerocket: An operating system designed for hosting containers
Honestly this reply was worth the karma loss.

Hi Matt!

jsolson··on Google recommends all North America employees work from home
Folgers crystals can address your caffeine addiction. They taste like shit, but you'll live and it'll address the chemical needs, speaking from experience. Alternatively, you can buy and freeze pounds of ground coffee and use a filter and hot water for a similar (but vastly better tasting) effect.

So sorry about the dishes, but perhaps your dish washing efficiency will improve? If it's directly impactful to your productivity now for a single individual's daily calories it suggests you've got a lot to learn here! Perhaps a G2G course once normal business returns?

In terms of the monitor I can't help you, my most productive years were on an 11" MacBook air with a single tmux session. A 13" laptop sounds heavy.

(Disclaimer: I work at Google)

jsolson··on “Just walk out” technology by Amazon
I did my most recent liquor run at the nearby Amazon Grocery.

They check ID entering the liquor section (and it is a human that checks), but otherwise it is bag it and just walk out.

jsolson··on Stop using Material Design text fields
> On android's slide down config screen (the one with wifi, bluetooth, etc toggles) I always have to stop and think a bit to determine if wifi is on or not.

I'm a bit surprised by this -- it transitions from dark gray on light gray when it's off to _bright_ white on _bright_ blue when it's on. This would seem to mirror many physical things that have status lights which are not illuminated when disabled and are illuminated when enabled. What's the source of initial confusion for you?

jsolson··on Measuring Latency in Linux (2014)
Intel's VMCS includes a TSC offset field as well as TSC scaling. These allow for a stable RDTSC across migrations between hosts, modulo actual time lost to migration blackout.

(I work on virtualization in GCE)

jsolson··on I've screwed up plenty of things too
I was once (partially) responsible for the deaths of dozens of virtual machines at a distance of about three and a half years.

Fun fact: none of these VMs had rebooted in that time, or they wouldn't have crashed.

Anyway, back in 2014 or so I dropped a bunch of transmit packet completions. In most cases I also double completed packets which was immediately fatal. Kernels get mad about that sort of thing.

Turns out, not all of the affected VMs died. Some of them lived on with head indices forever unequal to tail indices (until they rebooted).

In 2018 a developer realized there was a potential bug in waiting for VMs entering a quiescent state -- a truly idle networking stack had retired all Tx packets that it had admitted. Having unequal indices was impossible under correct operating conditions. They fixed the glitch.

This change rolled out gradually.

Gradually, the kernel panics appeared.

The change rolled back, halting the impact, but then the analysis began. What had we broken?

Another fun fact: Linux often includes an uptime in dmesg logs.

Slowly a pattern appeared. The dmesg logs included unusually large numbers for uptimes. Plotting these, there was a clear cliff in terms of a minimum uptime. Historical deployment logs showed a noteworthy release at that date, years past. Noteworthy in that it was rolled back for my bug, years prior.

On the plus side, I realized this was almost certainly my years prior fuckup slightly sooner than anyone else, so at least I got to call myself out :)

jsolson··on Volkswagen exec admits full self-driving cars 'may never happen'
Some humans, when seeking to end their own lives, end the lives of others: https://en.wikipedia.org/wiki/Germanwings_Flight_9525

That's an extreme example, but automotive suicides that kill other passengers, drivers, or pedestrians fall into the same category. Consider also deaths from accidents involving drunk driving or fatigue -- thousands of motorists take to the roads every day modified in one manner or another that reduces their driving aptitude.

Also, while it may be correct to say that computers don't "fear death", there's no reason that "risk to self" can't be part of the criteria for decision making by an autonomous system.

jsolson··on Volkswagen exec admits full self-driving cars 'may never happen'
In an extreme example I expect that's precisely what would happen. Consider what's currently unfolding around the 737 Max. In the automotive space there's a long history of serious flaws that resulted in loss of life, ranging from faulty airbag deployment systems to flawed designs for ignition systems.

We have precedent for how we qualify and evaluate things for safety: test them across a variety of conditions, accumulate driver-miles or operator-hours and incident frequencies. Then, using that data establish a bar for what constitutes an acceptable level of risk given the utility something provides. If we wanted to ensure nobody ever died in a car accident, we would ensure there were no cars, but collectively we've made a different choice.

jsolson··on Volkswagen exec admits full self-driving cars 'may never happen'
You may not be getting in that car, but I certainly will.

After all, every driver on the road today is an incomprehensible black box where not only do we not know the parameters, we don't even know the function they're parameterizing. Every instance functions differently, and our testing procedures have woefully low coverage.

jsolson··on Which Machines Do Computer Architects Admire? (2013)
I remember getting an MSDN CD wallet that included a build of Windows NT for Alpha. Of course, I never had an Alpha on which to run it.
jsolson··on Debian votes for Proposal B, “Systemd but we support exploring alternatives”
I'm also mostly disconnected from this, which honestly makes me a fan of systemd.

Today (literally) I wanted to launch Xilinx hw_server as a daemon that my peers could restart if it broke itself (it's... prone to doing so). While I could write an init script that knew about PID files, creating a systemd unit was _vastly_ easier.

Yeah, it's not Unix, but UNIX is also sort of terrible for anything that's not a one-shot, no?

jsolson··on Bazel 2.0
I don't know the team specifically, but I suspect the difference comes down in part to Go open sourcing early in development, thus finding a bunch of the rough edges that exist outside Google's walled garden early in the project's life when there wasn't much compatibility _to_ break. Blaze was a mature project within Google for years before Bazel was opened up. Many of the breaking changes seem to be taking one-offs built for features within Google (e.g., the handling of protocol buffer rules) and building those in terms of more general and composeable features.

The net result is a Bazel (and Blaze) that are less burdened by the baggage of legacy, but the cost is a faster treadmill to keep pace with changes.

jsolson··on Google claims copyright on employee side projects
I do, yes, and that Google grants that permission relatively (to some others companies) freely is a small part of why I work there. I could quit and work on whatever I want, but eventually I'd need to worry about how to pay my mortgage. I could also go work for a smaller company without such a clause in their contract, but part of what I like about my job is the resources available to tackle problems, and I'd likely lose some of that.

Employment often imposes lots of restrictions on the actions we can take. I'm fortunate enough to have some choice in the set of restrictions I have to live with, so for me this particular restriction is just part of the deal.

jsolson··on The hardest aspect of learning English as a second language
I mean, we already disagree on spelling for some words, the name for the letter 'Z', and what to call dwellings in multi-family properties, so we're not off to a great start...
jsolson··on Google claims copyright on employee side projects
> Maybe we should be happy it took google so long.

If you read the thread you'll note that this policy has been around for a long time (at least the ~7 years I've been with Google). The author also notes that, in terms of open source, things have become dramatically _more_ permissive over time.

jsolson··on Create-React-App 3.3
Swift?

Or, depending on how you feel about sending a message to nil simply returning nil, Objective-C?

Swift's variant is much more modern, though, as it has the ability to inline assert.

https://docs.swift.org/swift-book/LanguageGuide/OptionalChai...

jsolson··on How to do a code review
> This is a professional procedure in a professional setting, not warm words from an encouraging teacher at school...

Citing a good practice in someone's code as "yes, please do more of this" alongside "don't do this please" is not, in my opinion, fluff.

> I don't want to have to go through comments that do not add any value to the exercise of finding issues, and I have never seen people leave such comments in 20 years.

So only pay attention to unresolved comments?

jsolson··on How to do a code review
I disagree, with a caveat. Assuming your review tool supports "Resolved" or "No Action Required" comments (ours does), it's rather easy to distinguish something that is informative from something that is actionable. I am now recalling that some review tools don't distinguish open/unresolved comments from resolved or informative comments, which would make this more of a trade-off than an obvious win.
jsolson··on How to do a code review
> Unless there's something extremely interesting, overly cautious phrasing and "positive" comments are mostly noise.

Not really. Positive comments help teach engineers which things they've done that conform to local best practices (and why) without them having to meticulously dig those up (assuming they're even documented). A lack of positive comments leaves engineers to learn them only by running afoul of them. Effectively it provides direction only when some threshold of badness is crossed, while leaving positive comments on good code (especially for new team members or junior engineers) provides a beacon pointing away from the badness threshold entirely.

Put differently: commenting on good code makes for swifter and less eventful code reviews by steering engineers away from the bad practices that make code challenging to review in the first place.

Also, it's just nicer to spend forty hours a week with people who demonstrably appreciate their peers' good work.

jsolson··on QEMU VM Escape
Some of it is academic study (BS and MS at Georgia Tech) -- that helped with the theory and a bit of the breadth. That said, most of what I know of modern virtualization was more or less "on the job" from working on Google Compute Engine. Lots of practical experience understanding virtualization boundaries, and really really enormous access to historical domain experts in the space.
jsolson··on QEMU VM Escape
It's quite reasonable to distinguish QEMU from the hypervisor: QEMU typically runs with strictly lower privileges.

The "hardware acceleration" that helps QEMU emulate a computer operates in root VMX and non-root VMX (when actually executing guest instructions) modes. When operating in root VMX mode as a supervisor, it has approximately the highest privileges in the system (ignoring SMM), as it is operating in host ring 0. QEMU runs in host ring 3, making it a typical user mode process. If you manage to compromise QEMU (and only QEMU), you have only escaped into host user mode. In that case you haven't compromised the hypervisor, merely the virtual machine monitor.

If, on the other hand, you manage to compromise KVM (or vhost) you could reasonably claim to have actually compromised the hypervisor.

Type 2 hypervisors make this generally much uglier than type 1, but it's reasonable to say that the hypervisor in a type 2 deployment is limited to the kernel component rather than the user-mode VMM.

jsolson··on Why do developers at Google consider Agile development to be nonsense? (2016)
Beyond the literal fact that "everyone" is undeniably a giant mass of people, I suspect this term was intended to mean "to consumers", particularly given that the following example calls out enterprises.

Developing for a small but demanding set of enterprise customers with concrete (but sometimes arcane) requirements is a very different problem from developing for consumer markets.

jsolson··on The CUE Data Constraint Language
That's common to most projects released by Googlers. Options are either a getting the IP officially assigned over to yourself or releasing it under a Google copyright with a CLA: https://opensource.google.com/docs/creating/

The latter is lower friction, in my opinion.

This is especially common for "personal" projects that people end up working on or using on the clock.

I don't know if this was the case with Cue specifically (this is the first I've heard of the project, but less boilerplate in my JSON has some appeal).

← PreviousPage 6 of 24Next →