HNHacker News
TopNewBestAskShowJobs

alex-mohr

142 karma · joined June 22, 2015

Engineering Manager for the Seattle branch of Google's Kubernetes and Container Engine (GKE) team.

Previously, I was the overall technical lead for Google Compute Engine's initial public launch; also lead the design and launch of its second-gen VM instance manager subsystem.

And before all of that, ran a research lab at Stony Brook University [1], but I'm a reformed academic and prefer building real systems now.

[1]: http://alexmohr.com/papers/

submissionscomments
alex-mohr··on Magical systems thinking
Process is useful for raising the lowest deliveries quality, for making former-unknowns into knowns, and for preventing misaligned behavior when culture alone becomes insufficient.

If you have need for speed, a team that knows the space, and crucially a leader who can be trusted to depart from the usual process when that tradeoff better meets business needs, it can work really well. But also comes with increased risk.

alex-mohr··on AI at Amazon: A case study of brittleness
And you could write a similar blog post about why Google "failed" at AI productization (at least as of a year ago). For some of the same and some completely different reasons.

  - two competing orgs via Brain and DeepMind.

  - members of those orgs were promoted based on ...?  Whatever it was, something not developing consumer or enterprise products, and definitely not for cloud.

  - Nvidia is a Very Big Market Cap company based on selling AI accelerators.  Google sells USB Coral sticks.  And rents accelerators via Cloud.  But somehow those are not valued at Very Big Market Cap.
Of course, they're fixing some of those problems: brain and DeepMind merged and Gemini 2.5 pro is a very credible frontier model. But it's also a cautionary tale about unfettered research focus insufficiently grounded in customer focus.
alex-mohr··on Why did Windows 7 log on slower for months if you had a solid color background?
It was in the early days of Kubernetes and long since fixed. I don't recall the precise details, but it was likely the first official CVE we published: https://kubernetes.io/docs/reference/issues-security/officia...

Link to the patch fixing it: https://github.com/kubernetes/kubernetes/commit/7fef0a4f6a44...

Of course, we'd already fixed other issues like Kubelet listening on a secondary debug port with no authentication. Those problems stemmed from its origins as a make-it-possible hacker project and it took a while to pivot it to something usable in an enterprise.

alex-mohr··on Why did Windows 7 log on slower for months if you had a solid color background?
The code in question reminds me a lot of my favorite Kubernetes bug:

  if (request.authenticationData) {
    ok := validate(etc);
    if (!ok) {
      return authenticationFailure;
    }
  }
Turns out the same meme spans decades.
alex-mohr··on Comparing Fuchsia components and Linux containers [video]
As far as I could tell, its main goal was to have fun writing an OS. At that, it seems to have succeeded for a number of the people involved?

In terms of impact or business case, I'm missing what the end goal for the company or execs involved is. It's not re-writing user-space components of AOSP, because that's all Java or Kotlin. Maybe it's a super-longterm super-expensive effort to replace Linux underlying Android with Fuchia? Or for ChromeOS? Again, seems like a weird motivation to justify such a huge investment in both the team building it and a later migration effort to use it. But what else?

alex-mohr··on Intel doesn't know how to be a foundry, Tim Cook reportedly said in 2011
Clearly the next step after building your own CPU and SOC is to start Apple Foundry and become totally vertically integrated?
alex-mohr··on Intel doesn't know how to be a foundry, Tim Cook reportedly said in 2011
Also a myth for GCE.

From a technical perspective, App Engine and Compute Engine were built on top of internal infrastructure (borg), but did not expose borg directly. And there were a number of interesting mismatches between the semantics that customers expected of VMs and what borg offered to its containers that eventually resulted in dedicated borg clusters with different configs for cloud. And some retrospectives on whether building on borg was a better option than going bare metal directly.

Org-wise, the App Engine team was first and not part of the internal-focused Technical Infrastructure teams. GCS came next, and it too was not part of the canonical storage org. Then GCE, which was only possible because it was either written off or at least tolerated as an experiment by most, with a few key people providing behind-the-scenes support to make it happen -- especially in networking. It likely also helped that GAE was in SF and the rest of GCP in Seattle/Kirkland initially, so geo provided some insulation too.

The dominant perspective internally was that Google's technical infrastructure was its secret sauce, so why would they give it away to others? It took a long time to change that.

[Disclosure/source: I was on GCE and helped get it launched.]

alex-mohr··on Netflix buffering issues: Boxing fans complain about Jake Paul vs. Mike Tyson
If Netflix were working correctly and could handle the load, you'd absolutely be correct.

But it does seem the capacity of a hybrid system of Netflix servers plus P2P would be strictly greater than either alone? It's not an XOR.

And note that in this case of "live" streaming, it still has a few seconds of buffer, which gives a bandwidth-delay product of a few MB. That's plenty to have non-stale blocks and do torrent-style sharing.

alex-mohr··on Netflix buffering issues: Boxing fans complain about Jake Paul vs. Mike Tyson
Yes, the properties about scaling do hold even with near-real-time streams. [1]

The problems with using it as part of a distributed service have more to do with asymmetric connections: using all of the limited upload bandwidth causes downloads to slow. Along with firewalls.

But the biggest issue: privacy. If I'm part of the swarm, maybe that means I'm watching it?

[1]: Chainsaw: P2P streaming without trees, https://link.springer.com/chapter/10.1007/11558989_12

alex-mohr··on Show HN: Allocate poker chips optimally with mixed-integer nonlinear programming
J. Shallit (2003). "What this country needs is an 18c piece" (PDF). Mathematical Intelligencer. 25 (2): 20–23.
alex-mohr··on Willow Protocol
Willow appears closer to a "Protocol Construction Kit" than a protocol itself.

As a construction kit, it has value for people who want to make protocols where they'll control both ends, but don't have to re-implement basic table stakes.

alex-mohr··on What This Country Needs is an 18¢ Piece (2002) [pdf]
Maybe interesting aside: I saw a link to this when it was first published at the height of the p2p networks craze and noticed some similarities between the two.

One of my students at the time, Mahadev Konar, ended up writing a paper "Ring-like DHTs and the Postage Stamp Problem" [1] that shows how you can use solutions to the postage stamp problem (aka denomination-choosing problem) as a way to structure the finger pointers in Chord. And went on to co-found Hortonworks.

Sometimes random things on HN end up having implications in other areas!

[1]: https://alexmohr.com/papers/dht-postage-stamp-podc2005-exten...

alex-mohr··on Awesome Engineering Games
I get that, but IMO it's not an issue.

On one shard you can't use more than base CPU, so there's no advantage to a subscription there.

Other optional shards are almost entirely people who subscribe. Do that too if you decide to, or ignore them.

alex-mohr··on Awesome Engineering Games
Look at Screeps: World [0] for depth in a programming base-builder RTS.

You write the code for each of your units, either natively in Javascript or Typescript, or via WASM you can run Rust, Python, etc. You use a private server or join a shared MMO world. There's a free sim [1] to try out the basics, though the actual game has much more depth. And an active Discord for help [2].

There's also a variant Screeps: Arena [3] that focuses on 1:1 PVP battles with ranked ladders if you prefer short-lived matches to a long-running world.

[0] https://store.steampowered.com/app/464350/Screeps_World/

[1] https://screeps.com/a/#!/sim

[2] https://discord.com/invite/RjSS5fQuFx

[3] https://store.steampowered.com/app/1137320/Screeps_Arena/

alex-mohr··on Stopping at 90%
See also Chris Lema's "Done Done" essays[0] and book[1] for more about the value of taking things to finished.

[0]: https://chrislema.com/a-done-done-culture-habit-one/

[1]: https://www.amazon.com/Creating-Done-Culture-Habits-Performe...

alex-mohr··on AWS networking concepts in a diagram
> if the goal was to simplify things, then I'm not sure how successful they were at that.

Maybe if they have other goals that are tradeoffs vs. simplicity, then it's more understandable?

What if another goal is allowing enterprise customers to recreate a virtual enterprise network? Or a virtual data center network? Those are much more complex than a client TCP stack.

The defaults are simple for simple uses. And yes, for those complex cases, you'd need AWS specific product knowledge, but most of the underlying concepts are shared in common with other clouds and on prem networks. Like learning your Nth programming language.

alex-mohr··on Ask HN: How do I get back into the tech industry after 4 years of unemployment?
> I don't use any networking sites such as LinkedIn etc

Reconsider that. You can try to add everyone you remember from previous companies or school, and you'll be more available to recruiters. "Luck surface area."

alex-mohr··on WAN router IP address change blamed for global Microsoft 365 outage
It does seem like network configuration remains rather manual compared to other large scale systems that include more automation.

In Microsoft's case, the remediation is not to put in place higher level systems to safely accomplish the goal of the command. Instead:

- "We have blocked highly impactful commands from getting executed on the devices (Completed)"

- "We will require all command execution on the devices to follow safe change guidelines (Estimated completion: February 2023)"

Requiring commands to follow guidelines sounds suspiciously like they're requiring network ops not to break things.

alex-mohr··on Staff Engineer Archetypes (2020)
Member of Technical Staff, or MTS
alex-mohr··on Ask HN: How did you overcome perfectionism?
Perfectionism is too short a label, and could be one or many of multiple underlying issues. Ultimately, you seem aware of the tendency, so ask yourself: why can't I do "good enough" and move forward?

Maybe the fear of actually doing something outweighs the costs of not doing it under the guise of making it better? Maybe something else?

I found Bezos's thoughts on decision making useful: understand if the decision is a two way door, and if so, move forward with 70% of the data you wish you had. Or 70% of the product, implementation, whatever.

See e.g. https://www.aboutamazon.com/news/company-news/2016-letter-to...

alex-mohr··on ‘Positive deviants’: Why rebellious workers spark great ideas
For a decent take on that organizational deficiency, Rickover wrote a 1.5 page memo in 1953 addressing similar issues for the nuclear US Navy. At some point, it's a failing of process and structure rather than people. [1]

1: http://ecolo.org/documents/documents_in_english/Rickover.pdf

alex-mohr··on Open source projects should run office hours
Something similar is one reason early Kubernetes was successful: many of the core people in the project were in an irc/slack channel available for anyone with questions.

That direct two-way communication both unblocked adoption for users and was a source of feedback for the devs.

alex-mohr··on Early Retirement May Speed Up Cognitive Decline: Study
Choose. What you happen to be feeling at the moment can control you, or not. See e.g. https://www.google.com/search?q=discipline+is+freedom
alex-mohr··on USB-C Has Finally Come into Its Own
Chargers are finally coming: here's one that does 48w from one usb-c port or 30w+18w from 2 ports (plus 2 usb-a), plenty for a non-gaming laptop and phone: https://www.amazon.com/Universal-International-Worldwide-Mul.... I have one and it works fine.

Only battery pack with 2 usb-c seems to be the ZMI Ambi: https://www.amazon.com/ZMI-PowerPack-Ambi-USB-C-Power/dp/B07....

alex-mohr··on Talos: OS for Kubernetes
Generally seems like a great offering!

I see immutable, but also upgradable? Is that via in-place upgrades or do upgrades require a reboot?

Example: severe bug or vulnerability in kubelet or containerd/docker. Can I use the API to roll out a fix to existing nodes such that running workloads have no disruption?

alex-mohr··on Ask HN: Academics who switched to industry, what's your experience been like?
Matt Welch has written extensively about switching from tenured Professor at Harvard to Software Engineer at Google: http://matt-welsh.blogspot.com/, with the initial post http://matt-welsh.blogspot.com/2010/11/why-im-leaving-harvar... Matt's blog is great because (a) he writes it, (b) he's still connected to academia via program committees and (c) you can see how his thinking has evolved over the last 8 years.

And there are of course a number of other former faculty there too, but none that I know of who've blogged as much as Matt. In the systems space, off the top of my head: Amin Vahdat, Mike Dahlin, Steve Gribble, Craig Chambers, David Patterson, David Wetherall, Eric Brewer.

Personally, I've had way more impact (and fun!) building Compute Engine and Kubernetes than I had in academia. If in doubt, try industry for a summer or a year -- nothing we write can replace personal experience.

alex-mohr··on MFA issues lock out Office 365 and Azure users globally
A longer article with more details, including a first attempt to fix that didn't work: https://www.cbronline.com/news/azure-down-office-355-down
alex-mohr··on Google Cloud Platform Finland, No National Connectivity
The cpuid is virtualized and reports virtual hardware so that compatibility and consistency across regions are maintained.

The actual hardware is likely much more recent, and GCE allows you some control over which version is used and reported: https://cloud.google.com/compute/docs/instances/specify-min-...

alex-mohr··on Ask HN: Who is hiring? (December 2017)
Google, Inc. | Software Developers, Security Engineers, Eng Managers | Seattle, WA | ONSITE, full-time | Comp: Google-scale + relocation

Come work for Google's Cloud Platform in Seattle and help us build OSS Kubernetes [0] and our managed service, Google Kubernetes Engine (GKE [1]). I manage our Seattle engineering teams and am hiring software developers, security engineers who love writing code, and engineering managers.

We have three subteams in Seattle, where each subteam has responsibility for making both k8s and GKE great in its area of ownership:

(1) Cluster Lifecyle: Active in sig-cluster-lifecycle and sig-cluster-ops, its mission is to deliver a great cluster administrator experience across the entire lifecycle of a cluster: not only one-off install time, but managing a cluster over years: upgrades, config, machines, repairs, and teardown, for cloud and on-prem environments. We're building a Cluster API to drive all of that.

(2) Security and Auth: Active in sig-auth, its mission is to deliver platform features that enable our users to build secure apps on k8s and GKE (not a reactive vulnerability disclosure nor pure analysis role -- must ship implementation).

(3) GKE API Infrastructure: Drive the infrastructure that lets us offer a great managed as-a-Service product. It's the central core of GKE and plays a role delivering features across multiple subteams.

Ideally you have previous experience with Kubernetes and/or containers -- either building those platforms or running apps on top of them -- and further domain knowledge in one of our subteam's focus areas. You're passionate about delivering great products to end users, and focus on impact over implementation.

Google Seattle also hosts a number of other Google Cloud Platform efforts and many are hiring. Apply at https://careers.google.com/locations/seattle-kirkland/ and mention the particular products you're interested in, or contact me via hn@alexmohr.com.

[0]: https://kubernetes.io/

[1]: https://cloud.google.com/kubernetes-engine/

alex-mohr··on Snap commits $2B over 5 years for Google Cloud infrastructure
Yes, many of Google's technical leads working on Kubernetes and Container Engine are former members of the Borg and Omega teams, so Kubernetes and our hosted version, Container Engine, both benefit from what we learned building those other systems. (I think our 5 most-senior engineers have ~40 years of container management systems experience between them now?)

And it's not just the rather-large core team directly on GKE and k8s, nor the related products like Container Registry [1], Container Builder [2], and Container-Optimized OS [3]. GKE and k8s benefit in other ways too: Google's internal kernel team helps debug customer issues when we trace them to the kernel, and people like Kees Cook are helping with the upstream Kernel Self-Protection Project [4] that make container technology more secure. In addition to that kernel work, Google also has rather-decent security teams and they work with us to improve security in other ways too.

Finally, re: toomuchtodo's question, "Why opt for Google if you're going to use containers in Kubernetes?" Because we hope you find that Container Engine is the best place to run Kubernetes -- and benefit from the other parts of Google Cloud Platform. If you ever find GKE is not that place, and you don't derive value from the rest of GCP, then exactly as toomuchtodo puts it: "You can even move to your own datacenter at some point (relatively) easily."

[1] https://cloud.google.com/container-registry/

[2] https://cloud.google.com/container-builder/docs/

[3] https://cloud.google.com/container-optimized-os/

[4] https://www.linux.com/news/google-developer-kees-cook-detail...

Page 1 of 2Next →