If you have need for speed, a team that knows the space, and crucially a leader who can be trusted to depart from the usual process when that tradeoff better meets business needs, it can work really well. But also comes with increased risk.
142 karma · joined June 22, 2015
Previously, I was the overall technical lead for Google Compute Engine's initial public launch; also lead the design and launch of its second-gen VM instance manager subsystem.
And before all of that, ran a research lab at Stony Brook University [1], but I'm a reformed academic and prefer building real systems now.
[1]: http://alexmohr.com/papers/
If you have need for speed, a team that knows the space, and crucially a leader who can be trusted to depart from the usual process when that tradeoff better meets business needs, it can work really well. But also comes with increased risk.
- two competing orgs via Brain and DeepMind.
- members of those orgs were promoted based on ...? Whatever it was, something not developing consumer or enterprise products, and definitely not for cloud.
- Nvidia is a Very Big Market Cap company based on selling AI accelerators. Google sells USB Coral sticks. And rents accelerators via Cloud. But somehow those are not valued at Very Big Market Cap.
Of course, they're fixing some of those problems: brain and DeepMind merged and Gemini 2.5 pro is a very credible frontier model. But it's also a cautionary tale about unfettered research focus insufficiently grounded in customer focus.Link to the patch fixing it: https://github.com/kubernetes/kubernetes/commit/7fef0a4f6a44...
Of course, we'd already fixed other issues like Kubelet listening on a secondary debug port with no authentication. Those problems stemmed from its origins as a make-it-possible hacker project and it took a while to pivot it to something usable in an enterprise.
if (request.authenticationData) {
ok := validate(etc);
if (!ok) {
return authenticationFailure;
}
}
Turns out the same meme spans decades.In terms of impact or business case, I'm missing what the end goal for the company or execs involved is. It's not re-writing user-space components of AOSP, because that's all Java or Kotlin. Maybe it's a super-longterm super-expensive effort to replace Linux underlying Android with Fuchia? Or for ChromeOS? Again, seems like a weird motivation to justify such a huge investment in both the team building it and a later migration effort to use it. But what else?
From a technical perspective, App Engine and Compute Engine were built on top of internal infrastructure (borg), but did not expose borg directly. And there were a number of interesting mismatches between the semantics that customers expected of VMs and what borg offered to its containers that eventually resulted in dedicated borg clusters with different configs for cloud. And some retrospectives on whether building on borg was a better option than going bare metal directly.
Org-wise, the App Engine team was first and not part of the internal-focused Technical Infrastructure teams. GCS came next, and it too was not part of the canonical storage org. Then GCE, which was only possible because it was either written off or at least tolerated as an experiment by most, with a few key people providing behind-the-scenes support to make it happen -- especially in networking. It likely also helped that GAE was in SF and the rest of GCP in Seattle/Kirkland initially, so geo provided some insulation too.
The dominant perspective internally was that Google's technical infrastructure was its secret sauce, so why would they give it away to others? It took a long time to change that.
[Disclosure/source: I was on GCE and helped get it launched.]
But it does seem the capacity of a hybrid system of Netflix servers plus P2P would be strictly greater than either alone? It's not an XOR.
And note that in this case of "live" streaming, it still has a few seconds of buffer, which gives a bandwidth-delay product of a few MB. That's plenty to have non-stale blocks and do torrent-style sharing.
The problems with using it as part of a distributed service have more to do with asymmetric connections: using all of the limited upload bandwidth causes downloads to slow. Along with firewalls.
But the biggest issue: privacy. If I'm part of the swarm, maybe that means I'm watching it?
[1]: Chainsaw: P2P streaming without trees, https://link.springer.com/chapter/10.1007/11558989_12
As a construction kit, it has value for people who want to make protocols where they'll control both ends, but don't have to re-implement basic table stakes.
One of my students at the time, Mahadev Konar, ended up writing a paper "Ring-like DHTs and the Postage Stamp Problem" [1] that shows how you can use solutions to the postage stamp problem (aka denomination-choosing problem) as a way to structure the finger pointers in Chord. And went on to co-found Hortonworks.
Sometimes random things on HN end up having implications in other areas!
[1]: https://alexmohr.com/papers/dht-postage-stamp-podc2005-exten...
On one shard you can't use more than base CPU, so there's no advantage to a subscription there.
Other optional shards are almost entirely people who subscribe. Do that too if you decide to, or ignore them.
You write the code for each of your units, either natively in Javascript or Typescript, or via WASM you can run Rust, Python, etc. You use a private server or join a shared MMO world. There's a free sim [1] to try out the basics, though the actual game has much more depth. And an active Discord for help [2].
There's also a variant Screeps: Arena [3] that focuses on 1:1 PVP battles with ranked ladders if you prefer short-lived matches to a long-running world.
[0] https://store.steampowered.com/app/464350/Screeps_World/
[1] https://screeps.com/a/#!/sim
[2] https://discord.com/invite/RjSS5fQuFx
[3] https://store.steampowered.com/app/1137320/Screeps_Arena/
[0]: https://chrislema.com/a-done-done-culture-habit-one/
[1]: https://www.amazon.com/Creating-Done-Culture-Habits-Performe...
Maybe if they have other goals that are tradeoffs vs. simplicity, then it's more understandable?
What if another goal is allowing enterprise customers to recreate a virtual enterprise network? Or a virtual data center network? Those are much more complex than a client TCP stack.
The defaults are simple for simple uses. And yes, for those complex cases, you'd need AWS specific product knowledge, but most of the underlying concepts are shared in common with other clouds and on prem networks. Like learning your Nth programming language.
Reconsider that. You can try to add everyone you remember from previous companies or school, and you'll be more available to recruiters. "Luck surface area."
In Microsoft's case, the remediation is not to put in place higher level systems to safely accomplish the goal of the command. Instead:
- "We have blocked highly impactful commands from getting executed on the devices (Completed)"
- "We will require all command execution on the devices to follow safe change guidelines (Estimated completion: February 2023)"
Requiring commands to follow guidelines sounds suspiciously like they're requiring network ops not to break things.
Maybe the fear of actually doing something outweighs the costs of not doing it under the guise of making it better? Maybe something else?
I found Bezos's thoughts on decision making useful: understand if the decision is a two way door, and if so, move forward with 70% of the data you wish you had. Or 70% of the product, implementation, whatever.
See e.g. https://www.aboutamazon.com/news/company-news/2016-letter-to...
1: http://ecolo.org/documents/documents_in_english/Rickover.pdf
That direct two-way communication both unblocked adoption for users and was a source of feedback for the devs.
Only battery pack with 2 usb-c seems to be the ZMI Ambi: https://www.amazon.com/ZMI-PowerPack-Ambi-USB-C-Power/dp/B07....
I see immutable, but also upgradable? Is that via in-place upgrades or do upgrades require a reboot?
Example: severe bug or vulnerability in kubelet or containerd/docker. Can I use the API to roll out a fix to existing nodes such that running workloads have no disruption?
And there are of course a number of other former faculty there too, but none that I know of who've blogged as much as Matt. In the systems space, off the top of my head: Amin Vahdat, Mike Dahlin, Steve Gribble, Craig Chambers, David Patterson, David Wetherall, Eric Brewer.
Personally, I've had way more impact (and fun!) building Compute Engine and Kubernetes than I had in academia. If in doubt, try industry for a summer or a year -- nothing we write can replace personal experience.
The actual hardware is likely much more recent, and GCE allows you some control over which version is used and reported: https://cloud.google.com/compute/docs/instances/specify-min-...
Come work for Google's Cloud Platform in Seattle and help us build OSS Kubernetes [0] and our managed service, Google Kubernetes Engine (GKE [1]). I manage our Seattle engineering teams and am hiring software developers, security engineers who love writing code, and engineering managers.
We have three subteams in Seattle, where each subteam has responsibility for making both k8s and GKE great in its area of ownership:
(1) Cluster Lifecyle: Active in sig-cluster-lifecycle and sig-cluster-ops, its mission is to deliver a great cluster administrator experience across the entire lifecycle of a cluster: not only one-off install time, but managing a cluster over years: upgrades, config, machines, repairs, and teardown, for cloud and on-prem environments. We're building a Cluster API to drive all of that.
(2) Security and Auth: Active in sig-auth, its mission is to deliver platform features that enable our users to build secure apps on k8s and GKE (not a reactive vulnerability disclosure nor pure analysis role -- must ship implementation).
(3) GKE API Infrastructure: Drive the infrastructure that lets us offer a great managed as-a-Service product. It's the central core of GKE and plays a role delivering features across multiple subteams.
Ideally you have previous experience with Kubernetes and/or containers -- either building those platforms or running apps on top of them -- and further domain knowledge in one of our subteam's focus areas. You're passionate about delivering great products to end users, and focus on impact over implementation.
Google Seattle also hosts a number of other Google Cloud Platform efforts and many are hiring. Apply at https://careers.google.com/locations/seattle-kirkland/ and mention the particular products you're interested in, or contact me via hn@alexmohr.com.
And it's not just the rather-large core team directly on GKE and k8s, nor the related products like Container Registry [1], Container Builder [2], and Container-Optimized OS [3]. GKE and k8s benefit in other ways too: Google's internal kernel team helps debug customer issues when we trace them to the kernel, and people like Kees Cook are helping with the upstream Kernel Self-Protection Project [4] that make container technology more secure. In addition to that kernel work, Google also has rather-decent security teams and they work with us to improve security in other ways too.
Finally, re: toomuchtodo's question, "Why opt for Google if you're going to use containers in Kubernetes?" Because we hope you find that Container Engine is the best place to run Kubernetes -- and benefit from the other parts of Google Cloud Platform. If you ever find GKE is not that place, and you don't derive value from the rest of GCP, then exactly as toomuchtodo puts it: "You can even move to your own datacenter at some point (relatively) easily."
[1] https://cloud.google.com/container-registry/
[2] https://cloud.google.com/container-builder/docs/
[3] https://cloud.google.com/container-optimized-os/
[4] https://www.linux.com/news/google-developer-kees-cook-detail...