HNHacker News
TopNewBestAskShowJobs

negz

13 karma · joined April 2, 2016

submissionscomments
negz··on I built a fleet-scale inference control plane using Crossplane
I'm the post author, and a long-time Crossplane maintainer. To my knowledge this is the most advanced open source control plane anyone's built with Crossplane. Happy to chat about what worked well and what didn't building this - AMA.
negz··on Scaling Kubernetes to Thousands of CRDs
Why not both? ;) - https://github.com/crossplane-contrib/provider-terraform
negz··on Scaling Kubernetes to Thousands of CRDs
We certainly design Crossplane with the intent that end-users will interact with a "Composite Resource" API, not use MRs directly. Little bit like end-users using a Terraform module curated by their platform team rather than writing resources directly.
negz··on Scaling Kubernetes to Thousands of CRDs
We've thought about a client-side tool to diff desired from most-recently-observed actual state, similar to tf plan. Nothing built yet though.

FWIW though we never automatically delete-and-recreate in Terraform fashion. Our thinking is that once Kubernetes supports it (or once we build an admission control webhook for it) we'd like to explicitly mark immutable fields as such, which would require you to explicitly delete-then-recreate. Little bit less declarative at the expense of being a lot less surprising given the constantly reconciling nature of our system.

negz··on Scaling Kubernetes to Thousands of CRDs
Haha I noticed that also shortly after I made this graphic. I've used it in a few talks and posts now and always worried someone would notice. Good eye.
negz··on Scaling Kubernetes to Thousands of CRDs
Correct. We take the control plane of Kubernetes and extend it so that it can be used to configure/orchestrate anything, not just containers.
negz··on Scaling Kubernetes to Thousands of CRDs
Post author here - thanks for the kind words!

I wouldn't say Crossplane is _supposed_ to be used together with CD systems (that's optional) but that's certainly a common and practical use case.

negz··on Crossplane vs. Terraform
I wasn't familiar with Atlantis - thanks for the pointer!
negz··on Crossplane vs. Terraform
OP here - this is enabled by a mix of tooling consistency (e.g. everything can happen in the same Kubernetes API server) and carefully designing our "managed resources" - the custom resources that represent bits of infra - to be able to reference each other in an eventually consistent way. For example you can create a firewall rule in Crossplane before you create the database instance it applies to. The firewall rule controller will go into backoff until the database is created, at which point everything will fall into place.
negz··on Managing Machines at Spotify
I believe we expect moving to the cloud to be more expensive than running our own DCs as you suggest, but I don't believe that takes into account any 'wasted developer time' you might factor into this.

I believe we started building this platform when AWS was very new, and hadn't seen a compelling reason to transition from it to the cloud until now. There's a couple of posts with more details behind our decision to go to GCP, but primarily it was to leverage their data tooling.

negz··on Managing Machines at Spotify
It's strange (or perhaps rather unfortunate) to me as an SRE at Spotify as well. Helios is in many ways similar to Kube, so it was our hope that eventually it would lead us to scheduling multiple service containers per physical machine. We certainly have the service discovery framework to support that model.

However given Spotify's business position our priority has yet to shift from providing engineers compute capacity as fast as possible to optimising our usage of said compute capacity. It's all now somewhat of a moot point as we move away from our own hardware into Google's cloud.

negz··on Managing Machines at Spotify
FWIW a team of largely men using womens' names for servers always felt a bit icky to me personally for reasons I couldn't quite enunciate. We now use ungendered serial numbers.
negz··on Managing Machines at Spotify
Post author here. A bunch of stuff was glossed over as the post was more focused on the stack's history and evolution than specific technical details.

Ideally we hope to provide some followup posts that go deeper into technical detail about key pieces of the stack (DNS, initramfs framework, job broker, GCP usage, etc).