Kubernetes on Oxide: How customer needs shaped our integrations
oxide.computer
oxide.computer
FYI I've got `karpenter-provider-oxide` on my bingo card...
My colleague demo'd Karpeneter internally. We haven't committed releasing it yet but we're discussing it.
[0] https://github.com/kubernetes/cloud-provider/blob/master/clo...
RE Karpenter, it always seemed like a natural fit for Oxide, even more so than some of the currently implementors. And now you have the expertise in the team, I'd be really interested in the reason if you don't go down that path.
One implementation specific detail which makes Karpenter interesting on Oxide is the tunable CPU and memory parameters rather than strict instance type shapes. The prototype I built for Karpenter on Oxide generates all possible "instance type" combinations, so you can create some really specific nodes to bin-pack pods.
Another interesting area is multiple providers. This is becoming pretty common across public clouds too. A multi-provider Karpenter is something I'm interested in and I know some folks have already been gluing together, but the Karpenter story isn't great on maintaining those since you basically need to compile them together today. CAPIs multi-provider story is a bit cleaner since it only relies on CRDs.
If you have ideas, let's chat in the Kubernetes slack #karpenter-dev.
Autoscaling is a requested feature, and Karpenter fits that shape naturally. The nuance is to decide where something like the Cluster API provider ends and Karpenter begins since there's a bit of overlap in concerns. Specifically, both want to manage Kubernetes nodes but for different reasons. We're discussing it though. We have a Kubernetes watercooler meeting today where we'll likely discuss the comments in this post!
I’m not an oxide customer, just a fan. I have lots of ideas around how something like this should look, being disappointed with the complexity and also shortcomings of products on the market for this stuff, both in the cloud and on prem.
There's an opportunity to more tightly integrate at the network layer but we'll want to get our load balancer released first.
What a flex.
It's gotten to the point I'm starting to actively look for Oxide customers so I can apply with them.
A great example of completely misdirected marketing and/or engineering. Oxide should have developed home microcomputers, and they would have sold like cupcakes, with their ASCII art marketing. Instead they do million dollar "mainframes" that no one (except VCs) wants to buy.
In any event: from the beginning, we have always been targeted at the enterprise buyer who is looking at annual public cloud bills in the tens, hundreds or (in some cases) thousands of millions of dollars per year. We love the enthusiasm that Oxide engenders among the home lab set, but that's not how we have geared the business (for lots of good reasons).
Also, for whatever it's worth: people do want to buy them, it turns out.
Just because a few people on a niche tech site are excited by niche tech stuff doesn’t mean you should focus your entire business on them, and the implication that you’d be successful in marketing a consumer product with ASCII art is mind blowing.
Sadly, I agree.
> They barely want to pay [...]
Not so! I've conservatively dropped $10k on my homelab so far, and have cobbled together a passable VM host/network/home-auto system in a half-rack. I would happily have paid more for an integrated (hypothetical) MicroOx! As fun as it is to play with cage nuts and to crimp cables, I love doinking around with a functional system even more.
There must be literally dozens of us out there...
Lots of people saying "I would love this" is not a profitable market. There are a lot of implicit assumptions in that phrase. Would you buy the lowest level rack for $100k, paying extra for any support? (I don't have any insight into actual pricing, but I know enterprise consumers don't bat an eyelid at prices like that.)
If not then maybe the "I would love this" is not relevant for a real market.
It's so bad I can only guess that it's purposefully so. "sell like cupcakes"????
[0] https://rfd.shared.oxide.computer/rfd/0001
[1] https://oxide-and-friends.transistor.fm/episodes/rfds-the-ba...
And yeah, you also can use react with no client side JS as well.
ClusterAPI never got the love that it should. I spent a good amount of time on it at VMware, as Tanzu heavily leverages it for cluster deployment (mostly CAP-A and CAP-V). It's basically kubeadm + the spirit of Terraform, Kubernetes controller edition. There's lots of great enterprise-ready options for centralized k8s cluster fleet management these days, but it works really well for folks that are all-in on GitOps and such.
I talked with a colleague from Oxide in 2024 about your Kubernetes story and he said back then "not yet but soon-ish". Seems like soon-ish is now :)
We said we'd talk again when that happens but he's since left Oxide. If you (or well...your customers) are interested in a Kubernetes native data platform 100% open source we'd be very happy to talk about how that could work easily. As it's "just" Kubernetes it should be trivial but we'd be happy to test and add you to our list: https://docs.stackable.tech/home/stable/kubernetes/#supporte...
The offer stands. If you're interested, my mail is in my HN profile. https://stackable.tech
On AWS EKS Fargate, each pod runs in its own dedicated VM. With Oxide each k8s node is a VM, so you still need something like Talos.
On networking it looks like it is getting closer. Where you can have external subnet give pod routed IPs without overlay. But the gap, as marked by the article, is also the load balancing.
It would be nice if Kubernetes were a native feature out-of-the-box. Also integrated within the existing user/access control.
I don't necessarily want to match the public cloud experience if there's an opportunity for Oxide to exceed the public cloud experience. Eliminating the overlay is a good example of this. We have customers using external subnets to eliminate the overlay but we haven't integrated that into our controllers yet.
That's what we prototyped before local disk was released and before we started disk hot-plug work.
> Could you attach a single large volume, and do path-based provisioning on that?
Possibly. We'd still want disk hot-plug first. Otherwise, customers would have to create their cluster in a certain shape before using PVCs.
At first look, it feels like oxide is equivalent to proxmox or some virtualization tool, may be it uses qemu stuff underneath.
Just curious.
The reason I am asking this is that, we have lot onprem scenarios in our business. We are tightly coupled with k8s, to solve this we started building an internal project that is kubernetes API compatible [1] but runs containerd or WASM or our platform natively.
Just curious how oxides work in this scenario
The core primitive on Oxide is the instance (virtual machine). We could support some container primitive, but that's a larger product direction discussion. Our host OS is Helios (Illumos) so there are details to iron out there regarding what abstractions we would build and expose to the users. Not impossible but not something we're currently pursuing either given that we have other immediate product asks.
If you're at a scale where compute density, power efficiency, security, and rack-level API management matters then that's where Oxide makes sense for you.
Switching to Linux and KVM would gain the ecosystem benefits but Oxide would lose control of the host OS which can impact our security stance. Better to tightly integrate at this layer and expose the primitives that are needed. It's not impossible, but we're already maintaining the existing Helios host OS. Adding another wouldn't be beneficial.
HBOM
SBOM
Unparalleled support w/ vertical integration
Tons of companies will see you a "cloud". The other parts seem pretty unique.
oxide and coreweave?
Oxide is in-house custom everything. Switches. Racks. Power bus. BMC. Firmware. OS. Software and APIs. Virtualization. Etc.
4+2 Erasure Coding has only half the storage overhead of 3x replication but the same fault tolerance
Also re spdk
https://muratkarslioglu.com/blog/io-uring-spdk-kernel-bypass...
Edit: skimmed it, classic AI "analysis" - totally ignores all the crap io_uring went through
I guess it’s a bunch of computers. It says “AMD” in a bunch of places so, I guess they’re x86. Where are the GPU’s? If it’s for “frontier workloads”, wouldn’t those be sorta important? Is it its own OS? No idea. I see people adjacent to Oxide mention IllumOS sometimes[0], so does that mean it’s a solaris-like OS? Why would I want to run k8s (presumably with linux containers) on a not-linux OS? Or does it virtualize linux instances?
Do actual CTO’s go for this type of marketing, with zero details and nothing but pure fluff about “solutions”?
[0] This of course is from HN comments, I see no mention of any operating system whatsoever on the website, so if it is IllumOS, it’s not like you can easily find that info anywhere.
The host OS is an implementation detail since you, the customer, aren't running workloads directly on the host OS. The VMs running on Oxide run on our host OS, much like other cloud providers.
That rack-scale hardware and software co-design allows Oxide to provide higher CPU and memory density in the same footprint at lower power than competitors. That's super helpful for companies operating at scale where data center power matters. The co-design allows us tackle security problems by eliminating the BIOS, providing attestation from the firmware up to the guest VM, and actually updating the firmware that's on Oxide rather than letting it rot like most customers do today.
This is why when people ask which verticals Oxide targets we kinda say "all of them" because different customers benefit from different Oxide features, but all customers need the core VM, disk, VPC abstraction. Some customers come to us because they don't have an API to manage their on-premises compute and Oxide solves that. Others come to us because they need absolute confidence that there's no malicious firmware running in their compute stack. Others come to us because they are tired of paying exorbitant amounts of recurring money just to run on-premises compute.
Our job is to make Oxide an appealing on-premises computing platform for customers to run their public cloud provider workloads and on-premises workloads without it feeling like it's an entirely different platform than you're used to.
Re-assessing my criticism, I think I’m mostly complaining that I can’t figure out what software it’s running. But I guess the target customers don’t care, since it’s an implementation detail. I’ve only ever heard “oxide” in the past in reference to the OS they were (are?) developing, which a web search says is called “hubris”, so when I see an article like OP I’m wondering “are they running k8s on top of hubris? Wow!” But I’m imagining that’s very much not the case.
In the comments in this very section an oxide employee mentions they’re using Illumos, but I don’t see that anywhere on their website either.
But you’re probably right in that the target customer doesn’t give a shit what OS it runs, so long as they can get instances deployed (although I would question why, if you’re going to run k8s anyway, you don’t just run it on bare metal and skip the hypervisor, but that’s my bias showing up as someone who writes bare metal controller software.)
But to your point, I don't think I ever said machines don't need a management control plane (I develop a bare metal management control plane for $DAYJOB, I'm fully aware of what it entails!) I'm saying that you don't need to insert a VM layer between the bare metal and containers: Having k8s run on the bare metal OS would be ideal IMO.
But since Oxide seems to be Illumos-based, that's basically not possible... At least not for containers as most people know them (ie. with linux-based docker images.) Heck, even if running a VM layer between the OS and the containers, it looks like running your own hypervisor prevents you from (currently) having nested virtualization, which makes certain k8s workloads a lot harder.
Interestingly he has mentioned that in the podcast and basically said, the industry has standardized on VM and in the kind of environments they run, there just isn't a good way around it.
But yeah, a world where Illumos based bare-metal containers had become the standard would be a nicer. But its not where the market is.
I get what you're saying though, and you're right that certain workloads like Kubernetes on the metal that can take advantage of hardware wouldn't be a good fit for Oxide today. We could decide to change that in the future, but I don't see that as a priority for Oxide given all else we have to build first.