The first Oxide rack being prepared for customer shipment
hachyderm.io
hachyderm.io
[0] When we started: https://news.ycombinator.com/item?id=21682360
[1] On the Changelog podcast: https://news.ycombinator.com/item?id=32037207
[2] On our embedded Rust OS, Hubris: https://news.ycombinator.com/item?id=29468969
[3] On running Hubris on the PineTime: https://news.ycombinator.com/item?id=30828884
[4] On compliance: https://news.ycombinator.com/item?id=34730337
[5] On our approach to rack-scale networking: https://news.ycombinator.com/item?id=34976444
[6] On our de novo hypervisor, Propolis: https://news.ycombinator.com/item?id=30671447
[7] On our boot model (and the elimination of the BIOS): https://news.ycombinator.com/item?id=33145411
-George Bernard Shaw
Really curious to see if in 2023, Engineered Systems still have a market in this world of commodity cloud hardware.
Also, congrats.
Disclaimer: I am a fan, not affiliated.
Though the biggest advantage to Rust-on-bare-metal would have been running all the code in kernel space, which it seems Tasks don't allow for (in order to provide robustness at the OS level).
Q: Are you shipping Milan racks and if so what does the moving to Genoa entail for Oxide?
For funsies, it’s neat to look back at the original announcement and discussion from four years ago: https://news.ycombinator.com/item?id=21682360
What are the problems people are having with existing systems (like Vxrail), and how does Oxide fix those? What stories are you hearing?
For some organizations cattle-like pizza boxes or chassis with blade systems are still not cattle-like enough. By making the management unit the entire rack you can reduce overhead (at least compared to a rack of individual servers, even if they are treated like cattle).
There are vendors that will drop ship entire (multiple) racks pre-wired and pre-configured for various scenarios (especially HPC): just provide power, (water) cooling, and a network uplink.
My sibling commentor just left a great answer to your second question, so I'll leave that one there :)
From your referenced comment:
> The rack isn't built in such a way that you can just pull out a sled and shove it into another rack; the whole thing is built in a cohesive way.
> other vendors will sell you a rack, but it's made up of 1U or 2U servers, not designed as a cohesive whole, but as a collection of individual parts
What I'm curious about is how analogous or different is this cohesiveness to the days where vendors built the complete system? Is that the main selling point or there are nuances to it?
> What I'm curious about is how analogous or different is this cohesiveness to the days where vendors built the complete system?
To be honest, that was before my personal time. I was a kid in that era, using ultra hand-me-down hardware. I'd speculate that one large difference is that hardware was much, much simpler back then.
- Oxide is entirely open where Apple is extraordinarily secretive: all of our software (including the software at the lowest layers of the stack!) is open source, and we encourage Oxide employees to talk publicly about their work.
- Oxide is taking a systems approach where Sun sold silo'd components: I have written about this before[0], but Sun struggled to really build whole systems. For Oxide, we have made an entire rack-level system that includes both hardware and software: the abstraction to the operator is as turn-key elastic infrastructure rather than as a kit car.
We have tried to take the best of both companies (and for that matter, the best of all the companies we have worked) to deliver what customers want: a holistic system to deploy and operate on-premises infrastructure.
Outposts, too, started with a full-rack configuration only, but they eventually introduced an individual server configuration as well. It'll be interesting to see if Oxide eventually decides to serve the smaller market that doesn't have the scale for whole-rack-at-a-time.
Once installed, you plug in the network connection(s) and add power, then boot up. Then you can add your VMs and start running them.
When you cut out all that fat, you can fit a lot more muscle in its place. Then you can arrange everything on a rack scale for airflow, power, redundancy, etc.
The obvious downside is that you are locked into a single supplier for everything in the rack. Somebody please add a link here for Oxide's strategy to minimize this risk for their customers. (And feel free to correct any mistaken over-simplification.) Overall, they have seemed incredibly open to sharing in a way that allows others to play. That might just be because they would die if they didn't, but I think a lot of people are willing to accept that they are pretty all-around awesome folks trying to make the world better. And making money in the process is not a bad gig. We'll see which other vendors jump into the pool.
I wish them great success. But I'm wondering why they haven't coordinated my shipping address for their first shipment. First one's free, right?
This is not really true. Data center servers are highly optimized for density already. Like tens of thousands of hours on airflow CFD, tweaking cases and fans and internal layout, heatsink design, etc. chasing the 0.1%s. There are many significant tradeoffs to be made, but there is not a large factor improvement just sitting on the shelf due to a bunch of unused legacy leftovers.
The attack surface it leaves open is a different story.
[0] https://www.osfc.io/2022/talks/i-have-come-to-bury-the-bios-...
More to the point with Oxide, there's layer upon layer of software cruft as well: various levels of proprietary firmware blobs driving the boot process, and then the management stack that underlies a commodity PC platform OS.
All of that is gone, replaced with a bottom-up redesign, of both the hardware and the software, with the specific purpose of running large scale modern application loads.
This goes well beyond a blade server, and is a completely different beast than a rack of DL360s.
Proof of the pudding will be in the eating I guess. Does Oxide talk about performance advantages at all, or have numbers?
If it didn't matter, then there'd be no interest in blades, or OCP, and all the cloud vendors would use standard rackmount whitebox servers. At the scale of one rack, the difference is fairly small. But once you're building out cages full of racks, maybe it matters more?
And it's not just power and space (although they both matter), it's the attack surface, and the ability to actually have your hardware do what you want, and the benefits of a stack that's built for purpose, not cobbled together out of off-the-shelf parts.
It's not for everyone, to be sure, but I suspect that for its target market, it'll be a very successful product/family. As you say, we will see. I wish them well.
Not with the current state of the art.
> If it didn't matter, then there'd be no interest in blades, or OCP, and all the cloud vendors would use standard rackmount whitebox servers.
I'm talking about it mattering from the starting point of blades, OCP, "cloud" systems!
The thread is about where the oxide niche is and what advantages it has over competition. The idea there is huge amount to be won on southbridge chipset and IO ports in large scale systems is simply not true. Quite amazing that people who don't understand this are posting in this thread as though they are experts in the matter.
> Not with the current state of the art.
Perhaps you could educate me: what's the state of the art that makes watts free?
>> If it didn't matter, then there'd be no interest in blades, or OCP, and all the cloud vendors would use standard rackmount whitebox servers.
> I'm talking about it mattering from the starting point of blades, OCP, "cloud" systems!
> The thread is about where the oxide niche is and what advantages it has over competition. The idea there is huge amount to be won on southbridge chipset and IO ports in large scale systems is simply not true.
I don't think anyone has asserted that there's "huge amount"s to be won from the simplified hardware? It's my view that there is a small amount to be won (more if you're coming from DL380s, and a bit less if you're replacing blades, for instance). But those small amounts do add up, and contribute to the value proposition. I think there are other parts of the Oxide product that are more compelling, but that's the point: it's the total package that you're buying, not just the missing VGA socket on the front panel.
Imagine that I'm an enterprise computing user. I currently have a few small (wrt cloud providers) datacenters, stuffed with racks of any tier 1 vendor's rack mount, blade or even OCP systems. And they're running some combination of plain VMs, perhaps some orchestration platform (k8s, or Mesos, or whatever). And it's time for my 3/5/8-yearly refresh of a bunch of those systems.
A unified, coherent, targeted product (such as Oxide) might be a compelling offer. It promises to deliver cloud-type hardware and software efficiencies at a smaller scale. It gives me the savings of not renting from AWS/Azure/etc, while not dealing with the hodge-podge of the standard PC hardware and software stack.
> Quite amazing that people who don't understand this are posting in this thread as though they are experts in the matter.
Welcome to the Internet?
Doesn't seem to be taking. The issue is that there are not a couple of watts there. There are a couple here, and that's all.
> I don't think anyone has asserted that there's "huge amount"s to be won from the simplified hardware?
They did. In this thread you are participating in even. Perhaps you didn't really read the post of mine that you first replied to, or its context?
So, this isn't really chasing the 0.1% -- it's chasing much bigger wins, and it's doing it by changing the power distribution (shelf of rectifiers to a 54V DC busbar vs. redundant AC power supplies in every 1U/2U), changing the cooling (the larger fans), changing the networking (we have a blindmated cabled backplane, obviating the need for operator cabling), etc. etc. These are not just multipliers on efficiency, but also on manageability: once physically installed, our rack is designed to be provisioning VMs orders of magnitude faster than the traditional manual rack/stack/cable/SW install.
I was replying more to the idea there's all this costly legacy IO.
I guess you aren't the first to try different geometries or power delivery or cabling either. There's been lots of little opencompute-type efforts and startups come and go. I'm skeptical there's a lot in it in a significant niche that does not already do these things, but you don't need to convince me. Although if you did want to you could show some comparative density numbers. What can you fit, 2048 cores, 32TB, and push 15kW through a rack?
And not one in particular, there's just a bunch that have sprung up around OCP over the past decade. None that I'm aware of that are doing everything that Oxide does, but we were talking more about the mechanical, electrical, and cooling side of it there -- they do seem to do okay with power density.
I thought you were going for cloud DCs rather than enterprise. Seems like a big uphill battle to get software certified to run on your platform. Are any of the major Linux distros or Windows certified to run on your hypervisor platform? Any ISVs?
> The question on density is: how can you get the most compute elements into the space/power that you've got, and cramming towards highest possible density (i.e., 1U) actually decreases density because you spend so much of that power budget cranking fans and ramming air.
The question really is how much compute power, and electrical/thermal power ~= compute power. Sure you could fit more CPUs and run them at a lower frequency or duty cycle.
> And the challenge with OCP is: those systems aren't for sale (we tried to buy one!) and even if you got one, it has no software.
OCP is a set of standards. They certainly are systems sold. I guess the nature of the beast is that buyers of one probably don't get taken very seriously, particularly not a competitor.
Yes, and using larger fans decreases the proportional there. Centralized rectifiers reduces the proportional. Google can make it all the way down to 10% overhead power. That's the point they are making.
Now that racks are shipping it'd be awesome to see a top-to-bottom look at the hardware and software. They've given a lot of behind the scenes peeks at what they're doing via the Oxide and Friends podcast, but as far as I'm aware there is no public information on what it all looks like together.
* You buy your own servers instead of renting, which is what most people are doing now. They argue there's a case for this, but it seems like a shrinking market. Everything has gone cloud.
* Even if there are lots of people who want to leave the cloud, all their data is there. That's how they get you -- it costs nothing to bring data in and a lot to transfer it out. So high cost to switch.
* AWS and others provide tons of other services in their clouds, which if you depend on you'll have to build out on top of Oxide. So even higher cost to switch.
* Even though you bought your own servers, you still have to run everything inside VMs, which introduce the sort of issues you would hope to avoid by buying your own servers! Why is this? Because they're building everything on Illumos (Solaris) which is for all practical purposes is dead outside Oxide and delivering questionable value here.
* Based on blogs/twitter/mastodon they have put a lot of effort into perfecting these weird EE side quests, but they're not making real new hardware (no new CPU, no new fabric, etc). I am skeptical any customers will notice or care and would have not noticed had they used off the shelf hardware/power setups.
So you have to be this ultra-bizarre customer, somebody who wants their own servers, but doesn't mind VMs, doesn't need to migrate out of the cloud but wants this instead of whatever hardware they manage themselves now, who will buy a rack at a time, who doesn't need any custom hardware, and is willing to put up with whatever off-the-beaten path difficulties are going to occur because of the custom stuff they've done that's AFAICT is very low value for the customer. Who is this? Even the poster child for needing on prem, the CIA is on AWS now.
I don't get it, it just seems like a bunch of geeks playing with VC money?
This strikes me as at least as much of a leap as "everything has gone cloud."
Containers are... kind of a thing. And while there certainly are use cases for VMs over containers, they're comparatively niche.
This product seems as if it'd be a better fit for every real use case I've ever seen if it were a prebuilt kubernetes cluster rather than a prebuilt VM hypervisor.
While true, they also aren't buying $500k+ rack. They're running vmware on a much smaller scale.
A number of customers I've worked with in Biotech, Telecom, Retail, Defense, and Financial sectors have either completed their cloud transformation or have ongoing 3-5 year timelines to make such a move.
On the mid-market side, I've been seeing organizations increasingly offload the entire IT operation to MSPs instead of managing their own ESXi server. The MSP might often be using VMWare, but a number are also transitioning away from that.
I think it's premature to say container and cloud first deployments are the norm, but the change and transition is rapidly occurring.
A lot of this is highly dependent on the kind of app you are migrating as well as the scale needs.
IT departments typically love VMs (and vmware)- AD machines are most often hosted on VMs on VMWare.
We sell a product that used to be delivered as a rack mountable server, but now I think nearly all our customers want a VM instead.
I don’t remember utilization exactly, but I think most of the blades were reasonably committed. That said, when I left they were moving from a hybrid model to entire cloud, so apparently the cost:flexibility ratio wasn’t to peoples’ liking.
https://world.hey.com/dhh/we-have-left-the-cloud-251760fb
It's nice to have options. Cloud good. Self-hosting good. Middle options good.
Double that for companies that were loud about their cloud usage, usually that came with some usage discounts on the side- on-prem does not, and it also looks bad to pivot anyway; so why be loud?
This way, you can run your big static workloads on-prem to save money, and run your fiddly dynamic workloads in AWS and continue to use S3.
That said, if a hybrid cloud architecture is your plan and you desire a managed rack-at-a-time experience, AWS Outposts would seem to be the safer pick. They've been shipping for years and they have public pricing that you can look at. I'm not sure that Oxide specifically has an opening for customers who want to keep their cloud. I wish them luck.
This is an amazing plus in my eyes. Solaris systems are amazing.
Most every software engineer has worked in an org where spending six figures per month on AWS or GCP is totally normal and acceptable because the alternative, buying hardware, is this awful scary unknown that could be cheaper but could also blow the entire company up. If oxide can solve that, suddenly on-premise becomes much more attractive.
Yes, people are hooked on the cloud, but not because of data transfer… because it’s easy.
* well, the first few deployments might not be easy but that’s true of anything new.
Also, what this person said :) https://news.ycombinator.com/item?id=36552245
[0] https://oxide-and-friends.transistor.fm/episodes/tales-from-...
This is very much not true and seems to be a result of people in the valley thinking the rest of the world operates like the valley. In the rest of the world I've found mature businesses that bought into the "cloud is the best" quickly started doing the math on their billing rate and realized there is a VERY small subset of their business that has any reason to be in the cloud. Actually one of the very best use-cases of public cloud I've seen is a finance firm that sticks new products into the cloud until they hit maturity so they can properly right-size the on-prem permanent home for them. And if those products never take off, they just move on to the next one. They're willing to pay a premium for 12-18 months because they can justify it financially.
>* Even if there are lots of people who want to leave the cloud, all their data is there. That's how they get you -- it costs nothing to bring data in and a lot to transfer it out. So high cost to switch.
And yet company's do it all the time. I think you'll again find mature fortune 500s can do the math on the exit cost vs. staying cost and quickly justify leaving in a reasonable time window.
>* AWS and others provide tons of other services in their clouds, which if you depend on you'll have to build out on top of Oxide. So even higher cost to switch.
And as you've seen plenty of people here point out: most of those services tend to be overrated. OK, so you've got database as a service: except now you can't actually tune it to your specific workload. And $/query, even ignoring performance, is astronomically higher than building your own and paying a DBA to manage it unless you're a 2-man startup.
>* Even though you bought your own servers, you still have to run everything inside VMs, which introduce the sort of issues you would hope to avoid by buying your own servers! Why is this? Because they're building everything on Illumos (Solaris) which is for all practical purposes is dead outside Oxide and delivering questionable value here.
I don't know of a single enterprise that has run anything BUT VMs for the last decade. Other than Mainframe (which you can argue is actually just VMs in a different name), and some HFT-type applications that need the lowest possible latency at all costs, it's all virtualized. As for Illumos: why do you care? Oxide is supporting and maintaining it as an appliance. Tape has been "dead" for 2 decades. FreeBSD has been "dead" since the early 2000s. It's only dead for people that don't deal with enterprise IT.
>* Based on blogs/twitter/mastodon they have put a lot of effort into perfecting these weird EE side quests, but they're not making real new hardware (no new CPU, no new fabric, etc). I am skeptical any customers will notice or care and would have not noticed had they used off the shelf hardware/power setups.
I have no doubt they've done their research, and I can tell you from my industry experience there is a large cross-section of people who want an easy button. There's a reason why companys like Nutanix exist and have the market cap they do - but they could never actually get the whole way there because they got wrapped up in the "everything is software defined!!!". Which works really well until you realize that you're left to your own devices on networking.
>So you have to be this ultra-bizarre customer, somebody who wants their own servers, but doesn't mind VMs, doesn't need to migrate out of the cloud but wants this instead of whatever hardware they manage themselves now, who will buy a rack at a time, who doesn't need any custom hardware, and is willing to put up with whatever off-the-beaten path difficulties are going to occur because of the custom stuff they've done that's AFAICT is very low value for the customer. Who is this? Even the poster child for needing on prem, the CIA is on AWS now.
I mean no disrespect but I get the impression you haven't ever worked with a fortune 500 that's outside of the valley. This is EXACTLY what they all want. They aren't going to run their entire datacenter on this, but when the datacenter is measured in hundreds to thousands of servers, they've got plenty of workloads that it's a perfect fit for
Long story short, despite pressure to move from our parent company, there was literally no dimension - services, finance, performance, support - on which AWS was better for our workload. Our DC capex+opex was less than a quarter of the parent company’s opex spend on AWS, despite our workload being much greater - and we needed fewer staff to manage it.
So I for one am really happy to see Oxide producing an on premise private cloud solution like this.
Add in what appears to be, 2x 100Gb Ethernet switches, the 100Gbps NICs, cabling, other rack hardware; service processors that allow you to control the servers, whatever amount of NVME drives, etc. Assembly, burn in, test etc.
My guess is that a base config would be somewhere between $400K to $500K but could very definitely go up from there for a completely "loaded" config.
Most hardware companies have terrible software. If Oxide can handle manufacturing and logistics, then they'll be huge in about 10 years.
Would love to know what kind of uses this is being put to. In this age, everyone only talks about the cloud, with its roach motel model and all.
There is a minor trend of established companies reconsidering cloud because it turns out it doesn’t really save money if your work load is well understood and not highly elastic. In fact cloud is often far more expensive.
However if they’re aiming at the companies parachuting out of the cloud back to data centers and on prem then it makes a lot of sense.
It’s possible that the price comparison is not with comparable computing devices, but simply with the 9 cents per gigabyte egress fee from the major clouds.
If I was marketing director at Oxide I’d focus all messaging on “9 cents”.
Where is this magical company that needs exactly one rack of exactly one type of server? The vast majority of companies needing this much compute will also be interested in storage servers, servers filled with GPUs, special high-RAM nodes, etc. And at that point you'll also be using some kind of router for proper connectivity.
Why bother going for a proprietary solution from an unproven company for the regular compute nodes, and forcing yourself to overcommit by buying it per rack? Why not just get them from proven vendors?
They take the idea of "hyperconvergence" - building a software platform that's tightly integrated and abstracts away the Hard Parts of building a big virtualized compute infrastructure, and "hyperscaling" - building a hardware platform that's more than the sum of its parts, thanks to the idea of being able to design cooling, power, etc. at a rack scale rather than a computer scale. Then they combine these into a compute-as-a-unit product.
I, too, am a bit skeptical. I think that they will absolutely have the "hyperconvergence" side nailed given their background and goals, but selling an entire rack-at-a-time solution at the same time will be hard. But I have high hopes for them as it seems like a very interesting idea.
Is there anything in the works for using some of this in a homelab? Mainframes (unofficially) have Hercules, be good to see something similar for folks who want to experiment.
So this has to be targeted more at F500, or maybe even to colos?
I'd love to have a rack or two some day!
Still think they’ll succeed big but I don’t think they’ve fully dialed in what is important to people who may be able to pull the trigger on a decision like this.
possibly because I am in Europe and they want to focus on the NA market, not sure.
[0] https://oxide-and-friends.transistor.fm/episodes/the-power-o...
If I were to do a bare metal deployment I'd look at kubernetes + kubevirt + a software defined storage provider. Then you get a common orchestration layer for VMs, containers, or other distribute workloads (ML training/inference) but don't need to pay the Vmware Tax and you'd be using a common enough primitive that you can move workloads around to 'burst' to the cheapest cloud as needed.
I guess to be truly turn-key it would need to come with a special model or two adept at training and commanding an army of LLM agents.
Any plans for upgrades beyond single rack deployments and newer gen hardware? And what is the sell case over kube, mesos, or VMware, especially since the hardware and software appear tightly coupled?
Somewhat related if you are interested on companies which heavily uses on-Prem check out
None of this is exposed directly to the customer, though. You run VMs with whatever you want.
However. The kind of customer who spends this type of money can be conservative. They already have to go with on an unknown vendor, and rely on unknown hardware. Then they end up with a hypervisor virtually no one else in the same market segment uses.
Would you say that KVM or ESXi would be an easier or harder sell here?
Innovation budget can be a useful concept. And I'm afraid it's being stretched a lot.
Thanks!
> However.
Yes. These things are challenges, but that doesn't mean they're insurmountable. I am confident that we can overcome them.
> Would you say that KVM or ESXi would be an easier or harder sell here?
I am not in enterprise sales, so I couldn't tell you :) Oxide has those people, I'm not just one of them.
In addition, app teams (i.e. the folks who deliver revenue to the business) are more productive on public cloud. They don't file tickets, they make VMs. Oxide is intended to bring the hardware, software, and interface innovations of the cloud into enterprise data centers with greater reliability and lower TCO.
What about security updates and bug fixes? No platform is perfect, after all...
Or what about if (god forbid), you go out of business 5 years down the line... would the hardware be repurposeable? Or would it just become a large paperweight?
Security is an important part of the product. You'll get updates.
> Or what about if (god forbid), you go out of business 5 years down the line
All of the software, to the degree that we are able to, is open source. This is important precisely because you own the hardware; you as a customer deserve to know what we put on it.
The mainframe is optimized for reliability and compatibility foremost and does things that are quite unlike most other classes of machines even "high end servers", including RAIM (RAID for memory), and the ability to stop a failed processor and move its checkpointed state to another processor and resume it transparently from software point of view, and generally employing quite hardened circuits and strong error detection and correction throughout the system.
There are some guesses at $500k for this Oxide rack, not sure if that even gets your phone call returned for a low end mainframe with much less compute power and memory installed. High end configurations rumored to be many millions.
This thing is more a competitor to "scale-out" / "cloud" / "webscale" / etc., at least on the hardware front (they seem to do a lot more on software/firmware side than typical such hardware vendors).
Their own OS: Check Big tall box: Check Expensive: Check (well no price list i have seen so far) Proprietary hardware at least interfaces. (?) upgrades must be bought from the vendor (?)
Right to repair?
Like the Tandem Nonstop series. (Now part of Hewlett Packard Enterprise) I had a contract with a 9/11 call handler for a while. They had one of them (or possibly more). On display of sorts through a window of bulletproof glass.
The IBM Z series does not allow hot swapping the CPU and I dont believe it ever did. The parts on the Z Series are super easy to replace. All in nice components sitting on a rails to yank in or out.