Moore's Scofflaws
oxide.computer
oxide.computer
One day I would love to get all of their OSS up and running locally. Truly, why not try to run your own private cloud? How many old laptops, desktops, PIs, etc must we all have lying around?
Because until you're big enough to have local weather systems evolving inside your offices, that's called 'servers'. I'm not saying not to do it, it's great, and cost effective, and much safer, but calling a small-to-mid business's servers 'cloud' is like saying I have a 'private Uber' in my garage that I can drive myself.
I don't know what old laptops and PIs have to do with anything though, one $600 desktop is going to dwarf all of them combined in performance. There is no reason for anyone to do that unless they are a kid with only an allowance for income.
My first job ran their own datacenter. We started with a rack of about 50 nodes and some odd computers here and there. By the time I left we had a rack of 250 nodes and we had thrown away all individual servers (oh Solaris I won't miss you).
> Truly, why not try to run your own private cloud? How many old laptops, desktops, PIs, etc must we all have lying around?
Although VMs and Containers have done a world of difference on running your own datacenter, there are still reasons why just use the hardware that you have lying around for a business is a bad idea.
First, random laptops and desktops don't have a good ratio of price, performance, and energy consumption. Second, despite virtualization, all this heterogeneity in hardware and software is bound to cause an ongoing maintenance headache. Third, I hope your business doesn't have a sudden spike of demand, or you'll be struggling to find hardware to support it. Fourth, say good bye to quick iteration and experimentation when you have huge lead times and tons of capex.
Bottom line, cool for a hobbyist project, not cool for starting a business (and a maybe for a well establish and mature business).
But you can't buy integrated like this without software support.
So when 0xide have the (nice) problem of topping out whatever market share they can win, their shareholders will want to extract more revenue from existing customers. That can only come in the form of licensing updates (or selling new sleds) according to _how_ they're used: by core, by power, AI model, number of customers, or whatever else.
Either 0xide are such a roaring success that Amazon reduce their cloud prices to compete. Or they save their select customers so much money on the cloud that they can start raising the rent.
(not a negative judgment by the way, customers win either way. Just putting this kind of marketing in context and a reminder that you don't own anything any more, haha)
A friend of mine from Sun went to Santa Clara university for their MBA and I was always interested in running a company and we talked about what core knowledge the school felt that a "Master" of business administration should know. Not too surprisingly there was a lot about how you could organize a bunch of unreliable elements (employees) into a structure that reliably delivered results (products). And in one sense that is exactly what cloud computing did as well, organizing unreliable "cheap" PC type servers, into a fabric of reliable service delivery. It is the system of organizing the parts that is the solution.
Greg Lindahl was one of the founders of Blekko and his experience of managing hundreds of machines for the super computer types with big machine budgets and low IT budgets, was essential to Blekko's ability to deliver its search engine. With a DevOps group of six (one manager and five SREs) Blekko managed 2000+ machines in two data centers. IBM was astonished when they bought us that we could manage with so few devops engineers.
We had also done the "What would this cost us on AWS" calculation many times with many different variables and the 'break even' point, where the cost to self host at a Colocation facility fell below the cost to host with AWS was 120 servers. And that included reducing the DevOps group size to 3. So about 10 racks worth of servers. There was a really funny experience when IBM insisted we move our infrastructure to SoftLayer (an IBM business) when IBM acquired us and we pointed out that the "soft money" cost of running our infrastructure on SoftLayer was about $9M a month versus $120K a month. Which their finance group shut down that talk and just renewed the lease in the Colo.
But to successfully run a big distributed thing like that, you needed to both put your servers in a colocation facility and have Greg's software which allowed you to manage them with a small team. The total cost of that infrastructure depended on it.
I got a chance to go visit the 0xide team and they absolutely "got" that requirement. If I were running Engineering and Operations in Blekko today we totally would have kicked Supermicro to the curb and replaced them with the 0xide solution. I don't know if they have teamed up with Antithesis for testing their stuff but man, that combo of management software that never breaks and integrated server cabinets with all the things? That's a pretty good integration.
An MBA from the Sun Microsystems School of Hard Knocks, maybe. Otherwise I don't think so.
[0] https://www.gsb.stanford.edu/faculty-research/case-studies/n...
Maybe I'm stuck in 2015 but are people really out here rolling their own clouds these days?
[0] https://news.vmware.com/company/cpu-pricing-model-update-feb...
Different audiences have different expectations. In the place we are, that is, enterprise sales, the expectation is not that you can go to a website, get a price, and click a button. The expectation is that the two organizations will communicate over a period of time ("the sales cycle") to work through everything that goes into a deal, and eventually come to an agreement or not. This includes so many variables that putting the price on the website wouldn't make sense, as you're never going to end up at that exact price.
I used to find this attitude frustrating, as an engineer, but the longer I have been a professional, and worked at various organizations, the more I come around to that being a good thing, not a bad thing.
I feel bad burning salesman time for a solution that is way out of our range, but I can't learn that until after a meeting or two.
That quote (which may not even have been said by J.P. Morgan) is talking about luxury consumer goods, which is a completely different market than business to business sales.
By way of contrast, a payments provider we explored had four moderate sized boxes running everything. They were about six years old and fully depreciated, but more than adequate to run millions of transactions through a month.
If I don't know whether something will fit within the approximate budget for the project and can't quickly get an idea from other research I'm not going to mention it as an option. I'm used to spending about a million per rack but that's for a complete ESX cluster with storage and networking, if I have no idea how alternatives stack up against that it's hard to put it on the table.
Where I’m frustrated by this attitude is when I just want to buy 10 seats of something and it doesn’t have a price, not at the seven digit level.
I’m probably missing something obvious, so take this as a genuine question rather than an attempt to debate, but why aren’t they fundamentally connected?
If you sometimes need 10 servers and sometimes need 100 (elasticity) then with renting you can always have the proper number. If you own, then you have to own all 100 servers.
> why aren’t they fundamentally connected?
They are connected in the sense that capacity planning is always a thing. But that doesn't mean that the ownership model is inextricably tied to the deployment model, which is how I personally say instead of "elastic infrastructure." "I am making an HTTP call to request a new VM" is very different than "file a ticket with IT to procure me a new server, install it in the data center, and then send me the keys." The slogan "pets vs cattle" is an advanced form of this change in thinking. In some sense, it is hard to own your hardware yet treat your servers like cattle.
You basically have four basic options:
* deploy elastic, rent your hardware: this is the default today for many startups
* deploy to hardware directly, buy your hardware: this is the Old Times
* deploy elastic, buy your hardware: this is Oxide
* deploy to hardware, rent your hardware: this is a thing, though not nearly as popular as the other three
(I call them basic because hybrid is a thing: you can own usual capacity and then rent more for bursts, or as a fallback, etc)
Now, Oxide is not the only game in town when it comes to "deploy elastic, buy your hardware": this is what many IT departments do. You could, for example, buy some servers, toss OpenStack on there. Oxide's thesis is that doing this is less than ideal, and we can make it significantly better. In fact, it is so suboptimal that many people choose "deploy elastic, rent your hardware" because it is so much easier. And now we're back to Bryan's statement.
Does that makes sense?
In my industry (telco) we had two teams: my team ran our own hardware, the other team ran less than 10% of our workload on an AWS stack that cost as much per month as we paid per year - including annualised capital costs.
They also had double the ops team size (!!), they had to pay for everyone to be trained in AWS, and their solution was far more complex and brittle than ours was.
Assuming Oxide would have been price competitive with what we were already using, I would have jumped at the chance to use them, I could have brought the other team on board, and I think it would have given us a further cost and performance advantage over our AWS based competitors.
The cloud is good for many users, especially if they migrate to cloud native system design, but as a telco you would probably have facilities and connectivity which helps out a lot.
Companies like Mirantis, choosing a technology completely inappropriate for distributed systems (puppet) put a bad taste in the mouths of many people.
I implemented OpenStack at one previous employer to just convince them that they could run VMs, intending it to allow for a cloud migration in the future.
As they ran a lot of long lived large instances it was trivial to make it cheaper to run in our own datacenter. Well until I moved roles and the IT team tried to implement it with expensive enterprise gear and in an attempt to save money used FEXs despite the fact I had documented they wouldn't work for our traffic patterns.
Same thing during the .com crash. I remember our cage was next to one of the old Hotmail cages with their motherboards on boards. We were installing dozens and dozens of Netras and Yahoo was down the hall with a cage full of DEC gear...we went under because we couldn't right-size(cost) our costs.
A lot of the companies who save a lot in cloud migrations were the same, having decked out enterprise servers, SANs, and network gear that was wasted in a private cloud context.
Enterprise _ is often a euphemism for we are fiercely defending very expensive CYA strategy irrespective of value to the company or material risk.
What is a FEX? Feline Expedited eXchange?
Our production workload was pretty homogenous, and we were super cheeky - we’d use previous generation servers to keep capex down. We didn’t even use VMs, just containers; docker-swarm was good enough (barely). Our bottleneck was always iops, so we’d have a few decked-out machines to run our redundant databases. It worked fine, but I have subsequently enjoyed working with k8s a lot more.
We did use enterprise gear, but previous gen stuff is so much cheaper than current gen. So our perf per watt was not great, and we’d often make the decision to upgrade a rack when we hit our power limits, since we rarely had other constraints.
As mentioned elsewhere, we did use AWS spot instances for fractional loads like build and test. It’s not that we didn’t use cloud, it’s that we used it when it made sense.
All of that said, I do suspect the equation has changed - not with AWS, but with Vultr. I’ve deployed some complex production systems there (Nomad, Kafka+ZK, PG) and the costs are much closer to in house. They have also avoided the complexity of binding all their different services together. They also now provide K8s and PG out of the box, charging only for the cost of the VMs - as opposed to the wild complexity of AWS billing.
So maybe I’m coming around.
I honestly don't know who is on the right side of technology and economics here. One glance at Oxide or Nexus says these systems aren't cheap. But if those industries are right, and regular cloud genuinely won't cut it, then maybe these systems can find a niche.
One thing that would determine the fate of these platforms is if the major vendors, that everyone in a given industry uses, commit to it. For example, in telco that might be Ericsson, Nokia, and Huawei. I don't know if they have sufficient incentives to do so yet. They would, I guess, prefer to wait for a clear leader to emerge and then commit to only that rather than bet on three or four speculative horses?
Many applications don't depend on all that stuff, just vm, storage and db. Partly that is by design so the application can be deployed in different clouds. Also older application that moved to the cloud.
Definitely agree that you can easily pay too much for cloud services and that the hyperscalers have every incentive to put as much margin in their pockets as they can get away with.
> So these folks might be smart but they've not shipped very much product yet.
We just started shipping late last year, announcing our first two customers then: https://oxide.computer/blog/the-cloud-computer
So yes, still early :)
Why aren't there more players in this area?
If the difficulty and cost were the only challenge, this would be a candidate to do at a larger company, but the cross-domain nature of it makes it really thorny: you would need a lot of internal alignment to succeed -- and (more challenging) you need to maintain that internal alignment for a protracted period of time. I had done something not wholly dissimilar (though frankly much less ambitious) at Sun back in the day[0] -- and even though Sun was much more amenable to this kind of disruptive endeavor than any company of its size[1], we barely pulled it off. Indeed, some of the worst behavior I ever saw at Sun was from the people who trying to prevent us from succeeding because they felt it threatened them. I simply cannot imagine doing something more ambitious at Sun -- or as ambitious anywhere else.
We came to the conclusion that it really has to be new company formation -- which means raising money to do it, which means finding the right investors. Even though the upside is extraordinary, finding the right investors in hard tech is really, really tough[2]. So yes, there should be more options -- but there aren't, and there are frankly unlikely to be in the foreseeable future...
[0] https://bcantrill.dtrace.org/2008/11/10/fishworks-now-it-can...
[1] https://bcantrill.dtrace.org/2011/07/12/in-defense-of-intrap...
[2] https://oxide-and-friends.transistor.fm/episodes/deep-tech-i...
edit: and the current growth-area is GPU/ML processing, where Oxide has bupkis.
Your time, hopefully, isn't.
>The margins on server kit are tiny.
And man, does it ever show.
>Oxide gets plugged here non-stop, but the audience here still struggles to see what the point is.
I'd like someone to appreciate music with; but alas, my cat doesn't like Debussy, she likes tuna, and I must love my cat for what she is.
That the HN audience struggles to see the point of a properly-integrated systems offering isn't an indictment of that offering. They (and good for them!) may never have had the professional experience of having to manage a large enterprise IT resource on questionable hardware with indifferent software support from multiple vendors.
>Sounds like a vanity project to keep some of the old Solaris talent in-pocket.
Then that would be a vanity project for Oracle; but I believe that ship has sailed.
>the current growth-area is GPU/ML processing, where Oxide has bupkis.
Nobody but nvidia has anything hw-wise in that area (for the moment...).
Even on low-end hardware, the cost of competent deployment amortizes to zero. Like most industrial pursuits, the deployment cost of the first box is high; the cost for each additional box is orders-of-magnitude less. Add the end of Moore's Law into the mix--getting many years of economic service out of boxes that would have been obsolete within the year--the cost is lower still. There are people out there who've been running aisles of rock-bottom supermicro gear, for almost decade, without incident. Gabriel's Law: "worse is better". FAANG rode Gabriel's law to the bank.
> Then that would be a vanity project for Oracle
As for "in-pocket", I did not mean Frankenstein's Solaris. I meant Operation Paperclip. Knowledge was acquired on someone else's dime, and it would be imprudent to leave this talent disengaged--to be retrieved and used by one's competitors.
> Nobody but nvidia has anything hw-wise in that area
Exactly. That is the castle that needs to be stormed. That is where the treasure is. 0xide is running a war against an already-dying kingdom. They might as well be playing minecraft.
Have there? But not you I take it?
Bryan worked at joyant and ran a public cloud on dell and supermicro gear. And it seems they had plenty of problems.
I have heard the same sentiment im other places.
Additional evidence comes from the cloud providers themselves. Why did all of them develop their own stuff if operations with that stuff is so problem free?
Google tried it and rejected it quickly.
I have a variety of customers running a variety of different x86_64 gear. From an up-time perspective, there has been low correlation between price and reliability. If we're talking non-x86_64 systems--like zSeries or industrial control systems--then the price and the reliability are correlated, but that is not an x86_64 market, and probably never will be.
> Bryan worked at joyant and ran a public cloud on dell and supermicro gear. And it seems they had plenty of problems.
Never called it perfect. I wrote "worse is better"--the "best" losing out to the "good enough". The supermicros I mentioned were "good enough". Of course Cantrill et. al. had problems with them, since they were running Solaris/Indiana on x86_64 gear, expecting Sun-tier SLAs. They'd need Sun-tier hardware to accomplish that, but the Sun that made that kind of gear is dead, just like Joyent. They are both dead for a reason: worse is better. Google also used bare-bones discount servers in their early days (not wasting money on brands, SLAs, and fancy cases[0]).
> Additional evidence comes from the cloud providers themselves. Why did all of them develop their own stuff if operations with that stuff is so problem free?
Let's say there are three tiers of users 1) small-fries who spin up a few AWS images, 2) businesses who've grown to the point where public cloud is no longer economical, but not invested enough in computing to justify developing their own hardware, and 3) FAANG, who have so much money and so much need that it would be absolutely stupid for them not to develop their own hardware. At FAANG-scale it's no longer just about cost or quality; it's about being in control of your own business. #2 would seem to be 0xide's target market, but as someone who inhabits that space, I am unimpressed.
[0] https://en.m.wikipedia.org/wiki/File:Google%E2%80%99s_First_...
What good enough is depends. Sure maybe if you built new processors and everything it would be even better.
But using the same processor at the core as typical servers, and just sounding it with a slightly different architecture can improve things. That's exactly what google does too.
> Of course Cantrill et. al. had problems with them, since they were running Solaris/Indiana on x86_64 gear
Ah the old 'its the os fault' excuse. Classic.
> dead, just like Joyent
Joyent was bought, it didn't go bust and the actual infrastructure is still running.
> Google also used bare-bones discount servers in their early days
Again, that's exactly what I am saying. Why do you think google stopped doing that?
> At FAANG-scale it's no longer just about cost or quality; it's about being in control of your own business.
Except if you actually listen to interviews of the people who did those innovations at Google and co, they stopped working with the traditional vendors and did costume stuff, it was because they couldn't get the quality out of the traditional architecture.
That is also why google invested in coreboot, linuxboot and are running their internal infrastructure on things like NERF firmware, instead of the standard you get from the typical vendors.
> #2 would seem to be 0xide's target market, but as someone who inhabits that space, I am unimpressed.
Have you done an actual detailed comparison with real price quotes compared to a traditional system plus all the software and setup and so on? Because if you haven't its not worth much as an opinion.
The problem wasn't Solaris. The problem wasn't x86. The problem was Cantrill & Co.'s expectations. The industry began moving to availability in software instead of hardware back in the early 2000s. Premium iron vendors fell on hard times because of this. They had a premium product with a premium price, but demand for it dried-up. Cantrill came from one of those companies, and is right to think commodity x86 hardware is crap in comparison, but it doesn't matter for the mid-tier and below markets anymore. They've already cleared this obstacle.
> That is also why google...
What Google does is irrelevant to this discussion. That is not 0xide's market. Businesses seriously competing in the compute space are going to vertically integrate as much as possible--not just motherboards and lights-out management, but all the way down to the CPU: Google's Cypress & Maple, AWS's Graviton, Apple's M1. Switching to 0xide kit is not vertical integration.
If you read with a little less haste and hostility, and a little more humility--rather than just jumping in and getting personal, you will see that I am being perfectly clear.
I am saying that this battle already played-out, 15-20 years ago. The people who need primo hardware already have things like zseries, which 0xide is in no position to compete with. The people who are competing against Azure and AWS have to make their own hardware--partly to distinguish themselves in the marketplace, partly because they'd end up sharecropping for their vendor if they did otherwise. The rest is in the commodity band, where interchangeability is key. An okay server that I can swap-out with any other vendor is an asset. A fantastic server, with custom components & bespoke tooling, that I can only get from one supplier, is a liability. All of their open sourcing makes no difference if it only runs on their equipment.
>"People who need primo hardware" ...can buy a mainframe. That is quite the take.
>The rest is in the commodity band, where interchangeability is key. An okay server that I can swap-out with any other vendor is an asset.
Setting aside the trouble one has already incurred with a system that needs to be swapped out; switching vendors (even of erstwhile "commodity" systems)isn't a neat process in any medium-to-large organization. There's purchasing, contracts for maintenance, depreciation, validation by your DC people that power and cooling for the new box is good...
And...in the end, if you your hardware is truly commodity, swapping out for another vendor is not going to yield a better outcome in the long term. You'll be riding the crap-hardware merry-go-round all over again, just on a different horse.
Here's a nickel, kid. Go buy yourself a better server.
You literally called me a 'fanboy' but I am somehow I am hostile. I didn't say anything hostile in the slightest.
You even doubled down on your position. If you were ever gone be google or a mainframe you would have done it already, so be happy with the garbage provided. That literally your argument.
> I am saying that this battle already played-out, 15-20 years ago.
Markets change over time. Its never set in stone. The industry is continuing to grow, computing needs increase for all companies and many, many billions get spend each year on new servers.
The main reason the 'fancy' hardware like Sun and friends failed, was because of the processor on the hardware side, and because of open source linux on the software side. The costume motherboards that cost slightly more and firmware weren't actually the real problem.
Also, 20 years ago was before virtualization and scripted setup were universal, things were just totally different. Those decisions should effect what the right decisions are today.
> The people who need primo hardware already have things like zseries
Again, the market is growing, any money companies who might have needed 10 computers in the past now need 100. And those that needed 100, no need 1000. And so on.
There is a gigantic gap between a mainframe and your standard Dell PC Server, there is lots of space between those things.
> The rest is in the commodity band, where interchangeability is key.
And yet a huge amount isn't actually this perfect interchangeable commodity. Sure you can get any basic OS to boot and there is some standards, but the situation isn't close to actually being a true commodity.
It also depends on where you expect interchangeability. If you define much of your infrastructure with terraform scripts and a few services to monitor everything, then that's the interchangeability layer you actually care about. Sure if you buy a bunch of Oxide servers, you can just rip one out and replace with the server from some other vendor. You have to rip out the whole rack and replace it with some other rack.
But whatever that something else is, its just gone be another target for your terraform scripts that gets monitored in a comparable way (against hat part isn't neatly standardized anyway). Industry 'standards' like RedFish are not close to providing that even for pretty low level stuff. And once you want to do more complex things there are even less standards.
If you operate a number of VxRail Dell Rack, its not that easy to just rip it out and replace it with some rack of HP machines. Again, this is specially true if you use anything but the most basic features, because very quickly you are in proprietary non-standard software hell. Guess what, Dell doesn't actually want their racks to be perfectly easily replaceable.
And if you have an issue that your Dell rack and your HP rack behave inconsistently, guess who isn't gone care about that, Dell and HP. And guess what, neither of their implementation is open source either so good luck with that. There is a reason people stick to a low number of vendors.
It definitely shows who it's actually an indictment of though.
Yes and about 5000 startups plus many large going after that market.
The market Oxide is after is proven to be large, not very innovative and very few new competitors. And guess what, its not industrial maschines that survive 50 years. Servers are not cost effective after a few years, so there is a lot of turnover. Costumer also don't have a huge amount of brand loyality.
I can see worse buissness strategies then that.
But of course other things like egress, serverless services more than make up for it.
Per core maybe made a tiny bit of sense when a server was one or two cores. But that era's long gone.
What next? Let’s charge users based on how many instructions they have in flight for our application. These ILP cheaters have been mooching off the rest of us for ages!
I wish there was a much better way to compare vCPUs than just "count" because one core of some five year old server is not the same as a core of a modern performance beast.
The 'Oracle doesn't have customers, it has hostages'
The main purpose shifted from a hypothetical development cost offset to an effort to extract more fees from customers who were locked into your product.
Broadcom's changes to VMware, where they have publicly stated that increasing revenue while not attempting to build new revenue streams or through any value adds for the customers is happening now.
I get that any technology moves to an extraction phase after growth slows.
But for me, moves like per core licenses are an indicator we need to consider vendor mitigation.
Unfortunately capital markets seem to prefer these lower long term but predictable gains through extraction vs long term dividends etc...
Just a weird novelty impact that I hadn't seen before, not really a big deal. Also I retract "invert". I went back and looked again and it's definitely not inverted. Check it out!