Oxide builds servers as they should be [audio]
changelog.com
changelog.com
There's tons more where that came from; if you are a fan of On the Metal, I don't think you'll be disappointed -- and it has the added advantage that you join a future conversation!
[0] On The Metal: https://podbay.fm/p/on-the-metal
[1] Ken Shirriff: https://podbay.fm/p/on-the-metal/e/1611669600
[2] Friends of Oxide: https://podbay.fm/p/oxide-and-friends
[3] Tales from the Bringup Lab: https://podbay.fm/p/oxide-and-friends/e/1638838800
[4] Theranos, Silicon Valley, and the March Madness of Tech Fraud: https://podbay.fm/p/oxide-and-friends/e/1632182400
[5] Another LPC55 Vulnerability: https://podbay.fm/p/oxide-and-friends/e/1649116800
[6] The Pragmatism of Hubris: https://podbay.fm/p/oxide-and-friends/e/1639443600
[7] The Sidecar Switch: https://podbay.fm/p/oxide-and-friends/e/1638234000
[8] Debugging Methodologies: https://podbay.fm/p/oxide-and-friends/e/1652745600
[9] The Books in the Box: https://podbay.fm/p/oxide-and-friends/e/1632787200
And maybe do an addendum 5 minute announcement episode on the “on the metal” podcast?
Love your excellent work, thank you all
And the 5 minute announcement on the On the Metal feed is a great idea; thanks for the idea -- and for the kind words.
Finally, I hasten to add: we have a really exciting Space coming up on Monday[0], where we'll be joined by Jon Masters to talk about the importance of integrating hardware and software teams; join us!
[0] https://twitter.com/bcantrill/status/1545441245853495296
Would love recommendations for in depth, tech related podcasts that stand on their on, like On The Metal.
Some of them have even been delving into some of the participants personal politics as well which is just something I'm not interested in hearing.
While there may be some unfocused rants in there (sorry, I guess?), there's also a lot of extraordinary technical content -- certainly, if anyone else has described their board bringup experiences in as explicit technical detail as we have in [1] and [2], I would love to be pointed to it!
[0] https://podbay.fm/p/oxide-and-friends
[1] Tales from the Bringup Lab: https://podbay.fm/p/oxide-and-friends/e/1638838800
[2] More Tales from the Bringup Lab: https://podbay.fm/p/oxide-and-friends/e/1650326400
Another thing that ended up being frustrating is repeated references to some image and I had to pause and go digging through people's personal twitter accounts to try and find the image that was being talked about. There seemed to be an assumption that the listener follows all employees personal twitter accounts.
There's also audio quality issues in general as it's not nearly as good a quality of media equipment as was used for the original on the metal podcast. (Maybe it's similarly cheap equipment, but it was better quality before regardless.)
[I wouldn't have known if I hadn't stumbled onto these replies. :( ]
I ask genuinely … because I don’t understand who will spend the premium for a rack server that is easier to maintain?
In a world of cloud and dedicated hosted servers (where the users of the server is not the buyer of the server) - who would buy Oxide?
It seems like whoever can build the cheapest server wins now.
Said differently, cloud has completely changed the value-chain in server buyers
obviously AWS, azure, etc do their own hardware procurement and totally bespoke stuff in house.
so this would be targeted at any smaller players in hosting companies that are not at that massive size of azure, but need vast quantity of servers.
But my understanding Oxide is/will-be priced as a premium over a typical x86 server.
It seems unlikely that any hosting company is going to pay a premium just for a bit more convenience.
Note: I'm not trying to be a "hater". Just genuinely not understanding the market here.
have been on the manufacturing side of this before.
bulk x86-64 servers such as you could buy from the vendors you meet at the computex taipei are a big part of it.
things along the same general idea as the supermicro that is 2U high and has four discrete nodes in it, but cheaper.
long/narrow motherboards with dual socket xeons
and other similar stuff...
if the price for this is not market competitive with that, or just a bunch of 1U commodity dual socket servers, the only people buying it will be enterprise/specialty end user and not hosting operations.
What they seem to want to do is give you what a rack (or 2) full of Dell Servers plus the price of a VMWare license would give you.
And then you should also save money on maintenance, space, updates, compliance and so on.
If the approach to firmware and virtualization means you can have just 1/2 employ less then that's already a huge amount of money.
You can proxy this demand by looking at VMWare sales. Anyone running a large amount of baremetal VMWare virtualization has a large on premise deployment and could find managed rack products like Oxide interesting.
Dell manages an annual revenue of greater than $20B, so on-prem HW is clearly going strong, regardless of whether or not you think it should be.
You know how you can buy pizza in a cardboard box? What happens if you need to feed a party of ~80 people? You end up getting dozens of pizza boxes delivered.
Oxide Pizza delivers all that pizza in a single box.
Not my area, but based on other areas: "easier to maintain" can translate directly to "fewer engineer/admin hours consumed", which can translate directly to saved money (or resources that are now freed up to be put toward something more useful)
I haven't used RedFish but it seems it does not address the problems of IPMI in this regard: To be fully automatable there shouldn't be an assumption about the network environment - rather I would like that to have an L2-based protocol (one has to use the MAC address as identifier anyway) or if IPv4/v6 is used then at least demand mandatory link-local auto-addressing by default. However, RedFish seems to support emulating a CD-ROM drive for booting an installation media, this is a nice idea but what I really would like to see is directly writing an image to the hard disk. Then of coure the hardware manufacturers would need to write high-quality firmware and not the regular quirk- and bug-ridden stuff that we know from IPMI implementations.
Boot a "CD-ROM" via IPMI into a custom installer, flash image (generated from a Dockerfile!) to disk, reboot. Install takes ~15 seconds, full process (going via POST once) takes ~4 minutes. This also allows us to have a "hardware validation" image (that one doesn't get persisted to disk).
Not sure if there's any plans to make it public right now, but I'll ask around. Feel free to contact me via email (in profile)
Please do. There's a need for it. :)
Any hints about how to do this? I'd love to use something as nice to work with as a Dockerfile to build bare metal installs.
Fly.io[1] also does this (although they boot the result in a VM, the concepts are the same)
[0]: https://blog.davidv.dev/docker-based-images-on-baremetal.htm...
So, I share your excitement when it works but I'm not happy with the out-of-the box experience of IPMI. I wish there was something better but unfortunately this is just one case of a split/mismatch of what HW people provide and what software people need.
CPU: 2,048 x86 cores (AMD Milan)
Memory: Up to 32TB DRAM
Storage: 1024TB of raw storage in NVMe
Network switch: Intel Tofino 2
Network speed: 100 Gbps
That's a pretty meaty server!If they're following the typical 1/5x pricing model for support, that'd be roughly $500k/yr/rack. But it's also hard to do that while simultaneously describing Dell as "rapacious".
Cantrill is, shall we say, very opinionated. And does not suffer fools gladly.
Edit: but also, it's a YC (Anti)Application video, so being high-and-mighty is practically a prerequisite.
I write something longer in a previous thread: https://news.ycombinator.com/item?id=30678324
Even for servers in my building, I find it’s more convenient than a crash cart for most things (and of course my desk is a better work environment).
VCE used Cisco C2xx and C4xx rackmount and BLxxx blade servers (before the EMC/Dell acquisition). Now, I believe they use all Dell hardware.
Hmmm wrong Bryan? Breaking Bad wouldn't look the same if you swapped them, I'm afraid.
println!("{my_name}");
The host is using Illumos and bhyve for hypervisor stuff. I don’t work on that part of the stack personally so that’s the most detail I can give you.
[0] https://github.com/oxidecomputer/omicron
[1] https://github.com/oxidecomputer/propolis
[2] https://github.com/oxidecomputer/crucible
[3] https://artemis.sh/2022/03/14/propolis-oxide-at-home-pt1.htm...
Why?
Or a bunch of the the 4-servers-in-a-single-2U-package setup from Supermicro?
And then interconnecting them by your own choice of 100GbE switches at top of rack.
If you do as you suggest you will have to do a whole lot more, at the minimum you need to set-up (and pay) VMWare.
Then you will need to figure out how you do automated firmware upgrades, low level monitoring and if security is a concern you likely want to configure secure booting with attestation and things like that.
Consider how much work it is to do what you suggest, vs buying this rack. And then consider how good (and repeatable) the end result is.
That is at least the theory, see:
It’s the kind of thing I would expect Nvidia or Dell to become big at as well, and Bryan Cantrill has enough leverage to pique my interest significantly here.
An example is nvida's bluefield, there are others
That sort of thing is being used in a few areas, can even run ESXi on them.
There was a fun article on servethehome a while back using them to build an arm cluster (Link: https://www.servethehome.com/building-the-ultimate-x86-and-a... )
Same problem with Anthos.
You can work around the limitations by having a million points of presence with them and splitting up your workloads but there’s real costs associated with that model. Running 1 rack in 100 DCs is a lot harder than 10 racks in 10 DCs.
I think Outpost is a great hybrid-ish solution for those companies with workloads in AWS that want to supplement that with on-premise workloads using the same APIs and tools they are already used to.
The use cases I see for Oxide are, like other have said, are for hyperscalers or hosting providers who want to have their own on-premise infrastructure that isn’t tied to another entity like AWS. Whether that’s because you have the in-house talent to manage it or for compliance reasons it’s required it does have a niche.
Wikipedia has a short page with a different definition: https://en.wikipedia.org/wiki/Hyperscale_computing
But if outpost was nearly the same price as oxide, wouldn’t you always want access to the full AWS or GCP infrastructure and world class control plane versus something bespoke that only works in your own data center.
I think what Oxide has to bank on is that this isn’t a market that Google, Microsoft or AWS want to really compete in (e.g. it may require these companies to undercut themselves too much).
My gut is that they are for places that need on-Prem for compliance purposes for certain parts of their stack
But that seems entirely due to AWS not currently being interested in this market versus something fundamentally limiting about their tech.
https://en.wikipedia.org/wiki/Hyper-converged_infrastructure
Oxide racks won't be able to use anything other than Oxide computers (or maybe eventually third party computers that follow Oxide standards).
Furthermore pretty leading question to compare Brian to Steve Jobs forced me out.