Why Oxide Chose Illumos
rfd.shared.oxide.computer
rfd.shared.oxide.computer
{{citation needed}}?
When I ran the numbers in 2019, there hadn't been guest exploitable vulnerabilities that affected devices normally used for IaaS for 3 years. Pretty much every cloud outside the big three (AWS, GCE, Azure) runs on QEMU.
Here's a talk I gave about it that includes that analysis:
slides - https://kvm-forum.qemu.org/2019/kvmforum19-bloat.pdf
Google also uses KVM with a variety of userspace stacks: a proprietary one (tied to a lot of internal Google infrastructure but overall a lot more similar to QEMU than Amazon's) for GCE, gVisor for AppEngine or whatever it is called these days, crosvm for ChromeOS, and QEMU for Android Emulator.
[ec2-user][~]$ hostnamectl
Static hostname: ip-x-x-x-x.ec2.internal
Icon name: computer-vm
Chassis: vm
Machine ID: ec2d54f27fc534ea74980638ccc33d96
Boot ID: 6caf18b7ed3647819c1985c11f128142
Virtualization: xen
Operating System: Amazon Linux 2023.5.20240903
CPE OS Name: cpe:2.3:o:amazon:amazon_linux:2023
Kernel: Linux 6.1.106-116.188.amzn2023.x86_64
Architecture: x86-64
Hardware Vendor: Xen
Hardware Model: HVM domU
Firmware Version: 4.11.amazonhttps://www.theregister.com/2017/11/07/aws_writes_new_kvm_ba...
It could be that it's not all over and tied to specific machine types still, or there's something they've done to make it report to the guest still that it's xen based for some compatibility reasons.
So there existed known guest-exploitable vulnerabilities as recently as 8 years ago. Maybe that, combined with the fact that QEMU is not written in Rust, is what is causing Oxide to decide against QEMU.
I think it's fair to say that any sufficiently large codebase originally written in C or C++ has memory safety bugs. Yes, the Oxide RFD author may be phrasing this using weasel words; and memory safety bugs may not be exploitable at a given point in a codebase's history. But I don't think that makes Oxide's decision invalid.
Also QEMU's fuzzing is very sophisticated. Most recent vulnerabilities were found that way rather than by security researchers, which I don't think it's the case for "competitors".
But I still don't think that makes Oxide's decision or my comment necessarily invalid, if only because of an a priori decision to stick with Rust system-wide -- it raises the floor on software quality.
Languages cannot possibly do this.
It's also possible for a language to raise the ceiling of software quality, and Zig is an excellent example.
I'm thinking of "floors" and "ceilings" as the outer bounds of what happens in real, everyday life within particular software ecosystems in terms of software quality. By "quality" I mean all of capabilities, performance, and absence of problems.
It takes a team of great engineers (and management willing to take a risk) to benefit from a raised ceiling. TigerBeetle[0] is an example of what happens when you pair a great team, great research, and a high-ceiling language.
Cargo is widely recognized as low quality. The thesis fails within it's own standard packaging. It's possible for a language to be used by _more people_ and thus raise the quality _in aggregate_ of produced software but the language itself has no bearing on quality in any objective measure.
> to benefit from a raised ceiling
You're explicitly putting the cart before the horse here. The more reasonable assertion is that it takes good people to get good results regardless of the quality of the tool. Acolytes are uncomfortable saying this because it also destroys the negative case, which is, it would be impossible to write quality software in a previous generation language.
> TigerBeetle[0] is an example
Of a protocol and a particular implementation of that protocol. It has client libraries in multiple languages. This has no bearing on this point.
Can you point me to both of:
* why it's considered low quality
* evidence of this "wide regard"
Other than random weirdos who think allowing dependencies is a bad practice because you could hurt yourself, while extolling the virtues of undefined behavior - I've never heard much serious criticism of it.
Other software providing the same features produce better results for those users. It's dependency management is fundamentally broken and causes builds to be much slower than they could otherwise be. Lack of namespaces which is a lesson well learned before the first line of Cargo was ever written.
I could go on.
> evidence of this "wide regard"
We are on the internet. If you doubt me you can easily falsify this yourself. Or you could discover something you've been ignorant of up until now. Try "rust cargo sucks" as a search motif.
> random weirdos
Which may or may not be true, but you believe it, and yet you use your time to comment to us. This is more of a criticism of yourself than of me; however, I do appreciate your attempt to be insulting and dismissive.
You're right, I'm being dismissive of weasely unbacked claims of "wide regard". It's very clear now that you can't back your claim and I can safely ignore your entire argument as unfounded. Thanks for confirming!
>* why it's considered low quality
>* evidence of this "wide regard"
> Acolytes are uncomfortable saying this because it also destroys the negative case, which is, it would be impossible to write quality software in a previous generation language.
Not impossible, just a lot harder. It's as if you're thinking in equations that are true/false, while I'm thinking in statistical distributions.
Have you used Macintosh System 5? How about Windows 3.1? Those were considered quality systems at the time, but standards are up, way up since then.
Why are modern systems better? Is it because we have better developers today? -- I don't think so. It took a "real" programmer to write quality apps in Pascal for early Macintosh systems, or apps in C for Windows 3.1.
I think the difference is in the tooling that is available to us -- and modern programming languages (and libraries) are surely a very large part of that tooling.
If you disagree, I challenge you to find a seasoned modern desktop app developer who can write a high-quality app for MacOS or Windows that looks and functions great by modern standards and doesn't use any modern languages or directly invoke any non-vendor libraries built after the year 2000. It's possible[0]. They may be able to do it, but you must certainly concede that doing a great job requires a much better developer than the average modern desktop app developer to be able to work well under those kinds of constraints.
That's what I mean by "raising the floor" -- all software gets better when languages, libraries, and tooling improve.
[0] https://stackoverflow.com/questions/30269329/creating-a-wind...
That's totally true, and I never contradicted that.
By "ceiling" I mean the limit of what is possible. A great team can do a lot more when they have great tools.
2. OP (bonzini) has given specific and valid arguments that that statement is wrong.
3. You're not answering to that specific arguments, but defending Rust and bashing C++ generally without giving any prove.
4. bonzini again provides specific arguments that your generalization is not correct in that context. That despite Firecracker being written in Rust it had a security issue.
5. You still insist without given any solid argument. You just insist that Rust is superior. Not helpful in any discussion. Think about it.
I'm not bashing C++ beyond saying "any sufficiently large codebase originally written in C or C++ has memory safety bugs". I did not say those bugs are exploitable, just that they're present.
I'm also not insisting Rust is superior, except to say that it raises the floor of software quality, because it nearly eliminates a class of memory safety bugs.
Do you disagree? Neither of those statements implies C++ sucks or that Rust is awesome. Just 2 important data points (among many others) to consider in whatever context you're writing code in.
Google – "KVM-based hypervisor"
Azure – Hyper-V
You can of course assume that all of them heavily customize the underlying implemenation for their own needs and for their own hardware. And then they have stuff like Firecracker, GVisor etc. layered on top depending on the product line.
Oracle Cloud - QEMU/KVM
Scaleway - QEMU/KVM
QEMU typically uses KVM for the hypervisor, so the vulnerabilities will be KVM anyway. The big three all use KVM now. Oxide decided to go with bhyve instead of KVM.
Usually QEMU runs heavily confined, but remote code execution in QEMU (remote = "from the guest") can be a first step towards exploiting a more serious local escalation via a kernel vulnerability. This second vulnerability can be in KVM or in any other part of the kernel.
This isn't true - Azure uses Hyper-V (https://learn.microsoft.com/en-us/azure/security/fundamental...), and AWS uses an in-house hypervisor called Nitro (https://aws.amazon.com/ec2/nitro/).
I thought Azure was moving/moved to KVM for Linux, but I was wrong.
>AWS uses an in-house hypervisor called Nitro
Nitro uses KVM under the hood.
https://www.brendangregg.com/blog/2017-11-29/aws-ec2-virtual...
How many reliability bugs has QEMU experienced in this time?
The man power to go on site and deal with in the field problems could be crippling. You often pick the boring problems for this reason. High touch is super expensive. Just look at Ferrari.
The difference between
"Buy our Illumos/Bhye solution! Why? I have been an Illumos/Bhyve Maintainer!"
and "Buy our Linux/KVM solution! Why? I have been an Illumos/Bhyve Maintainer!"
should make my point a bit clearerAnd finding people to heir that know Linux/KVM wouldn't be a problem for them.
This evaluation was done years ago and they added like 50 people since then.
Saying 'We have a great KVM Team but our CEO was once an Illumos developer' is perfectly reasonable.
And as I point out in my other comment, the former Joyant people like know more about KVM then anything else anyway. So it would be:
"Buy our KVM Solution, we have KVM experts"
But they evaluated that Bhyve was better then KVM despite that.
Of course, but that is less of unique selling point.
> But they evaluated that Bhyve was better then KVM despite that.
If you are selling Bhyve you better say that whether it's true or not. So why should I, as a reader or employee or customer, trust them?
Who cares about uniqueness? That's not a goal.
> If you are selling Bhyve you better say that whether it's true or not. So why should I, as a reader or employee or customer, trust them?
They are not selling Bhyve. This is an internal document. Their costumers don't care about the implementation details. And if they do, then they will do their own evaluation based on their own evaluation.
As an employee you trust it because you know how the company heirs and who wrote these RFDs.
As a reader, its literally like any other thing on the internet.
In the long term if $A is actually better then $B, then it makes sense to start with $A even if you don't know $A. Because if you are trying to building a company that is hopefully making billions in revenue in the future, then long term matters a great deal.
Now the question is can you objectively figure out if $A or $B is better. And how much time does it take to figure out. Familiarity of the team is one consideration but not the most important one.
Trying to be objective about this, instead of just saying 'I know $A' seems quite like a smart thing to do. And writing it down also seems smart.
In a few years you can look back and actually say, was our analysis correct, if no what did we misjudge. And then you can learn from that.
If you just go with familiarity you are basically saying 'our failure was predetermined so we did nothing wrong', when you clearly did go wrong.
Upstreaming those changes into FreeBSD bhyve is a more complicated situation, given that illumos has diverged from upstream over the years due to differing opinions about certain interfaces.
Mosdef.
IIRC, these RFDs are part of Oxide's commitment to FOSS and radical openness.
Whatever decision is ultimately made, for better or worse, having that written record allows the future team(s) to pick up the discussion where it previously left off.
Working on a team that didn't have sacred cows, an inscrutible backstory ("hmmm, I dunno why, that's just how it is. if it ain't broke, don't fix it."), and gatekeepers would be so great.
Even if you think it's a foregone conclusion given the history of bcantrill and other founders of Oxide, there absolutely is value in putting decision to paper and trying to provide a rational because then it can be challenged.
The company I co-founded does an RFD process as well and even if there is 99% chance that we're going to use the thing we've always used, if you're a serious person, the act of expressing it is useful and sometimes you even change your own mind thanks to the process.
Bryan Cantrill is CTO of Oxide [1].
I assume that has no bearing on the choice, otherwise it would be mentioned in the discussion.
[1] https://bcantrill.dtrace.org/2019/12/02/the-soul-of-a-new-co...
And before that, they used to run FreeBSD.
Mentioned for example in this comment by Bryan Cantrill a decade ago:
https://news.ycombinator.com/item?id=6254092
> […] Speaking only for us (I work for Joyent), we have deployed hundreds of thousands of zones into production over the years -- and Joyent was running with FreeBSD jails before that […]
And I’ve seen some other primary sources (people who worked at Joyent) write that online too.
And Bryan Cantrill, and several other people, came from Sun Microsystems to Joyent. Though I’ve never seen it mentioned which order that happened in; was it people from Sun that joined Joyent and then Joyent switched from FreeBSD to Illumos and creating SmartOS? Or had Joyent already switched to Illumos before the people that came from Sun joined?
I would actually really enjoy a long documentary or talk from some people that worked at Joyent about the history of the company, how they were using FreeBSD and when they switched to Illumos and so on.
https://www.youtube.com/watch?v=eVkIKm9pkPY
This is about as good as you are gone get on the topic of Joyant history.
As I recall they were also the original host of Twitter, which if I recall was Rails back in the day.
Up until 2008:
* https://web.archive.org/web/20080201142828/http://www.joyeur...
https://www.youtube.com/watch?v=cwAfJywzk8o
As far as I know, Bryan didn't personally work on the porting of bhyve (this might be wrong).
So if anything, that would point to KVM as the 'familiar' thing given how many former Joyant people were there.
Keeping a KVM port up to date is a huge effort compared to bhyve, and they probably had learnt that in the years between the porting of KVM and the founding of Oxide.
Given all of that, and taking into account building a product on top of it, and thus needing to support it and stand behind it, Linux wasn't the best choice. Looking ahead (in terms of decades) and not just shipping a product now, it was found that an alternate ecosystem existed to support that.
Culture of the community, design principles, maintainability are all things to consider beyond just "is it popular".
Exciting times in computing once again!
1. Xen Type-1 hypervisor is smaller than KVM/QEMU.
2. Xen "dom0" = Linux/FreeBSD/OpenSolaris. KVM/bhyve also need host OS.
3. AMZN KVM-subset: x86 cpu/mem virt, blk/net via Arm Nitro hardware.
4. bhyve is Type-2.
5. Xen has Type-2 (uXen).
6. Xen dom0/host can be disaggregated (Hyperlaunch), unlike KVM.
7. pKVM (Arm/Android) is smaller than KVM/Xen.
> The Service Management Facility (SMF) is responsible for the supervision of services under illumos.. a [Linux] robust infrastructure product would likely end up using few if any of the components provided by the systemd project, despite there now being something like a hundred of them. Instead, more traditional components would need to be revived, or thoroughly bespoke software would need to be developed, in order to avoid the technological and political issues with this increasingly dominant force in the Linux ecosystem.Is this an argument for Illumos over Linux, or for translating SMF to Linux?
I don't mean FUD in a disparaging sense, more like literal fear of the unknown causing people to be excessively cautious. I wouldn't have any problem with Oxide saying "we went for what we know best", there's no need to fake that so much more research went into a decision.
I think arguably the bhyve over KVM was the more fundamental reason, and bhyve doesn't run on linux anyway.
I am obviously biased as I am a KVM (and QEMU) developer myself, but I don't see any other plausible reason other than "we know the Illumos userspace best". Founder mode and all that.
As to their choice of hypervisor, to be honest KVM on Illumos was probably not a great idea to begin with, therefore they used bhyve.
While it's true that I'm a dyed in the wool illumos person, being in the core team and so on, I have Linux desktops, and the occasional Linux system in lab environments. I have been supporting customers with all sorts of environments that I don't get to choose for most of my career, including Linux and Windows systems. At Joyent most of our customers were running hardware virtualised Linux and Windows guests, so it's not like I haven't had a fair amount of exposure. I've even spent several days getting SCO OpenServer to run under our KVM, for a customer, because I apparently make bad life choices!
As for not discussing the social and political stuff in any depth, I felt at the time (and still do today) that so much ink had been split by all manner of folks talking about LKML or systemd project behaviour over the last decade that it was probably a distraction to do anything other than mention it in passing. As I believe I said in the podcast we did about this RFD recently: I'm not sure if this decision would be right for anybody else or not, but I believe it was and is right for us. I'm not trying to sell you, or anybody else, on making the same calls. This is just how we made our decision.
In other words, I don't think that the social or technological reasons in the document were that strong, and that's fine. Rather, my external armchair impression is simply that OS and hypervisor were not something where you were willing to spend precious "risk points", and that's the right thing to do given that you had a lot more places that were an absolute jump in the dark.
That's just fine, as long as they're not choosing a clearly inferior long term option. The technically superior solution is not always the right solution for your organization given the priorities and capabilities of your team, and that's just fine! (I have no opinion on KVM vs bhyve, I don't know either deep enough to form one. I'm talking in general.)
I don't know why you think none were mentioned - to name one, they link a GitHub issue created against the systemd repository by a Googler complaining that systemd is inappropriately using Google's NTP servers, which at the time were not a public service, and kindly asking for systemd to stop using them.
This request was refused and the issue was closed and locked.
Behaviour like this from the systemd maintainers can only appear bizarre, childish, and unreasonable to any unprejudiced observer, putting their character and integrity into question and casting doubt on whether they should be trusted with the maintenance of software so integral to at least a reasonably large minority of modern Linux systems.
Unfortunately this makes modern Linux not reliable.
https://github.com/systemd/systemd/pull/554
What's your suggested alternative?
Using pool.ntp.org requires a vendor zone. systemd does not consider itself a vendor, it's the distros shipping systemd which are the vendor and should register and use their own vendor zone.
I don't care about systemd either way, but your own false representation of facts makes your last paragraph apply to your "argument".
That if they do not wish to ship a safe default, they do not ship a default at all.
There is no place documenting how to integrate the Dom0less/Hyperlaunch in a distribution or how to build infrastructure with it, at best you will find a github repo, with the last commit dated 4 years ago, with little to no information on what to do with the code.
Some preparatory work shipped in Xen 4.19.
Aug 2024 v4 patch series [1] + Feb 2024 repo [2] has recent dev work.
> hard to get actual documentation
Hyperlaunch: this [3] repo looks promising, but it's probably easier to ask for help on xen-devel and/or trenchboot-devel [4]. Upstream acceptance is delayed by competing boot requirements for Arm, x86, RISC-V and Power.
dom0less: ELC2022 slides [5] and video [6].
[1] https://lists.xenproject.org/archives/html/xen-devel/2024-08... [2] https://github.com/FidelisPlatform/xen
[3] https://github.com/apertussolutions/initrd-builder [4] https://groups.google.com/g/trenchboot-devel
[5] https://www.slideshare.net/StefanoStabellini/static-partitio... [6] https://www.youtube.com/watch?v=CiELAJCuHJg
However, two things are an issue:
1) The CDDL license of SMF makes it difficult to use, or at least that’s what I was told when I asked someone why SMF wasn’t ported to Linux in 2009.
2) SystemD is it now. It’s too complicated to replace and software has become hopelessly dependent on its existence, which is what I mentioned was my largest worry with a monoculture and I was routinely dismissed.
So, to answer your question. The argument must be: IllumOS over Linux.
With some effort, Devuan has managed to support multiple init systems, at least for the software packaged by Devuan/Debian.
> SMF is superior to SystemD ... [CDDL]
OSS workalike opportunity, as new Devuan init system?
> The argument must be: IllumOS over Linux.
Thanks :)
Maybe 15 years ago, not by a mile now. systemd surpassed SMF years ago and it's not even close now. No one in their right mind would pick SMF over systemd in 2024.
I don't really want to litigate the systemd vs. everything else argument, but as someone that has issues with systemd but is not particularly in love with sysvinit derivatives, I wouldn't mind SMF as an alternative.
You don’t lose socket activation or supervison. SMF is designed to help work in the event of hardware failure too, which systemd definitely cant handle.
It's simple and straightforward to use any other logging or network configuration system you wish.
>doesnt ever force any reload of itself are all reasonable reasons to prefer it.
It doesn't force reload it self.
>You don’t lose socket activation or supervison.
You don't lose that in systemd either.
>SMF is designed to help work in the event of hardware failure too
So is systemd.
>which systemd definitely cant handle.
It definitely can handle that. One of systemd's core functions is handling hardware events.
Did AI write this? its completely incorrect.
Tends to happen when facts are on your side.
>Did AI write this?
"Everything I dislike is AI."
>its completely incorrect.
No it's not.
Blisteringly fucking moronic that you double down.
“not losing socket activation” was a reference to the fact that systemd actually gives you that.
No, it isn't
>Systemctl reload
Yes, that is a feature.
>journald binary logging being forced on
Again, no one is forcing you to use that, this is old FUD.
>Blisteringly fucking moronic
Ad hominem means I'm right and you can't handle it.
>that you double down.
Yes, the facts haven't changed, I double down on the facts.
>“not losing socket activation” was a reference to the fact that systemd actually gives you that.
Good feature.
I'd certainly like that! I had spent some time working with Solaris a lifetime ago, and ran a good amount of SmartOS infrastructure slightly more recently. I really enjoyed working with SMF. I really do not enjoy working with the systemd sprawl.
I will note the distinction between type-1/type-2 hypervisors never really made sense, and makes even less sense today. http://blog.codemonkey.ws/2007/10/myth-of-type-i-and-type-ii...
curious about what bugs are being thought of there. Sounds like a very interesting situation to be in
Practically speaking, its hard to do it completely objectively and the in-house expertise probably colored the decision.
In general, being on your own private tech island is a tough thing to do, but many engineers would rather do that than swallow their pride.
It's a small operation, but https://openbsd.amsterdam/ have absolutely proven that OpenBSD's hypervisor is production-capable in terms of stability - but there are indeed other problems that rule against it on scale.
For those who are unfamiliar with OpenBSD: the primary caveat is that its hypervisor can so far only provide guests with a single CPU core.
I will say, though, that single VCPU guests would not have met our immediate needs in the Oxide product!
Could Oxide not have helped push multi-vcpu guests out the door by sponsoring one of the main developers working on it, or contributing to development? From a secure design perspective, OpenBSD's vmd is a lot more appealing than bhyve is today.
I saw recently that AMD SEV (Secure Encrypted Virtualization) was added, which seems compelling for Oxide's AMD based platform. Has Oxide added support for that to their bhyve fork yet?
Being that vmd's values are aligned with OpenBSD's (security above all else), it is probably not a good fit for what Oxide is trying to achieve. Last I looked at vmd (circa 2019), it was doing essentially all device emulation in userspace. While it makes total sense to keep as much logic as possible out of ring-0 (again, emphasis on security), doing so comes with some substantial performance costs. Heavily used devices, such as the APIC, will incur pretty significant overhead if the emulation requires round trips out to userspace on top of the cost of VM exits.
> I saw recently that AMD SEV (Secure Encrypted Virtualization) was added, which seems compelling for Oxide's AMD based platform. Has Oxide added support for that to their bhyve fork yet?
SEV complicates things like the ability to live-migrate guests between systems.
If I were Oxide, though, I’d be sprinting to seamless VMWare support. Broadcom has turned into a modern-day Oracle (but dumber??) and many customers will migrate in the next two years. Even if those legacy VMs aren’t “hyperscale”, there’s going to be lots of budget devoted to moving off VMWare.
Broadcom also isn't all that dumb, VMware was fat and lazy and customers were coddled for a very long time. They've made a bet that it's sticky. The competition isn't as weak as they thought, that's true, but it will take 5+ years to catch up, not 2 years, in general. Broadcom was betting on it taking 10 years: plenty of time to squeeze out margins. Customers have been trying and failing to eliminate the vTax since OpenStack. Red Hat and Microsoft are the main viable alternatives.
Linux has a massive advantage where it comes to hardware support for all kinds of esoteric devices. If you don't need that, and you've got engineers that are capable of patching the OS to support your hardware, yep, have at it. Good call.
* https://www.youtube.com/watch?v=UvEKSqBBcZw
Certainly they already had experience with ZFS (as it is built into Illumos/Solaris), but as it was told to them by someone they trusted who ran a lot of Ceph: "Ceph is operated, not shipped [like ZFS]".
There's more care-and-feeding required for it, and they probably don't want that as they want to treat product in a more appliance/toaster-like fashion.
There are different levels of scalability needs. CERN has over a dozen (Ceph) clusters with over 100PB of total data as of 2023:
* https://www.youtube.com/watch?v=bl6H888k51w
Certainly there are some number of folks that need more than that, but I don't there are many.
> Like Ceph it is also vulnerable to single points of failure.
The SPOF for ZFS is the host (unless you replicate, e.g., zfs send).
What is SPOF of Ceph? You can have multiple monitors, managers, and MDSes.
Do you know how many administrators CERN has for its Ceph clusters? Google operates Colossus at ~1000x that size with a team of 20-30 SREs (almost all of whom aren't spending their time doing operations).
You can also tell Ceph to use a single disk as your failure domain. No one does that either. Homelabbers maybe, but then why are you comparing such setups with Google?
We run Ceph with a failure domain of an entire rack. We can literally take down (scheduled or unscheduled) an entire rack of 40 servers, and continue to serve critical, latency sensitive applications, with no noticeable performance loss.
We have a Ceph footprint 5x larger than CERN run by a team of 4-5 people.
What?
> A Ceph cluster must contain a minimum of three running monitors in order to be both redundant and highly-available.
* https://docs.ceph.com/en/latest/glossary/#term-Ceph-Monitor
> Our Configuring ceph section provides a trivial Ceph configuration file that provides for one monitor in the test cluster. A cluster will run fine with a single monitor; however, a single monitor is a single-point-of-failure. To ensure high availability in a production Ceph Storage Cluster, you should run Ceph with multiple monitors so that the failure of a single monitor WILL NOT bring down your entire cluster.
* https://docs.ceph.com/en/latest/rados/configuration/mon-conf...
In my experience you need something like GlusterFS which I wouldn't call "light".
Oxide is shipping an on-prem 'cloud appliance'. From the customer's/user's perspective of calling an API asking for storage, it does not matter what the backend is—apple or orange—as long as "fruit" (i.e., a logical bag of a certain size to hold bits) is the result that they get back.
> To mitigate all this, we’re intending to stick with the OSS build, which includes no CCL code.
[0] https://news.ycombinator.com/item?id=41256222
[1] https://rfd.shared.oxide.computer/rfd/0110It is so sad that we've ended up with designs where this is the case. There is no intrinsic reason why nested virtualization should be hard to implement or should perform poorly. Path dependence strikes again.
That said, it does add a lot of complexity.
That's with virtio, the virtual intel "card" is even slower.
They went with Illumos though, so curious if the poor performance is a FreeBSD-specific thing.
[0] https://code.fizz.buzz/talexander/machine_setup/src/commit/20768edcf69eddae5cf65e30a0bde869f9ddd19b/ansible/roles/bhyve/files/bhyve_netgraph_bridge.bash#L214
[1] https://code.fizz.buzz/talexander/machine_setup/src/commit/20768edcf69eddae5cf65e30a0bde869f9ddd19b/ansible/roles/firewall/files/mrmanager_pf.conf#L35But I think you might be right about something because, playing with it some more, I'm seeing an asymmetry in network I/O speeds; when I use `iperf3 -R` from the VNET jail to make the host connect to the guest and send data instead of the other way around, I get very inconsistent results with bursts of 2 Gbps traffic and then entire seconds without any data transferred (regardless of buffer size). I'd need to do a packet capture to figure out what is happening but it doesn't look like the default configuration performs very well at all!
I know NetApp (stack based on FreeBSD) contributed significantly to Bhyve when they were exploring options to virtualize Data ONTAP (C mode)
https://forums.freebsd.org/threads/bhyve-the-freebsd-hypervi...
The section about Rust as a first class citizen seems to contain references to its potential use in Linux that are a few years out of date; with nothing more current than 2021.
> As of March 2021, work on a prototype for writing Linux drivers in Rust is happening in the linux-next tree.
Bryan Cantrill, ex-Sun dev, ex-Joyent CTO, now CTO of Oxide, is the reason they chose Illumos. Oxide is primarily an attempt to give Solaris (albeit Rustified) a second life, similar to Joyent before. The company even cites Sun co-founder Scott McNealy for its principles:
https://oxide.computer/principles
>"Kick butt, have fun, don't cheat, love our customers and change computing forever."
>If this sounds familiar, it's because it's essentially Scott McNealy's coda for Sun Microsystems.
Frankly I don't understand why they blogged that at all. It reeks of desperation, like they feel they need to defend their choice. They don't.
It also should not matter to their customers. They get exposed APIs and don't have to care about the implementation details.
> It also should not matter to their customers. They get exposed APIs and don't have to care about the implementation details.
Yes, the whole product is definitely designed that way intentionally. Customers get abstracted control of compute and storage resources through cloud style APIs. From their perspective it's a cloud appliance. It's only from our perspective as the people building it that it's a UNIX system.
The reason I continue to invest myself, if nothing else, in illumos, is because I genuinely believe it represents a better aggregate trade off for production work than the available alternatives. This document is an attempt to distill why that is, not an attempt to cover up a personal preference. I do have a personal preference, and I'm not shy about it -- but that preference is based on tangible experiences over twenty years!
I don't think working at Oxide would be for me, but I respect the team's values and process.
[1] https://rfd.shared.oxide.computer/rfd/0001#_shared_rfd_rende... [2] https://github.com/oxidecomputer/rfd/blob/master/src