It's Time for Operating Systems to Rediscover Hardware
usenix.org
usenix.org
An anecdote I like to share is that while I was TA'ing CMU's 15-410 operating systems course (a great course - build a preemptive multitasking operating system from scratch) was that the student projects could run on real - if old - hardware. There was a PC that could boot your stuff if you put it on a disk.
This PC had to have PS/2 interface to the keyboard, though. A newer PC would be all USB. Apparently, the complexity of the codebase to talk to the USB device was around the similar level of complexity to the entire preemptive multitasking operating system the students were building (and, of course, considerably less educational). I mean, this wasn't a full-functioned OS, but it allowed multitasking and preemption of kernel calls, so seriously non-trivial.
Multiply this story by all the devices a reasonable OS would be expected to talk to and it's a scary prospect for the OS researcher... and if you don't have those devices up and running, good luck supporting most workloads anyone cares about.
Plus, if we ever are to transition to the world of microkernels (something that OS researchers have been pushing for decades now), we're gonna need a massive amount of people writing drivers just to get us to the point where things are barely functional.
Reading from a keyboard is much simpler. Probably makes more sense to run the computer headless with a serial terminal and make the students poke at a UART. PC standard UARTs are pretty easy, and I think it's easier to find a motherboard with a COM1 header than a PS/2 port (at least my latest hardware matches that description).
PCI is nothing in comparison to the complexity of USB, especially since the BIOS will usually have set up the I/O ranges for the devices already. It's literally a few dozen bytes of machine code to scan the PCI(e) bus for the device you want.
Conversely, the work they did in kernels has a huge amount of transfer to thinking about concurrency, which is truly valuable - and deep. It also meant that they acquired a far less magical idea of where things like processes and threads come from, having actually built the mechanisms to make them happen (you build a thread library in a warm-up project).
As for microkernels - they often punt a lot of the hard concurrency stuff up to user space servers, which will wind up needing exactly the kind of concurrency that the students learned to build.
I agree you don't learn deep principles from it, but I disagree that it's not broadly applicable based on your description. Grinding away with spec in hand, "one damn thing after another" sounds exactly like most programming people are likely to encounter in their career.
Also, this course's OS kernel presumably also has a spec, and implementing such a kernel is also "simply going through the spec".
I think the point you're trying to make is that the contents of the spec have be relevant to an operating systems course, and most hardware specs are maybe only tangentially related to the kind of information you have to teach. I'm not sure that's fully true either though, because isolation and safety are core OS properties, eg. DMA has all kinds of security implications. Maybe not stuff for a beginner OS course, but hardware interfacing is critical.
Sure, but that's not the business of a university course to teach.
> Also, this course's OS kernel presumably also has a spec, and implementing such a kernel is also "simply going through the spec".
Per some of the other discussions, this doesn't seem to be the case (i.e. the spec that was offered to students seems to have been quite open ended), as people mention that a huge challenge as a TA was in adapting to the large diversity of approaches that students were trying.
There is also a huge difference between a spec that says "the task scheduler must accept tasks in this format and ensure they are scheduled fairly with an O(n) algorithm" and a spec that says "to enable PVM set bits 1 and 3 in register 7; to issue a new read cycle clear all bits in register 21". Device drivers deal mostly with the latter, unless you move to extremely complex devices like video card drivers, that are probably way more out of scope than a simplistic kernel.
But what you propose is preprofesional drival. It's bad enough that education is subsidized for employers already, we don't need to stoop to them further.
That class bordered on legendary. Awesome that you got to TA it. I still have an irrational fondness for AFS.
I believe that the course number changed at some point - or some shallow aspect of the course was set to change - and a bunch of industry people called the university to express concern. It is/was (I can't speak for the current iteration of the course, although I have no indication that quality has slipper or anything) truly transformative.
I offered to run a 'franchise' of the course at the University of Sydney a while back and was informed that anything quite that transformative wasn't really an option; our job was at least in part to pass engineering students who didn't care that much about computing.
Thats very disappointing. You might have more luck joining the University of New South Wales OS team. (For context, UNSW is a local rival of USYD).
When I was a student there, the OS course was legendary. I only did the "basic" OS course. Our assignments led us to implement a bunch of syscalls for handling file IO, and write a page fault handler for a toy operating system. I'm not sure what they do in the advanced OS course - but it has a reputation for a reason. And it looks like[1] its still run by Gernot Heiser, who's a legend. He's the brain behind SeL4 - which is the world's first (and only?) formally verified OS kernel.
I'm kicking myself for not doing his advanced OS course while I was a student.
https://news.ycombinator.com/item?id=17692499
That said, BIOSes can also emulate PS/2 with USB (no doubt via SMM). This is usually a setting called "legacy keyboard/mouse support".
[1] The notable exception being Apple. Think Different. https://news.ycombinator.com/item?id=12924051
That used to be true, but the PS/2 is so much slower (lower frequency) that by the time a package containing the mouse or keyboard info is finished transmitting the usb port started and finished a poll
"Hardware" (read: device drivers) is not notably complicated IMO.
That's not to say they're simple, but driving devices from a software point of view is pretty similar to interacting with other software. You read the spec (APIs), write code to set up data structures or registers a certain way, parse responses, etc.
Writing a modern full fledged USB stack is very complex, but would not be more complex than writing a modern full TCP/IP stack for example.
A lot of device programming models have gotten simpler too. Device drivers used to be notorious deep voodoo magic e.g., in cases of the IDE disk driver, but that was not really "complexity" of the software logic so much as hundreds of special cases and deviations from standards or odd behavior in all manner of IDE controllers and devices. To the point where the real wizards were the ones who had access to internal data sheets, errata, or reverse engineered firmware from these devices, or otherwise spent countless hours poking at things and reconstructing quirks for themselves, probably bricking a lot of silicon and spinning rust in the process.
But systems and devices are now getting to the point where things don't work that way anymore, silicon power is so cheap and firmware is on everything. The NVMe device for example sets up pretty simple packet-type command queues and sends requests and receives responses for operations - query device information, perform a read or write, etc). It's not quite that simple, there are some quirks and details, but it's quite a lot like writing a client for a server (HTTP, perhaps). I think other device interfaces will evolve toward cleaner simpler models like this too.
One exception to this is GPUs of course. The other thing is with moore's law continuing to slow down there will be increased incentive to move more processing out to devices. It's long been happening with high end network devices, but with technologies like CXL.cache coming soon, I would expect that trend to keep ramping up and come to other accelerators (crypto, AI, even disk).
So.. it's a mix. Things are definitely getting more complex, but AFAIKS the complexity of interacting with hardware is not increasing at a much greater rate than the complexity of interacting with other software. And, as always this complexity is made possible by abstractions and layers and interfaces. It's not clearly exceeding our ability to cope with it.
I said the software interfaces to them have not exploded in complexity, i.e., specifically refuting the suggestion that we don't "have the resources to cope with the monstrous complexity of hardware now", from the point of view of driving them with software. You utterly failed to refute anything that I actually wrote, but after your tantrum is over feel free to have an attempt.
And I work at the line and on both sides of it logic, architecture, and software, so don't bother with the vapid appeals to authority.
Utter rubbish. And he goes on to talk about programming for an embedded SoC, etc. Hah. Some of the most utterly wretched code, vhdl, and development practices I've ever had the displeasure of being acquainted with have been in embedded devices, devices, device driers, firmware, etc.
And first level interrupt handlers? He's saying that like it takes some genius to do it, or the process of doing it confers some deep understanding on the writer. It doesn't. It's not particularly complex, outside buggy or stupid architecture but even that's just grinding work.
And the there is some horrific buggy crappy interrupt handler code around and I can say that because I've written some. And it sounds like you may well have too, if your interactions with David Miller relating to the the performance of Sun OS system calls is anything to go by.
Arguably it's more valuable to understand how the C environment (stack, etc) is set up and how to wrangle the toolchain into emiting code at particular locations and such. But you can get all that many other ways. And even then, I don't actually agree that you need to know the minutiae of those details in this day and age which is a great thing. This is not where most of the interesting work is happening, like it or not.
This kind of elitist gatekeeping attitude is just sad. It reeks of the has-been (or maybe never-was) mindset. I would just as well trust code written by someone who deeply understands what they do writing a word processor or video game or distributed database, than someone who has hacked out an interrupt entry vector. And the sad fact is that being an OS developer does not require or confer a deep understanding of how to get the most out of hardware, understanding caches or multiprocessing or branch predictors or performant scalable data structures and synchronization techniques. No more than writing a web server prevents a person from understanding all those things deeply.
Yes, it’s hard to do research if you start at “let’s design a kernel from scratch,” but you’d never need to do that. You can just hack Linux or a bsd, or even use something like bpf to extend it.
The thing that annoyed me about 410 is that when I switched to Linux I realized that pusha/popa didn’t matter at all. It was a good course for writing reentrant C and learning the very basics of hardware, but the really hard stuff is in the weird dynamics of memory management on NUMA systems and work conserving io, which you can’t get anywhere near if you are starting from scratch.
Why not? Linux has plenty of core design flaws that can't be solved without starting from scratch.
So you wind up with a research operating system that has finger-painting graphics, can't talk to modern SSDs or networks cards... or you somehow conjure up the resources to build an enormous number of drivers.
If you "hack Linux or a bsd" - or even more timidly - use ebpf - the constraints of what you build are going to result in you basically building "more UNIX stuff". This isn't bad research, but the sheer mass of constraints you're taking on if you accept all the design tradeoffs of Linux/bsd/eBPF/whatever means you're certainly not doing the sort of OS research that the original author talked about.
For the complete USB stack, that's very likely. However, for minimal code needed to read keys from keyboard USB shouldn't be that bad.
These students only need USB 1.1. They don't need to support bulk transfers (USB protocol has 2 distinct transports, one low latency another one high throughput). They don't need to support USB hubs between computer and the keyboard. The only device type they need is HID.
USB consortium wouldn't allow to call that protocol USB due to the missing features, but it's good enough for education purposes.
BTW, USB specifies a small subset of HID interface called “boot interface”. It’s normally consumed by UEFI firmware who needs mouse and keyboard for the setup GUI which runs before OS launches. The subset only supports basic keyboards and mice.
I think for the USB in that OS, a C API with intentionally very narrow scope would be OK for the job.
The main reason why the real-life code is so complex is that scope being very wide. We have USB 2 and 3, mass storage, two-way audio, cameras and GPUs, hubs and composite devices, numerous wireless protocols on top, OTG, power delivery, power saving features, and more.
If we were talking about a serious project there are a lot of best practices that should be rarely (and then only carefully) questioned, but for something intended to be thrown away at the end of the semester challenging best practices is a great way to learn.
That said, PS/2 is a simplifying assumption as well.
It would certainly be a more beautiful system that I would love to see, but I need help justifying the engineering cost. I hope to see some open hardware passion projects along these lines, but I doubt they will ever be mainstream.
The whole combined experience just makes things run so smoothly and allows security issues to be patched since Apple has the resources that maintain each part internally.
Though maybe we'll have a RedHat of hardware someday.
One method could be, say, extending microkernel capability systems to incorporate these remote processors. Or perhaps even the ability to provide WASM or ePBF control blocks to customize the power management system. Though its true that'd require more engineering resources.
Actually USB is somewhat like this in that there's a hardware api and spec that the OS knows about. Linux has troubles with even that in my experience (often requiring a reboot to fix.
Maybe, instead of opening up all that complexity, simply making better ways to fake a dumb component would be more effective. (That is with verified, memory safe parsers and API interfaces. Then it doesn't matter what other gunk it has behind that interface.)
And this could be implemented as yet another managed layer between the main CPU that runs user space and the hardware SoC. (Basically a low-level "WAF".) Yeah, it's a herculean task, but at least we have access to the Linux drivers.
You have good questions in your comment, and I think he is calling out for searching answers to them and other similar questions. Apparently, the operating system researchers haven't been doing this enough, or at all. His speech is a lament for this state of affairs, and a call to get excited about finding out what benefits would be gained by doing things (very) differently.
But that's just my take, and I'm not working on operating systems research.
That's what basic research does. It discovers new things, and new ways of doing things.
On the radio side of things, Dash7 firmware will run on many Lora devices, meaning you don't have to use the proprietary Lora software.
I think a “true” OS would run its code on all these processors, implementing both sides of the communication protocols connecting these pieces of hardware into a computer.
For one, such hypothetical OS may upgrade or replace these communication protocols with OS updates, or because user changed some system preferences.
This also allows to re-configure these things based on runtime information. For instance, when AC power is disconnected from a battery-powered device, it makes sense to do stuff slightly differently, like stop all processes on all CPUs who are polling to get events faster, and switch to slower but energy efficient interrupts.
Actually, come to think of it extending DTS syntax to describe the pointer swizzling the OP video mentions. Then both OO'es could derive correct physical address handling.
1: https://www.linuxfoundation.org/blog/the-power-of-zephyr-rto... 2: https://docs.zephyrproject.org/latest/guides/dts/index.html
I wasn't around when the decisions at my employer when the decisions made around OS options, but this would've been a viable one if it was mature enough at that time. Thanks for the info!
P.S. huh looks like Linux and Zephyr _are_ doing this via RPMsg subsystem as well. Wow, much nicer than the last time I looked into it!
> I don't see how this might be changed. Most of those little red circles are not just "tiny computers" but also different implementers in different companies, all trying to protect their little fiefdom against their customers
There are of course initiatives like libreboot/freeboot/whatever, and efforts to reverse-engineer the firmware of microcontrollers in peripherals, and there exists fully documented hardware such as the Raspberry Pi 4 (I think?), but even with full documentation, it's incredibly hard to write custom firmware, and the benefit is very slim (better security perhaps?)
The problem is the proprietary nature of this stuff and the fact that a lot of it is outside of the scope of scrutiny that goes into Linux. It's held together with what basically amounts to glorified duct tape. It's complicated. It has weird failure modes and occasionally this stuff has expensive issues. Like for example security issues.
Very relevant if you want to build an open source phone platform or laptop. Not impossible; but it requires addressing a few things. Like connecting to a 5G network without relying on Qualcomm's proprietary software and hardware.
The real solution is not merely integrating these things as is and trying to reverse engineer them but engineer them from the ground up to work together and make sense together. That larger, open operating system does not really exist. And many proprietary equivalents on the market right now have an emergent/accidental design rather than something that was designed from the ground up.
I remember how it was with programmable GPUs.
When programmable shaders were introduced in GeForce 3 (2001), nobody really understood what to do with them. The consensus was that it’s incredibly hard to write custom shaders. OpenGL consortium refused to support shaders for almost a decade (GLSL 1.0 released in 2008), that’s one of the reasons of the success of Direct3D https://softwareengineering.stackexchange.com/a/88055
Yet look at them after 20 years of evolution. People no longer coding shaders in assembly, we have quite a few higher-level languages for them. People have built better APIs to interface with programmable GPUs, both low level close to metal (D3D12, Metal, Vulkan) and high-level (CUDA, TensorFlow, game engines). People built debuggers and profilers for these shaders. Overall, people embraced programmable GPUs, leaned to use them, and the results are amazing. Visual quality of games and other real-time 3D skyrocketed, despite the increase of pixel count and density in the displays.
I can see how the same path is possible for the rest of the programmable chips and pieces inside modern computers.
> and the benefit is very slim
Better products in very broad sense. For instance, when idle, a smartphone OS could power off not just the screen, but also CPU, RAM, and a few other components, and downscale itself to the slowest and the most power efficient core in the system.
Don't the proprietary versions already do that, though?
With different OS design it might be possible to power off more components of the phone when idle, while being able to wake up fast enough.
One one hand there is great vertical disintegration, on the other hand, even Apple is too lazy to actually do this right, still treating their software stack as a black box where breaking changes are fine but elegance is lost on the suits.
I suppose the folks at https://oxide.computer/ are smirking if they see this video.
I suspect much of the problem is "wall of confusion" differences: operating systems are primarily concerned with compatibility and reliability while computer architecture seems to be about rapid innovation and change.
I don’t know though if I fully agree with the characterization that Linux has been shuttled off to the corner as chip vendors work around it. It’s more that the computer has actually always been a distributed system and OS/hardware developers have ignored that for a very long time. You wouldn’t say that Linux isn’t the OS for distributed clusters even though a lot of work goes into making those clusters run and that all typically runs in user space outside the OS boundary.
It’s possible that there’s a better model out there that manages to unify things but I’m skeptical. All this code runs on different chips with different clock speeds. Code also resides in drastically different memory spaces and that’s unavoidable - an M3 might be able to load the firmware over PCIE, but it’s not executing out of main memory. Additionally these chips typically have drastically different cache coherency properties.
Now maybe we need to revisit this all holistically, which I think is the actual pitch being made and one I can support. That needs to look like defining interfaces for how this stuff works (& yes, making it possible to integrate into Linux) and getting chip vendors to adopt. Sunk costs are real and trying to rework not just HW architecture but also the entire SW ecosystem can be a dead end. I’d certainly be keenly interested to hear the speaker’s ideas. He’s certainly far more knowledgeable about this space than I am.
There is a message to build your own computers and that's probably what people will do and it could be interesting research for custom servers and future heterogeneous architectures. I suspect people in this industry are already doing it.
Historically, it started as a unikernel called OSv designed to run directly on hypervisors. But they eventually realized that it can run just as well on linux with kernel bypass working and io_uring and is easier to deploy.
Probably future machines should have lots of FPGAs, and easy access to claim and program them. We will need good languages for the latter, although probably C++ will end up being that.
Among other things, advocating for more stable & open standard HW interfaces, rather than OS-level APIs.
There's very little (interesting) you can do when application developers are (for the most part) incapable and unwilling to embrace anything even remotely different from the DOS application model.
The OS is, at its heart, a mapping from that application model onto the hardware. If a hardware feature or characteristic can't be expressed / surfaced within the application model, then there's little to no value in attempting to exploit it.
Someone like at least has internal customers, and a vertically-integrated embedded system can sometimes do better, but general-purpose systems are shackled to some minor variation on the POSIX application model for the forseeable future.
If Linux is running inside a sandbox, it can't secure anything.
-- Dan Ingalls
Or that you open two programs at once and they try to access the same device at once causing some kind of error and likely data loss or system instability.
I don't even want apps to directly have access to the file level interface of my storage but rather an extra abstraction which lets me pick which files it can access.
I build computer kiosks and I agree wholeheartedly. So many devices/peripherals require some sort of custom implementation. Most have some sort of SDK, but not all. Better hope that SDK comes in your tech stack's language(s) of choice.
Why isn't it just data in / data out?
Even printers and scanners are a pain to deal with. Sure, Windows (and TWAIN) abstracts some of the printer/scanner communication, but it all seems forgotten. Like, why is using new hardware in my software so difficult and how come I can't just implement new hardware the same way every time? To me, it's all just data in/data out from my language runtime to the device. I couldn't care less about the implementation details in between.
I thought we had more time..
Cant we global warm postpone this?
Let the next generation sort it all out?
Edit: And it turns out the title in video was actually without capital T and R.