I'm thinking systems designed based on the assumption that there are tens, hundreds or even thousands of processors, and design assumptions are made at every level to leverage that availability
I'm thinking systems designed based on the assumption that there are tens, hundreds or even thousands of processors, and design assumptions are made at every level to leverage that availability
I'm re-implementing it as a metacircular adaptive compiler and VM for a production operating system. We rewrite the STEPS research software and the Frank code [2] on a million core environment [3]. On the M4 processor we try to use all types of cores, CPU, GPU, neural engine, video hardware, etc.
We just applied for YC funding.
[1] https://github.com/smarr/RoarVM
You are doing God's work. Thank you.
I played with Squeak a bit [1] and several friends like [2] were also active in converting Squeak in (also) a OS.
[1] https://web.archive.org/web/20231205061256/http://swain.webf...
But mainstream servers manage hundreds of processor cores these days. The Epyc 9965 has 192 cores, and you can put it in an off the shelf dual socket board for 384 cores total (and two SMT threads per core if you want to count that way). Thousands of core would need exotic hardware, even a quad socket Epyc wouldn't quite get you there and afaik, nobody makes those, an 8 socket Epyc would be madness.
building better abstractions - kuberenetes is an example, although i certainly hope we dont keep being stuck there - is probably a better use of time
Ultimately, the OS has to be designed for the hardware/architecture it's actually going to run on, and not strictly just a concept like "lots of CPUs". How the hardware does interprocess communication, cache and memory coherency, interrupt routing, etc... is ultimately going to be the limiting factor, not the theoretical design of the OS. Most of the major OSs already do a really good job of utilizing the available hardware for most typical workloads, and can be tuned pretty well for custom workloads.
I added support for up to 254 CPUs on the kernel I work on, but we haven't taken advantage of NUMA yet as we don't really need to because the performance hit for our workloads is negligible. But the Linux's and BSD's do, and can already get as much performance out of the system as the hardware will allow.
Modern OSs are already designed with parallelism and concurrency in mind, and with the move towards making as many of the subsystems as possible lockless, I'm not sure there's much to be gained by redesigning everything from the ground up. It would probably look a lot like it does now.
On SGI, the CrayLink HW had perfomance counters visible via the Performance CoPilot (nee PCP).
On Linux, NUMA arch has similar things (numastat, Intel's PCM, other tools). Depending on the workload, it may matter, but if the OS/tooling does not expose the counters, it isn't even possible to quantify the impact.
SGI's IRIX, due to the sheer physical size of their larger ccNUMA systems (AFAIK, AMD's NUMA is from SGI's ccNUMA), had the option to auto-migrate the workloads when certain CPU to working memory latency thresholds were reached.
Plan 9 is very close. Its kernel was designed from the start with multi-processing and is all channels internally. It's small and highly portable so it runs on a Rpi and there was an IBM Blue Gene project. It's architecture is distributed so one machine can serve disk and auth and a hundred more can boot and auth to them. Then they can share resources seamlessly as everything happens through the foundational 9P protocol.
There was never a framework or API to fully abstract this for HPC but could be implemented with minimal effort as the foundation is there (e.g. You wouldn't need mpi as you have 9p.) Though general work loads can be pushed to CPU servers using ssh like commands to do thin me line run programs in pipes across machines.
It's not a true OS--but it's a platform on top of an arbitrary number of nodes that act as one.
The cool thing is that from the program's perspective you don't have to worry about the distributed system running underneath--the program just thinks it's running on an arbitrarily large machine.
Chuck Moore (of Forth fame) and his company springs to mind. They make a chip with "144 independent computers" on it. I'm not sure if that is totally orthogonal to what you're saying, but there it is, just in case!
https://kerlabs.com/ https://en.wikipedia.org/wiki/Kerrighed