The next operating system: building an OS for hundreds, thousands of cores
web.mit.edu
web.mit.edu
It feels to me that is the next iteration of refinements of technology that is 40 - 50 years old. Caches are what, from 1970? Interconnect issues date from the same time. I think machine cycles are in abundance, the scarce resource is interconnections for data flow. So one of the first things you want to do is organize your data flow so that processing is local. Think simulations like vision processing, weather prediction, rendering, where a processor can work locally and pass on a reduced amount information to its neighbors. The interesting problems arise when the results have to be delivered non-locally. If you store them in main memory for the recipient to pick them up, you run into bandwidth problems.
So what I see as needed is gazillions of low level worker bees with modest bandwidth requirements that have semi-permanent connections to the consumers of their output. Think the human brain, Google search, image rendering.
Apologies for the rambling, lack of citations, etc. etc, but I am interested in HNer's views on these issues.
However I agree that it'll likely produce interesting research advances and usable components, and some up-front overselling is probably unavoidable... saying something like, "we're initiating a collection of research projects to produce technologies that will be needed for a future generation of operating systems" isn't as good PR as saying, "we're building the next-gen operating system".
The ideas failed in a relative fashion. They failed to yield that many result relative relative to the effort and resources that were put into them. And they especially failed to yield as many results as you could get by just increasing raw clock speed. And their cost was not just their raw complexity but the training required for programmer to understand parallelism (the low cost of today's "entry level" programmer is a huge bogus to the IT industry. If companies had to spend a year on system-specific training, the cost would be vast).
But given that such projects were only failures relative to the alternative of just coming up with a simple architecture with a higher clock speed, if the alternative is going away, there's no reason they can't become relative successes. Watson was a relative success - if you could have Watson-level processing-power on a single chip programmable with Ruby, the effort IBM put into the project would look silly. But since it looks you can't, Watson seems like a productive use of resources.
In a lot of ways, the last thirty years have involved substituting low-cost, high-power chips for high-price, high-skill programmers yielding a huge, de-skilled programming workforce. The end of Moore's-law-for-speed would seem to mean things will move differently in the future.
I won't address the particular architecture points you make, which all sound good but are unrelated to the general of previous parallelism efforts "failing".
For the criticism that the project will "fail", I think there's some merit. Will this produce the next computation system that the world uses? Probably not. Will this push the boundaries of engineering and improve on the state of the art? Almost certainly. And due to its scale, and the fact that Angstrom was created by combining multiple separate projects, there's also room for individual projects or ideas to succeed even if others don't.
My background is that I worked on the FOS project, which is the part that is most related to the title of this post. If you were looking for more technical meat, I encourage you to check out the project page: http://groups.csail.mit.edu/carbon/?page_id=39
http://www.tilera.com/about_tilera/press-releases/linux-appl...
( June 2009, Erlang (BEAM) on Tilera 64-core )
Go might be feasible, but forcing system-wide GC at random times for the entire system? GC is very hard to make concurrent and a single random-alloc GC'd memory space can't possibly scale to thousands of cores.
Since a kernel is something that is expected to run forever, it can't afford to leak anything over the long term. For most GCs, collecting that last little bit of garbage (in deterministic time) requires O(committed address space) memory bandwidth. A full-copy style GC may take O(object memory), which could be an improvement.
Now that memory and applications are routinely many gigabytes, this is a big deal.
It's hard enough for an ecommerce web server to maintain responsiveness, I couldn't imagine trying to respond to hardware IO interrupts in real time while running a collector like that.
Erlang (and its way) is a much better fit there, I think it'd be a delightful apps language: the GC runs at the (erlang) process level, each process has its own heap, so even though the GC is a vanilla generational GC by the magic of the Erlang VM it turns into a highly concurrent pauseless GC (only needs to pause a single Erlang process at a time, and you generally have tens of thousands chugging along).
Not so much as you might think. Both Smalltalk and Lisp were OSes early on. If you dig around in some early Smalltalk images, you'll find the 4 stubs for "put the drive head down" "pick the drive head up" "move the drive head out" "move the drive head in."
In all seriousness, when you have to support hardware controllers whose only interface to the CPU is a memory block or I/O instructions, what other choice do you have besides C? I guess you have C++...
*Assembler
*Forth
*Any language which allows inline assembly.
*Any statically typed language compiled to C: Haskell, Ocaml, Go, etc. *Any language that can reference memory locations directly.Using a higher-level language can also help with many other things, for example integrated garbage collection, built-in serialization and process migration, intelligent message passing (pass objects around rather than bytes), etc.
A good example of an implementation of these ideas is Singularity (http://research.microsoft.com/en-us/projects/singularity/), which is written in a superset of C#.
I cannot apprehend the confusion of ideas that would lead you to believe this is caused by C. It's assembly, not C, that runs on your machine, and assembly is not memory safe. Microsoft's Singularity runs managed code, so there's a software layer providing memory protection.
An operating system is not defined by its preferred programming language.
There is no confusion of ideas, but I like the Babbage reference :)
In Singularity the code is compiled first to CIL, then to x86, x64, or ARM by an AOT (ahead-of-time) compiler. Now here's the thing: the OS loader does not load Assembly code, it loads CIL code, which it then compiles further down to Assembly. Since CIL is verifiably memory safe (like Java bytecode), and assuming the AOT compiler is not buggy, the OS is memory safe. Hence no need for MMU/MPU to do memory protection.
So even though you're right that Assembly is not memory safe, the final compiler stage (in this case CIL can be seen as an intermediary language) is implemented in the OS (the loader), and the OS will reject any non-memory safe code. No code that can write outside array boundaries, violates type safety, and attempts pointer arithmetic will be allowed to execute.
So no, Singularity does not provide a software layer for memory protection, it's purely an (intermediary) language feature.
The nice thing is that these protections can be applied ahead-of-time rather than at runtime, as the MMU does.
It has been shown (see Native Client) that even a slightly restricted subset of x86 assembler can be made verifiably-safe. In principle, you could do the same thing with C - in fact, you likely wouldn't even have to change the language definition, since the egregious abuses in C mostly result in "undefined behaviour", which allows the implementation to detect and trap them rather than hose the machine. (It is a curious fact that C itself isn't inherently "memory unsafe" - merely almost every implementation of C ever written is).
If you can come up with a program that verifies whether a C program is memory safe or not (without constraining the language spec), well, hats off to you sir, I think I know a couple of security experts who might want to have a word with you :)
Does the x86 still use a page-based memory model? I was under the impression it went to a full range addressing model with the P6.
Page-based memory was what stopped me learning x86 assembly back in the day. I loved the Z-80 (and Rodney Zacks).
Compiled, concurrent, and memory-safe, with no global GC.
(Full disclosure: I work on Rust.)
Cyclone would be another candidate: http://cyclone.thelanguage.org/
Cilk would be another interesting option, with its built-in support for concurrency: http://software.intel.com/en-us/articles/intel-cilk-plus/
That's exactly what Microsoft Azure claims to be in their marketing literature.
(I think it's wise to pay attention to how such forces change language.)