Xv6
en.wikipedia.org
en.wikipedia.org
Grokking xv6: http://experiments.oskarth.com/unix00/
What is a shell and how does it work?: http://experiments.oskarth.com/unix01/
What's on the stack?: http://experiments.oskarth.com/unix02/ (video of tracing a system call from user space to kernel space and back: https://www.youtube.com/watch?v=TWksEdn5eoA)
Page tables and virtual memory: http://experiments.oskarth.com/unix03/
Locks and concurrency: http://experiments.oskarth.com/unix04/
A short overview of the file system: http://experiments.oskarth.com/unix05/
Grok LOC? http://experiments.oskarth.com/unix06/
It was a very educational experience for me, and I highly recommend the journey for other people.
In your original post, you mention one of the goals being able to contribute a patch to a modern OS such as BSD or Linux. After your experiment, do you feel like this is the case?
I feel like I could after maybe ~30h (really I don't have enough information to determine beyond it probably being between 10 and 100 hours) of focused effort. I did read a bunch of The Design and Implementation of FreeBSD and I felt like I could understand it reasonably well, using the mental models I picked up studying xv6.
The reasons why I didn't do it (yet) are interesting in themselves. There was nothing in particular that I felt was missing from FreeBSD, and I hadn't used the "advanced" parts of the OS enough to run into any bugs. Because of this I think I lacked a sense of ownership, and the artificiality of a bug hunting activity caught up to me, motivation-wise. To do this I think you need to be motivated to dive into the nitty-gritty of a particular part of the OS, in addition to having a good understanding of the OS as a whole. Instead I ended up exploring some of the more advanced parts of FreeBSD (Jails and ZFS) with http://experiments.oskarth.com/netpowder/ - hopefully I will find something useful that I want to either extend or a bug I want to fix, or perhaps I get to test the hypothesis with something completely different (such as Mirage OS).
xv6 is pretty awesome for learning, and lecture is generally spent reading through the source code, some of which is...cute.
Most of us have the source code printed out. If you're interested in reading the source, this document is well-formated: https://pdos.csail.mit.edu/6.828/2014/xv6/xv6-rev8.pdf (and can actually be generated by the Makefile!)
I only ask because I have no idea how much is the kernel compared to everything else; and if it has its own custom compiler or whatever.
[0] http://www.ccs.neu.edu/course/cs3650/unix-xv6/HTML/S/64.html [1] http://www.ccs.neu.edu/course/cs3650/unix-xv6/HTML/S/42.html
For my final project I implemented a simple threading library based on the interface that pthread() uses. It's amazing how beautifully simple kernels can be.
Unless I'm missing something, the entire codebase is just 102 files (including README and such), no subdirectories, and 12,000 lines.
What a fantastic tool.
* Trace the execution of one of the user land programs (e.g. ls) to see how it uses the syscalls.
* Add a new syscall - there are a few open-courseware labs online which have this as in introductory exercise (e.g http://moss.cs.iit.edu/cs450/assign01-xv6-syscall.html)
* In a similar vein - find some other labs for other courses using this - and follow them.
* Read through the companion book (can be compiled from the source or downloaded pre-compiled) and follow it through the source code.
The Linux kernel itself is really well documented. Between the docs you can pull, KernelNewbies.org, the IRC channel, and the O'Reilly text "Understanding the Linux Kernel, 3rd ed", you can get up to speed re: conventions and develop a pretty good top-down view.
Then use strace/ltrace excessively on everything to get a bottom-up view. Strace is a great skill to have in general to help identify bugs in live code by attaching to the PID. Way more granular than a standard GDB attach. It's got a learning curve about on par with reading pre-C++11 template error msgs, but after a few months you'll start to recognize patterns, just as in tmpl error msgs, to the point where you can just skim the last few hundred lines strace piped to stdout and you'll know what sort of bug to deal with.
I haven't played with FreeBSD in nearly a decade but I do remember their documentation was second to none, including the IBM Red Books. So if you want to see another approach to a POSIX implementation (obviously Tanenbaum talks about MINIX which is semi-POSIX, so yeah, read that supplementary text), going through the Handbook from an end-user perspective then the mailing lists + source will give you a real interesting view as to why certain engineering decisions were made. I.e., why ipfw was replaced, the gradual progression of standard file systems from UFS all the way up to modern day ZFS, etc. The list offers a rare view behind the curtain exploring engineering decisions that end-users aren't often privy to, and the caliber of conversation is ridiculously high
You can find it here: https://github.com/NewbiZ/xv6
Considering the original build system, I think having a look at this just for the sake of the cleaner makefiles is worth it.
There is also a mailing list for xv6, to share its understanding: http://www.freelists.org/list/xv6
I'm not sure what I'm supposed to take away from your comment aside from "I went to MIT".
edit: here's an analogy - say the article was about a famous restaurant, and someone chimed in to say "i ate there when i was visiting hong kong, and it was one of the best meals i've ever had". now we already know that it's a good restaurant; that's what the article was all about after all. nevertheless, i like that sort of comment; it adds a human touch to the story because (rightly or wrongly) i perceive a fellow commenter as less remote than a newspaper food writer, and therefore their opinion has a certain anecdotal quality to it that the article lacks.
1. Class 828 was a really good class, best he took. Others might want to try it too!
2. There's another operating system that's used in the class called JOS you might like to check out also if you are into learning about operating systems.
2. b) JOS is similarly quite simple to understand and doesn't take that long to build. Try it out also.
Hope this is helpful!
As useful as it is, especially outside of i386, QEMU has grown to be a massive program with many dependencies that requires GB of RAM/swap during compilation.
If you want to write an oh-cool blog piece for general programmer audiences about how OSes work, sure, go ahead and write it in Python (seen that done, and pretty well, in fact).
But if you want to prepare people to be able to work on real operating systems, they need to know how to do it in C.
Most other languages hide away implementation details that are critical to writing an operating system--how would you handle, say, interrupts?
In addition, you can annotate a protected object (essentially, a group of shared procedures protected by an implicit mutex) as being interrupt-safe, and then any access to that object will be automatically protected by the appropriate instructions.
Plus, if you're in a Posix environment, you can use the exact same mechanism for interrupt handling. It's all remarkably elegant.
Alas, like everything Ada, the documentation is opaque in the extreme, but:
https://www2.adacore.com/gap-static/GNAT_Book/html/aarm/AA-C...
Note that at the bottom they're defining a parameterised interrupt handler structure and then instantiating it multiple times on multiple IRQs, each of which is in its own isolation domain...
Sometimes education is about pure theory, and sometimes it's about how people do things out in the world. I think this is a case where the second approach is much more valuable.
Are you implying that another language would be more clear?
Quite the opposite, C doesn't hide anything from you. When you want to understand what the computer is doing you can tell directly from the C code. Unlike other languages where you need to understand what the language is doing first, and only then can you understand the computer.
There are time where "hiding" the computer is useful, but not here.
Yes, it does.
C hides cache, SIMD, registers, the stack (no multiple return values for you!), the details of the heap (malloc() either succeeds or fails, and you can't know what it's going to do until you call it), SMP, instruction-level parallelism, and the details of atomicity, all of which are relevant to OS programming.
C is a nice language. Don't pretend it's how the hardware really works.
C is a nice wrapper around assembly. In a few cases, you wish it were a bit better specified to control the actual assembly/ABI binding a bit better. The problem I see is a lack of contract enforcement which makes introducing hard to diagnose bugs really easy. Maybe Rust will fill the gap.