According to cloc run against 3.13, Linux is about 12 million lines of code. 7 million LOC in drivers/, 2 million LOC in arch/, and only 139 thousand LOC in kernel/. (http://unix.stackexchange.com/a/223753)
Edit: Would be nice to have the numbers for the latest Minix release for comparison. Does anybody know how big their core team is? (The kernel is something like 15k LOC IIRC.)
Steve, thanks for all your work on this and with Rust. Your attitude toward teaching and your writing ability go a long way in encouraging people to learn and in having the lesson be worthwhile.
I understand that the amount of lines of driver code comes from the variety of devices. But still, it looks completely unbalanced, when compared to the kernel code itself. So I have some questions here:
1) Shouldn't there be common interfaces / abstractions for most of the devices? 2) If they exist, could they be improved somehow? 3) A bit unrelated, but how fun / interesting is it to develop driver code?
Like, once it seemed a good idea to have it like that and later people found out there are a bunch of things missing. Now many of them have to be implemented on the driver side over and over again.
http://www.usb.org/developers/docs/
Just making a standards compliant implementation takes a huge quantity of code. And with a standard this complex people inevitably get it wrong, so you'll need workarounds for all the broken devices too.
In the end, you understand as much as you can, implement what you have to, and do your best to get through it... and even then, someone will mess things up on some end or another... to this day, I'm surprised that SCORM was synchronous... hell, it feels like they're half the reason XHR has a sync option.
I feel the same way when looking at terminal emulators... sigh, so many things to implement to get something useful, even if you're only looking to get a small subset working.
3. Interesting? A lot. You get to learn how do low-level parts of your system work, deal with complex structures and interactions between kernel/hardware/userspace and.
But the most interesting thing is the programming ability you need. You have to really understand the code, be able mentally follow the execution paths (which are not simple, believe me) and write as few bugs as possible. Why? Because you don't have many tools. No debuggers, no unit tests... My main debugging tool is a set of debug macros that print to the kernel log. Debugging is a real pain in the ass and crashes can lead to a system restart, so I have to try to get it right as soon as possible.
I hope that also answers the "fun" part ;)
2) ... but the meat of a driver is putting the right values in the right registers of some chip, and sometimes work around bugs in said chip, or sometimes take into account that popular variant that's almost the same that the original but not quite... It takes a lot of boilerplate code.
3) It depends on your definition of fun. A driver by definition is a middle man between the OS and the hardware; so programming there is a lot about doing what you're told (and sometimes get punished harshly because you misunderstood something), either by the kernel documentation or by the datasheet of the chip you talk to.
2. Not... really. Standards are supposed to do that. Many hardware manufacturers don't really implement them in a proper manner, and "whoever designed this flash controller should die a slow, painful death" is not something users care about.
3. If the device is well-documented and relatively standard-abiding, it's extremely fun if you're passionate about it and relatively painless if you're not. Otherwise, it's about as fun as digging a tunnel under Mordor with a toothpick.
In the end, people tend to do the best they are able to in a given situation. Often that means ignoring some parts, and making assumptions for others so that your product can ship instead of waiting over a year for a standards body that might clarify something.
Right now it's about 2.
There are a lot of things about machines that you need to understand, but the technology itself is relatively well-understood and sane practices are encouraged.
It blows my mind that less than twenty people can string up something non-trivial with Node.js et co.. Operating systems are trivial in comparison.
As someone who is working on a simple os and a college student, this saddens me.
I'll have to see if my library has one.
You'd do better to look at xv6. That's like v6 UNIX, minus permissions, plus SMP, running on 64-bit x86. It's very clean and modern C99.
Matter of fact, Wirth's people make OS's with type-safe code that are so straight forward I've always recommended starting with an AOS or A2 Bluebottle port to your language of choice.
I'm saying all this in the past tense because the company was sold off a few years ago and I have no idea what happened to the IP, if anyone is still developing it and so on.
I only wanted to mention it because operating systems have this intimidating reputation, which is not only a little unfair, but seems to hold people back. I've seen a lot of very good programmers thinking they couldn't be part of a team that develops one, when it was in fact well within their possibilities.
And, historically, this has been true for a very long time. Many applications that run on operating systems are far, far more difficult to program than the system they're running on, or at least require vastly more knowledge in order to get right than writing the OS.
> Your website is interesting but a tad hard to follow due to the "organization" scheme. ;)
The joy I get out of doing it again vastly outweighs the pain of writing HTML by hand, but barely :-).
"The joy I get out of doing it again vastly outweighs the pain of writing HTML by hand, but barely :-)."
Haha. I can't talk given how disorganized my stuff is. Hell, I don't even have a blog: I have text files and PDF's I send to people that ask. Gopher is on another level compared to how I do things.
Nonetheless, it's good to know people that did embedded systems as I'm slowly accumulating knowledge on that sort of thing in my exploration of secure hardware/software systems. Really wish I spent more time on embedded and HDL's back when my brain worked reliably. That's where best bang-for-buck in security and reliability are. The more people write openly on such topics the better. So, consider organizing anything you have on the magic you used to do stuff like that web server.
Side note. I'm peripherally gaining information on doing 8- or 16-bit software for microcontrollers. Hell, I even have some docs (and maybe HDL source) on a 1-bitter although with 8-bit ALU. Do you have any resources I can give to new people showing the tricks people use to reliably do complex or fast stuff with them? SymbOS would probably be high-end of that but I intend to use them in peripheral controllers, monitoring, and such.
Absolutely, that's why we managed to get something that worked, despite being at least six times as few people who also had to work on other projects :-). Redox is an order of magnitude more complex, because it's trying to solve problems that are an order of magnitude more complex. I just want to dispel some of the "black magic" air around these things.
> Side note. I'm peripherally gaining information on doing 8- or 16-bit software for microcontrollers. Hell, I even have some docs (and maybe HDL source) on a 1-bitter although with 8-bit ALU. Do you have any resources I can give to new people showing the tricks people use to reliably do complex or fast stuff with them? SymbOS would probably be high-end of that but I intend to use them in peripheral controllers, monitoring, and such.
You mean re. digital design, or microcontroller software?
If it's the latter, I'm not sure what to recommend... most of what I know is stuff that every programmer knows + experience you get out of working on embedded devices (and, of course, failing a lot). Fundamentally, programming these devices is no different than programming other computing devices (save, perhaps, for the Harvard architecture, but that's fairly inconsequential), you just work based on other assumptions regarding what's acceptable in terms of failure, performance degradation and so on.
I guess related resources that might help a lot would be:
* Jean Labrosse has a series on his uC/OS kernels. I think the latest is uC/OS-III. I haven't read it, but Labrosse's work in general is extraordinary, so this book may be what you're looking for.
* If you don't mind Ada, Building Parallel, Embedded, and Real-Time Applications with Ada is a pretty good book on the subject, too. I've skimmed it and at least the aspects related to reliability are well-treated.
* If you haven't already read it, van der Linden's "Expert C Programming" is still a very good read. I don't doubt that languages such as Rust and Go are the future for programming higher-end systems, but I don't see C going away on the lower-end in the next 10-15 years, and frankly, I don't think Rust solves too many of the problems I encounter on these systems.
* For dealing with resource-constrained systems in general, Bitsavers.org has exceptional stuff, but it requires some digging and extrapolation. Old computers are a semi-serious hobby of mine -- it's fun, but I also learned a lot of interesting things from studying old systems.
Good recommendations may come from Ganssle's (www.ganssle.com) Embedded Muse, too.
Oh, speaking of which...
> I have text files and PDF's I send to people that ask.
...can I ask :-)?
I also share nickpsecurity ideas on memory safe systems programming.
The last treasury I got, was getting an edition of "Systems Programming with Modula-3" in very good state.
The first problem is the traditional use of uppercase for the keywords, that many dislike, but any suitable IDE would "convert as you type" kind of thing.
The chapters about multi-threading, file IO and graphical system are interesting and it contains multiple references to contemporary work like Topaz.
But while it is very interesting from historical perspective, I am not so sure if it would be useful for modern audiences.
I think we should start a club or something.
Yeah, embedded was what I was talking about and trial&error was what I was worried about. What I've seen people write up is a lot more difficult than software in general. I remember one had severe application issues when it ran with instant on that were eventually resolved as the PLL's not synced up. They put a delay in to let them warm up. Program worked fine. Seen some stuff in immunity-aware programming talking of similar issues. A comprehensive, free collection of stuff like that with tips would be a boon to hobbyists and pro's alike.
re book recommendations
Thank you. I'm already on Ganssle's newsletter. Good stuff in it. I've been published in one or two. Ancient hardware factored into my recommendations for two:
http://www.ganssle.com/tem/tem277.html
http://www.ganssle.com/tem/tem262.html
Ganssle confirmed that one SOC was taking the approach where it had a badass ARM Cortex plus a M0 for interrupts and such. It's claims were like a watered down version of Channel I/O. Cheap, too. :)
Btw, here's your 1-bitter. The link saying datasheet and VHDL design has the PDF's and code. The manual might have tricks worth remembering on other architectures.
http://www.linurs.org/mc14500.html
re Bitsavers
I already have over 10,000 papers from CompSci on my computer, a subset I read instead of skimmed. I keep planning to go through BitSavers to find more interesting stuff but afraid of accumulating useless junk. I just keep putting it off. I was able to pull interesting data on some interesting systems. They even had detailed books on OpenVMS drivers and internals. Something I'd have had to pay for when trying to clone its legendary reliability. Similar stuff on Tandem's NonStop architecture, patents on which should be expired by now or close. I'm keeping an eye there.
My own work was research and occasional work in high-assurance systems. I focused on security as it was hardest and most important with safety or predictability next. I have all the key papers from the past with many of today's best. The lessons learned papers had the most wisdom showing me every step of the way how they tried, failed, and/or succeeded in specific ways. I applied those lessons to my own work. After a brain injury, I'm like INFOSEC's Jason Bourne where I don't remember shit but what's left kicks in to help on forums and such. My main role is evangelist of high-assurance and old wisdom to embed it into more projects. Successes are few but keep it worthwhile. Best, recent example is Tinfoil Chat: Ottela applied every bit of feedback we gave him on Schneier's blog to make one badass design that deserves a rewrite in systems language.
I'll send you my stuff later today. Some people find it useful. Especially since the advice pre-empted the Snowden leaks defeating around 90% of their attacks. Funny that security "pro's" still argue with the shit while pushing what got defeated. In high assurance, it's mandatory to learn from the past. In retrospect, I wish I did lots of embedded or digital design before my memory loss as hardware/software interface is where best results are at. The analog and RF levels, too. I've mostly completed a secure ASIC design methodology and RAD strategy, though, so that will come as quickly as CompSci decides on right HW architecture. :)
It's no rush :-). I'm always looking forward to this kind of stuff. Nothing keeps one's mind fresh the way someone else's well-informed opinions do.
A long time ago, I thought about doing something similar to Dijkstra's letters -- bringing a few colleagues together and beginning to circulate small notes whenever we had something interesting and cohesive enough that it might be wort putting into writing. I don't remember what stopped me, but I still think this is a great way to keep innovation alive. Perhaps it's an idea that I ought to revisit :-).
> Yeah, embedded was what I was talking about and trial&error was what I was worried about.
I'm worried about this, too, and it occasionally drives me insane to see how many people have a "well, let's just get something working and see what happens" approach. It's not just the adversity towards doing some nothing more than simple math first that worries me, it's the fact that I see a lot of people doing this with no regard to how they're going to "see what happens". No serious test methodology, no attempt to at least document assumptions first. It's a wonder we're not at the point where a computer kills someone every day yet.
Trial and error is a natural way to learn things, but it should generally be done just once, ideally by as few persons as possible. We're... not only are we not there yet, we're doing the precise opposite of it.
I've thought about writing down some of these things, but I realized a lot of the "trial and error" spirit by which I learned them still lurks in my understanding of them. And progressing past that is, I've learned, anything but trivial.
> MC14500
Ah, I remember reading about that! I don't remember in what context but I'm sure I've seen the page you mentioned before. It may have been in the context of an article describing OISCs and other minimalistic CPU architecture.
I think these concepts would be great to revisit in the context of the latest developments in microelectronics. One could have thousands of MC14500s on a single chip with today's technology; granted, they could not all talk to each other at the same time due to the limits of interconnects, and not all could be independently interfaced with the outside world, but a hierarchical architecture built out of reliable nodes (and with plenty of room for redundance) might at least be worth investigating.
Even if we go past the realm of on-chip, things have changed dramatically lately. A workstation built out of a hundred Raspberry Pi Zero-grade devices is pretty much on the conceivable side, if not necessarily on the "good idea" side. (The RPi is anything but my favourite system but it's a good example in terms of price, capabilities etc.)
Interesting idea. They sort of do that already with ACM letters, online articles, and so on. If anything, we have so much of this going on that each community silos. That's partly why high assurance INFOSEC and the ITSEC field are two different things. ;) I do occasionally bring up the idea of creating a site hosting top papers and developments in software/security engineering that only invites people that do the research or job. Just people that know shit with a track record and something to bring. They can read the papers, discuss what's in them in moderated forum, and so on. Quite a few like the idea but it will be hard to bootstrap.
"o serious test methodology, no attempt to at least document assumptions first. It's a wonder we're not at the point where a computer kills someone every day yet."
People scratch an itch by modifying an embedded RTOS's kernel. Yeah, I'm amazed we're all still here too.
"Trial and error is a natural way to learn things, but it should generally be done just once, ideally by as few persons as possible. We're... not only are we not there yet, we're doing the precise opposite of it."
Excatly. One example was Burrough's including stack protection into their CPU's. That should've showed up in Intel's stuff as soon as stack attacks became more prevalent. Instead, you see all this trial and error research into tactics that all got beaten. Anything but actually managing one's stack or modifying a CPU to do so. Intel eventually puts it in their off-brand, but good, Itanium CPU. I know Secure64's OS uses it but idk about rest. People still countering stack attacks with tactics to this day despite it solved in 1961.
Ask me about setuid if you want another one even more clever.
"One could have thousands of MC14500s on a single chip with today's technology; granted, they could not all talk to each other at the same time due to the limits of interconnects, and not all could be independently interfaced with the outside world, but a hierarchical architecture built out of reliable nodes (and with plenty of room for redundance) might at least be worth investigating.""
Funny you say that because that was one of my first thoughts. Look up Kilocore for one of those. Other thought was reimplementing ABC Thinking Machines 65,000 CPU design on one or a few chips. The chip's were barely functional: just ALU's or whatever. Still did amazing things in genetic algorithms. One 256 8-bit core design on 500nm accelerated neural networks well in past, too. So, 8-16-bit MPP on a chip are conceivably useful to this day.
http://www.cfbsoftware.com/modula2/Lilith.pdf
Lilith was a custom computer with all kinds of hardware. The mouse eventually inspired Logitech. The software included a safe language (Modula-2), its compiler, an OS (Medos-2), a relational DB, and some other stuff. The trick was safe language, modularity, and simplicity in implementation details. Two or 3 people over a few years.
Later, that turned into the Oberon series of languages, compilers, and OS's. The extensions or ports usually took 1 or 2 students months to a year or two to pull off. Last one they released IIRC was A2 Bluebottle OS. Primitive, but usable. I keep telling people to rewrite it in other safe, C alternatives to save them work. Students got it done fast so I'm sure hobbyists could too.
This is obviously much more full featured (includes a GUI), so this is actually a very nicely sized team for the task.
There's a wiki to get you started if you're interested: http://wiki.osdev.org/Main_Page
I have been thinking hard about a simple plain text file format and associated tools for the last couple months. Sounds nowhere near as impressive. Yet I hope the end result could eventually make a practical difference.