/* You Are Not Expected to Understand This */ (2018)
community.cadence.com
community.cadence.com
Anyway I came upon this comment in the swap in code a bunch of times, never understand it, until I came to it from the right place and it was obvious - a real aha! moment.
So here's the correct explanation - V6 was a swapping OS, only needed a rudimentary MMU, no paging. When you did a fork the current process was duplicated to somewhere else in memory .... if you didn't have enough spare memory the system wrote a copy of the current process to swap and created a new process entry pointing at the copy on disk as if it was swapped out and set that SSWAP flag. In the general swap-in code a new process would have that flag set, it would fudge the stack with that aretu, clear the flag and the "return(1)" would return into a different place in the code from where the swapin was called - that '1' that has "has many subtle implications" is essentially the return from newproc() that says that this is the new process returning from a fork. Interestingly no place that calls the swap in routine that's returning (1) expects a return value (C had rather more lax syntax back then, and there was no void yet), it's returned to some other place that had called some other routine in the fork path (probably newproc() from memory).
A lot of what was going on is tied up in retu()/aretu() syntax, as mentioned in the attached article, it was rather baroque and depended heavily on hidden details of what the compiler did (did I mention I was porting the compiler at the same time ....) - save()/restore() (used in V7) hadn't been invented yet and that's what was used there.
You can click on the message timestamp, which will take you to a page just with that message, which has more options. One of them is "favorite" which will add it to your personal favorites list.
https://www.google.com/search?ei=eMSRX56COtaX-gTVuqPoDQ&q=ha...
One of these years we'll start doing that again. Or maybe we'll just publish the feed as an HN page.
The easier part I was trying to explain was the meaning of the SSWAP and the return(1) which effectively bring you back somewhere different returning a 1 where another routine returns a 0 (a bit like setjmp/longjmp) - this routine when called to do a context switch usually doesn't return anything
> Reading real programs is a part of learning computer science that should be emphasized more. As an undergraduate, we mostly programmed in BCPL, a local Cambridge language that is the fore-runner of C and C++. We had access to the full source code of the compiler and reading it was as valuable as the more theoretical aspects we learned in our compiler-writing course. Today, with open source, it is easy to read the source code for many systems actually in use (Linux, the Apache web server, Hadoop, TensorFlow, and thousands more) but these are hundreds of thousands, if not millions, of lines of code. As I said above, Linux is somewhere between 15M and 20M lines of code, depending on just what you include.
This is an interesting problem. Modern software is so huge and complicated that it's not feasible to go read it just as a learning exercise. Linux devs can probably spend years contributing to the project without having even visited all of its dark corners, much less understanding them.
Not sure what can be done. Though I am tempted to go see what a 9,000 line OS looks like.
Nothing. Any sufficiently complex system or body of knowledge is too large for any one person to understand. Does any physicist know ALL physics that was ever done? Does any one engineer know how to build a Ferrari? Let's be honest, we take for granted a lot of stuff that we don't fully understand, Linux is a major piece of infrastructure, it's natural nobody can fully wrap their head around it.
> Reading real programs is a part of learning computer science that should be emphasized more.
I had never thought of this in an educational sense, but I definitely agree. I sure wish I was taught more of that, toy problems get boring quickly.
Joe and Jane don't need to understand what you are doing, and you don't need to understand what they are doing. Of course partial understanding, especially for related technology is helpful, but the farther away you get the less you know, and the more you need to rely on others.
It might even span across to other niches. "Baking your own bread" is not that hard, if you get to it. But, becausethe entire process (fertilising ground, selecting seeds, growing grain, milling it) is too much for one person, lot's of people just give up alltogether and conclude that "today it is impossible to make your own food". Especially once you learn that bread mostly comes from large, complex industrial complexes today. Yet "just doing it" is very rewarding, as all the home-bakers will tell you.
And it spans across time as well: Sure, no-one can make a toaster from 0: mine own ore, distill your own plastics etc (there's a brilliant documentary from someone attempting just this, forgot its name). So people conclude that when the toaster stops working, there's nothing one can do, and they turn to Ali-express or Amazon and just buy a new one.
In short: I think labor alienation is toxic, in that it seeds itself into areas that don't need hyper specialisation.
The Linux Core Kernel Commentary [1]
It takes only the parts of the kernel you need to understand, leaving out the drivers, device classes, and much of te ancillary pieces. For the parts it chooses, it goes through them line by line and follows the boot process, explains the concepts of a Unix and SYSV-like kernel.
The 15M lines becomes much less intractible.
[1] https://books.google.com/books/about/Linux_Core_Kernel_Comme...
I suppose you could say many things have failed due to no one seeing the big picture among all the tiny details but without people focusing on the details none of that would have been possible to begin with.
I'm not a historian but I can imagine that a medieval specialist, such as for example blacksmith, could very much fix his own house (or maybe even build a new one), grow food, take care of animals, make wooden tools, make some simple herbal medicine etc. I'm guessing a lot of he was consuming was produced by himself and his family. Today, for engineers it's the total opposite - we only consume things produced by others, and essentially by strangers whom we've never met.
Regarding fragility specifically, if a major crisis comes (say a war), the desk jockeys, esp. ones didn't have the wisdom to accumulate resources, will fare much worse than someone who's say running a homestead. If they did accumulate resources (provided they don't become worthless by hiperinflation etc.), they will be able to buy the essentials and hope their stash lasts until the end of the crisis.
You should read Collapse Of Complex Societies by Professor Joseph Tainter.
Tl;dr a step change in complexity of ancient societies more often than not lead to their downfall. They were unable to train/retrain specialists quickly enough to deal with surprises.
>> The social impact, for one, is in general a large extent of estrangement from our human essence. Tangibility is intuitive, abstraction, not so much.
Specialization does introduce fragility, but that's something you brought to the discussion, not something that was present in the comment. This seems more like the common complaint that working on abstract fuzzy intangible things is damaging to your soul.
Using it as a counter argument against specialization isn’t very insightful though. Not even Marx tried to argue that (to Marx specialization was only alienating when you did it for the benefit of the bourgeoisie). Specialization is one of the most fundamental components of a developed economy. If you look at any ancient agrarian society, they had barely any specialization compared to a modern developed economy, and if you track the economic development of any society, you’ll notice specialization ticks upward along with it.
On the other hand, you make a fair point - tangible contribution to Things is probably built into us from our origin as tool-makers.
But isn't the gradual separation from direct contact with Made Things inevitable? At some point the body of knowledge required to make a meaningful contribution grows so large that learning it approaches the human lifespan. How can we avoid too much hyper-specialization in that case?
As a kid, it was an easier thing to reason about (wrongly). As an adult, I’ve concluded that it was never possible to both have something we’d think was civilization and for one person to know substantially everything humans knew.
The basic idea that Alan Kay has presented in his many similar YouTube presentations (most of which run with that office application) is what we call today domain-specific languages, you write the rules for antialiasing a pixel in whatever the ideal language would be for those rules: you just imagine that you have a dream language and "wishful think" the solution: then you go and you reimplement that DSL in another language that you are wishfully conjuring, a language-language, and that final one can become self-describing and can compile and optimize itself.
This is also a method of programming advocated in MIT's old SICP lectures, which will also cover the implementation of Lisp in itself, which is sort of a prerequisite for it.
That all said, I would love to read the source code someday but I can't seem to find it anywhere, the software product may be the KSWorld that was described in [1].
More careful reading of the proposal and reports will easily reveal this.
The 20,000 lines of code was a strawman, but we kept to it, and got quite a bit done (and what we did do is summarized in the reports, and especially the final report).
Basically, what didn't get finished was the "bottom-bottom" by the time the funding ran out.
I made a critical error in the first several years of the project because I was so impressed with the runable-graphical-math -- called Nile -- by Dan Amelang.
It was so good and ran so close to actual real-time speeds, that I wanted to give live demos of this system on a laptop (you can find these in quite a few talks I did on YouTube).
However, this committed Don Knuth's "prime sin" : "the root of all evil is premature optimization".
You can see from the original proposal that the project was about "runnable semantics of personal computing down to the metal in 20,000 lines of code" -- and since it was about semantics, the report mentioned that it might have to be run on a supercomputer to achieve the real-time needs.
This is akin to trying to find "the mathematical content" of a complex system with lots of requirements".
This is quite reasonable if you look at Software Engineering from the perspective of regular engineering, with a CAD<->SIM->FAB central process. We were proposing to do the CAD<->SIM part.
My error was to want more too early, and try to mix in some of the pragmatics while still doing the designs, i.e. to start looking at the FAB part, which should be down the line.
Another really fun sub-part of this project was Ian Piumarta's from scratch runnable TCP/IP in less than 200 lines of code (which included parsing the original RFC about TCP/IP.
It would be worth taking another shot at this goal, and to stick with the runnable semantics this time around.
And to note that if it wound up taking 100K lines of code instead of 20K, it would still be helpful. One of the reasons we picked 20K as a target was that -- at 50 lines to a page -- 20K lines is a 400 page book. This could likely be made readable ... whereas 5 such books would not be nearly as helpful for understanding.
Well anyway, thanks a lot for these presentations. I don't think it's as good as a software course where someone breaks all my preconceptions of what computing is, to leave me truly free, but they are helping to widen my thoughts.
Did some digging and found this site with some code, but I haven't looked closely enough to determine what it is specifically for... http://tinlizzie.org/dbjr/ And it is not exhaustive of what the PDF describes. Maybe someone is able to contact one of the original group working on this?
Here you go: http://v6.cuzuco.com/
I believe most of the Linux kernel source is actually drivers, or otherwise code for which you don't have the hardware to use on, so the actual relevant amount of code in the Linux kernel to e.g. a typical x86 PC might be an order of magnitude less.
I think you are going to like STEPS:
>The overall goal of STEPS is to make a working model of as much personal computing phenomena and user experience as possible in a very small number of lines of code (and using only our code). Our total lines of code target for the entire system -- from user down to the metal is 20,000, which we think will be a very useful model and substantiate one part of our thesis: that systems which use millions to hundreds of millions of lines of code to do comparable things are much larger than they need to be.
I can't agree. Linux is certainly huge, and yes it's size means it's beyond any newbie's capacity to understand it all.
But nonetheless lots of newbies contribute to it regularly. Most of them build up the understanding they need by reading the kernel code. Obviously, that means you don't have to understand all of it to get a lot of benefit from just reading a small subsection.
So it is indeed possible to go read the Linux kernel just as a learning exercise. The key point is you don't have to read all of it to get a benefit out of it. Reading just a small subset works just as well as reading the entire source of a small project.
https://jacquesmattheij.com/task.cc
3500 lines, micro kernels can be much smaller than that even.
But how much of it is the actual OS versus device drivers and other "extensions" of the main kernel? I assume there is somewhere a core that everything plugs into that has more or less the same functionality as what was done in the 9000 lines of the original Unix v6 code and only uses a fraction of all of the code of the kernel.
https://unix.stackexchange.com/questions/223746/why-is-the-l...
So it's still substantially heavier than would be easy to study, but you could probably still fit the core into your head (I've been able to deeply understand systems with similar LOC counts, though there's probably substantial differences in 'density' between codebases)
A debugger can also help. You can step through code as it runs. You can also often identify which code is associated with a feature, by putting breakpoints in likely places (dispatch functions!) and triggering the feature.
Wow, Unix is a work of art. I have hobby projects at 2x SLoC which are literally just small web services :(
File system: about 15Kloc
Network driver + IP stack: about 10 Kloc
Task manager, posix compatible libc odds and ends, maybe another 50Kloc, Graphics driver: 10Kloc, window manager 2500 lines.
So not that bad, you could do it by yourself in about 2 years of hard work, probably less than that if you use a VM instead of actual hardware if you're a halfway competent programmer. I've done it.
Perhaps we need something more modular, like a microkernel architecture :)
There are exactly zero reasons why this is the case other than a stupid ego war between Andrew Tanenbaum and Linus Torvalds who was gung-ho on recreating an OS from the 70's.
This irritates me as well, but blaming it on the kernel strikes me as absurd. Blame it on all the layers of abstraction and general lack of caring in applications instead.
It's true that there are cases where the Linux scheduler gets in the way of a good desktop experience, but desktop latency is too high even when the scheduler isn't a bottleneck. And if anything, good scheduling gets harder in a microkernel, not easier, because you have less insight into what is going on at a system level. (Ultimately, the distinction simply shouldn't matter though -- not for desktop experiences, that is.)
Yes, there are plenty of problems in user space. But those are not the root cause and without addressing the root cause you will not be able to solve the problem even if you don't have bloated software.
Lisps are hardly a modern invention. "Modern" languages do seem to be ever so gradually working their way towards full compile time metaprogramming though, gaining a great deal of expressiveness in the process.
Why can't distros be a package of microkernel + drivers?
The first week or so of the process was learning how to go from a new Debian machine (with the expectation we'd only used Windows or Macs before) through "git clone git://git.kernel.org/pub/scm/linux/kernel/git/stable/linux-stable.git" (not on Github at the time) to finding the function called by "man 2 fork" and navigating the source tree/unmasking preprocessor stuff to show the actual implementation. Yes, you can easily make a wrong turn if you're trying to do that on your own. But the actual linux/kernel directory, which has most of the parts you'd want, isn't that much larger, and a lot of the difference is modern requirements like power saving and security.
Performance. Basically, the required amount of context switches (since all drivers run in user mode) impacts performance and cache use in a negative way compared to monolithic kernels. Whether this is still a significant issue with modern designs, I don't know, but that was the argument back in day (i.e. the early 90s when Linux came about).
They are built (usually) at the same time as the kernel, yes, due to the lack of ABI stability guarantees, but most drivers you don't use won't take up any RAM.
related from 2008 https://news.ycombinator.com/item?id=203437
dmr's comments are at https://web.archive.org/web/19990501224738/http://cm.bell-la..., followed by Comments I do feel guilty about.
Fun fact: At one point in history, Lions' book had the status of being a forbidden book of knowledge due to copyright problems. After AT&T commercialized UNIX v7 in 1979, the permission for using the source code in classrooms was withdrawn, it could no longer be read without a UNIX license. But students continued to make illegal copies and study it illegally, who would later become the next generation of hackers and developers. This is also the initial motivation for Andrew Tanenbaum to write MINIX - as a classroom substitute. And he had to be as careful as possible not to derive anything from the original Unix code, such as Lions' book.
https://upload.wikimedia.org/wikipedia/commons/5/50/Lions_Co...
This picture from an Unix history exhibition recreated the historical atmosphere by showing Lion's book in front of a huge legal warning [0].
I found this period of history was amusingly similar to the worldbuilding of some fantasy fictions - In the world of magic, the original magic scriptures are forbidden to be learned, the books are sealed, reading them is a taboo [i.e. UNIX source code]. Instead, one must read "safe" textbooks to learn magic. These textbooks are written by great wizards, and it basically contains a second-hand and distorted rehearsal of the knowledge in the original scriptures, sanitized to the level that deemed safe for human, but it is not as powerful [i.e. Tanenbaum's MINIX]. And there are outlaws who want power and they ignore the taboo, however, attempting to read the original magic scriptures will cause irreversible psychological injuries [i.e. cleanroom violation, AT&T lawsuits are coming].
[0] This document contains proprietary information of the Bell System. Use is restricted to authorized employees of the Bell System who have a job-related need-to-know. Duplication or disclosure to unauthorized persons is prohibited. Outside the Bell System, distribution of this document has been restricted to holders of a license for the UNIX Time Sharing Operating System, Sixth Edition from the Western Electric Company. Use, duplication, or disclosure of this document is subject to the restrictions stated in such license from Western Electric Company.
https://en.wikipedia.org/wiki/Linux
And then Bill and Lynne Jolitz released 386BSD:
https://en.wikipedia.org/wiki/386BSD
And from that point forward us 'hobbyists' had access to more serious operating systems than bootlegged Xenix or QnX copies.
Minix always felt like a toy, nice but not quite unix.
> Thanks for putting a version of MINIX inside the ME-11 management engine chip used on almost all recent desktop and laptop computers in the world. I guess that makes MINIX the most widely used computer operating system in the world, even more than Windows, Linux, or MacOS... -- Andrew S. Tanenbaum
They are based on the idea of a computer programmer being summoned/teleported into a world of magic and then using programming principles to become a powerful wizard.
I did at one point understand that. The toughest thing for me to understand was kt assembly code.
// PLEASE DO NOT ATTEMPT TO SIMPLIFY THIS CODE.
https://github.com/kubernetes/kubernetes/blob/master/pkg/con...
* You are not expected to understand this.
*/
if(rp->p_flag&SSWAP) {
rp->p_flag =& ~SSWAP;
aretu(u.u_ssav);
}
I don't understand what `=&` is. Is that mistyping `&=` in the book? This should fail to compile, right?http://www.tom-yam.or.jp/2238/src/slp.c.html#line2238
It's what we'd use `&=` for nowadays. Note that the "B" language had these sorts of operators around the opposite way for how they appeared in "C" (e.g. read https://www.bell-labs.com/usr/dmr/www/kbman.html).
EDIT: Turns out that pre-standardized C is documented:
http://www.math.utah.edu/computing/compilers/c/Ritchie-CRefe...
Page 8 documents `=&`.
Edit for people who haven’t used this stuff: SKILL is the scripting language baked into Cadence tools. It’s lispy, but also allows you to move function names out of the parens. So you can (frob) or frob()
If you can explain clearly and precisely but it's a long explanation, then do so. The longer the explanation, the more it's needed. If it's that complicated to explain the first time round, it's gonna suck when you come back to that code five years later.
The OP says:
> Interestingly, years later, Ken Thompson admitted that the reason that it was so hard to wrap your head around what is going on is that it is wrong.
The original writers actually got it wrong, it was buggy code -- or kind of accidentally not buggy due to a quirk of the particular architecture it ran on, if I get the gist of what they are explaining, I definitely don't get the details!
# If you are reading this, I am so sorryI get the sense not enough of this happens. i didn't do too too much of it in my formal computer science education. If you get to work on open source as a significant part of your time, you might at least spend a lot of time looking at someone's code, but it isn't always the best code. If you want to be a good writer, you read the greats (in various ways) not just the backs of cereal boxes.
In [1] Brian Kernighan interviews Ken Thompson, who's wearing a t-shirt which contains the relevant snippet along with a comment "ΕΠΙΤΕΛΟΥΣ ΤΟ ΚΑΤΑΛΑΒΑ!" (greek for 'I finally understood it!').
Much--but not all--of the code available for people to learn from fails this test today.
what's the Shakespeare? Or The Adventures of Huckleberry Finn?
i think rob and/or ken also discuss this in one of those history of unix presentations they did in the past couple of years
A lot of Engineering is built around models of the world that describe physical behavior. Usually the equations we use are simplified and contain a lot of subtleties when you dig into them.
In fluid dynamics for example the Navier–Stokes equations are infamous because we don't understand all the theory behind why they work. They are used all the time despite not being fully understood.
Semiconductor behavior is another area, the Band Gap models used to describe them incorporates a lot of quantum physics nuances involving Fermi–Dirac Statistics and wave functions and the like.
At some point when you do work in Engineering you are "high enough" up in the stack you can use simplified abstractions and trust that what is underneath is acceptable.
From what I've seen software engineer has these abstractions as well, process switching sounds like one of these cases.
> I often feel ashamed of software engineering
Agree. There are many examples of how software development is unable to constantly deliver predictable results, and how software development has yet to become "engineering".
But this is not one of them.
Much engineering work also has better defined interfaces than in software development, at least where working with others is concerned. If I'm designing a PCB to fit into a die-cast enclosure, I don't need to know why the mechanical designer chose the draft angle they did to enable tool release. My consideration of mechanical issues can focus on a) satisfying the mechanical interface to the enclosure, b) meeting the static and dynamic physical constraints and c) enabling efficient and reliable assembly.
Is that what you mean?
You're right. Such a thing would never happen. Never!