Dennis Ritchie’s first C compiler (c. 1972)
github.com
github.com
Having talked my way into using the local university EE dept's 11/45, I found the source tree was mounted (they had a 90MB CDC "washing machine" drive) and decided to print it out on the 132-col Tally dot-matrix printer.
Some minutes into this, the sysadmin bursts into the terminal room, angrily asking "who's running this big print job".
I timidly raise my hand. He asks what I'm printing. I tell him the C compiler source code because I want to figure out how it works. He responds "Oh, that's ok then, no problem, let me know if you need more paper loaded or a new ribbon".
if (peekc) {
c = peekc;
peekc = 0;
} else
if (eof)
return(0); else
c = getchar();This triggers me because many people jump on “bad c0de” in the forums, but then you read some practical code on github, and it is (a) not as perfectly beautiful as they imagine at all and (b) is still readable without nightmares they promised and (c) the algorithms and structure itself requires programmer’s perception and understanding levels far beyond the “it’s monday morning so i feel like missing an end of statement in a hand-written scanner” anyway.
dmr was one of, if not the first C programmer(s).
The C language (and all of Unix) was designed to be very terse as a consequence.
And there are people claiming that computer scientists are not conservative :)
To what extend this explanation is correct is another question... The article by Denis Richie says "Other fiddles in the transition from BCPL to B were introduced as a matter of taste, and some remain controversial, for example the decision to use the single character = for assignment instead of :=".
It's a kind of butterfly effect :) Mr. Richie prefered "=" over ":=" and fifty years later a server crashes somewhere because somebody wrote a=b=c instead of a=b==c.
Similar story with the design of QDOS / MS-DOS / Windows and nowadays with Android. Both designed for super underpowered machines that basically went away less than a decade after they were launched and that will be hobbled because of those early decisions for a long, long time.
If they had gone for wart-free on properly powered hardware, they would be stuck back in Multics land or living with the gnu Hurd—cancelled for being over budget or moving so slowly that projects that actually accomplish what the users need overtake them.
Do I wish that C had fixed its operator precedence problem? Sure. But the trade offs as a total package make a lot of sense.
Why is this precedence weird? Bitwise AND tends to be used to transform data while a logical AND tends to be used for control flow.
As in:
if (x & 2 == 2)
...is actually parsed as: if (x & (2 == 2))
...which isn’t intuitive.“In retrospect it would have been better to go ahead and change the precedence of & to higher than ==, but it seemed safer just to split & and && without moving & past an existing operator. (After all, we had several hundred kilobytes of source code, and maybe 3 installations....)“
https://multicians.org/history.html
Instead we pile mitigations on top of mitigations, with hardware memory tagging being the last hope to fix it.
Perhaps this is why a programmer would want to rewrite a system & tout "funny success stories" about the effort & results?
https://news.ycombinator.com/item?id=25844428
> Why couldn't you just upgrade the dependencies once then set up the same CI/CD you're presumably using for Svelte so that you can them upgrade versions easily?
Because the existing system was painful & time/energy intensive to upgrade. It happens with tight coupling, dependency hell, unchecked incidental complexity, architecture churn, leaky abstractions, etc...
Maintenance hell & > 1 month major version upgrades tend to occur with large, encompassing, first-mover frameworks, often built on a brittle set of abstractions, as they "mature" & "adapt" to the competition. e.g. Rails, Angular...
Speaking of terseness, I love the the fact that C does not have 'fn'.
Obviously, there's no way to go back and edit anything mid-stream, you have to write the whole thing out in one shot.
$ cat foo bar > baz
will join the files foo and bar together into a single file called bazE.g.:
if (peekc) {
c = peekc;
peekc = 0;
} else
eof ?
return(0) :
c = getchar();
The first else clause sill looks weird, but the final part isn't nearly as out of place (well i guess assigning in a ternary would be weird, but in terms of indentation) and its not like we actually changed anything. a = return b; return(0); else
makes a bit of sense.“auto” exists in the initial version.
Edit : I know it’s also storage specifier but it also does type deduction here, hope I’m not confused with the terminology
In original C you could declare variables without a type, and these variables with auto have no type.
So this "auto" is essentially saying "this variable without a type uses automatic storage". That's a completely different language feature from "deduce the type of this variable".
Interestingly, long was commented
If you don't believe me, ask a non programmer friend what kind of thing an "integer" is in a computer program. Then ask them to guess what kind of thing a "long" is.
The only saving grace of these terms is they’re relatively easy to memorise. int_16/int_32/int_64 and float_32/float_64 (or i32/i64/f32/f64/...) are much better names, and I'm relieved that’s the direction most modern languages are taking.
(Edit: Oops I thought Microsoft came up with the names WORD / DWORD. Thanks for the correction!)
"Word" as a term has been in the wide use since at least 50s-60s, you can't really blame MS for that
You can _absolutely_ blame Microsoft for using it to mean something that it isn't.
(I know it's floating point, but it's the same as long/double).
What made you say that?
It is probably rather by chance that we ended up with the term "string (of characters)", as opposed to for example "sequence of characters". In a different universe we might be talking about charseqs (rhymes with parsecs) instead of strings.
Also, don’t forget about short, it feels so sad and alone!
"long" was commented out.
Why should a non programmer understand programming terms? Words have different meanings in different contexts. That's how words work. There is no need to make these terms understandable to anyone. The layman does not need to understand the meaning of long or word in C source code.
Ask a non-golf player what is an eagle or ask a physicist, a mathematician and a politic the meaning of power.
Word and long may have been poor word choices, but asking a non-programmer is not a good way to test it.
The word 'long' is ambiguous and unmemorable. And the type "word" is actively misleading. If you called a variable or class 'long', it wouldn't pass code review. And for good reason.
'Power' is an excellent example of what good technical terms look like. "Power" has a specific technical meaning in each of those fields, but in each case the technical meaning is suggested by the everyday meaning of the word "power". Ask a non-physicist to guess what "power" means in a physics context and they'll probably get pretty close. Ask a non-programmer to guess what "integer" means in programming and they'll get close. Similarly computing words like "panic", "signal", "file", "output", "stream", "motherboard" etc are great terms because they all suggest their technical meaning. You don't have to understand anything about operating systems to have an intuition about what "the computer panicked" means.
Some technical terms you just have to learn - like "GPU" or "CRDT". But at least those terms aren't misleading. I have no problem with the term "double precision floatingpoint" because at least its specific.
"long", "double", "short" and "word" are bad terms because they sound like you should be able to guess what they mean, but that is a game you will lose. And to make matters worse, 'long', 'double' and 'short' are all adjectives in the english language, but used in C as nouns[1]. They're the worst.
[1] All well named types are named after a noun. This is true in every language.
auto long sum;
Declares a variable named "sum" of type "long int" and of automatic storage duration.The "double" comes from double-precision floating point. In the 1970s and 1980s anyone who came near a computer would know what that meant. Anyone who ever had to use a log table or slide rule (which was anyone doing math professionally) would know exactly what that meant.
There are good sound reasons for the well-chosen keywords. Just because one is ignorant of those reasons does not mean they were not good choices.
"long", "short", "signed", "unsigned", "int", "float", and "double" are all type specifiers. The language specification includes an exhaustive list of the valid combinations of type specifiers, including which ones refer to the same type. The specifiers can legally appear in an order. For example, "short", "signed short", and "int short signed", among others, all refer to the same type. (They're referred to as "multisets" because "long" and "long long" are distinct.)
By your own standards, "word" and "long (word)" are actually excellent terms, because they're conventions used by the underlying hardware, and using the same conventions absolutely makes sense for a programming language that's close to the hardware:
All those fields could do with less jargon. Especially since in many cases there is a common word.
Lawyers especially give me the impression that they use jargon to obscure their field from regular folks.
Our field, being new, should not make the same mistakes, but yet, here we are, where the default "file viewer" on Unix is cat, the pager is called "more", etc.
I don't see any reason why those types could not have had descriptive names from the start. Well, there was one, lack of experience, so ¯\_(ツ)_/¯
Jargon lets you summarize whole concepts in one word. A lot of jargon is functional in the math/programming sense: you can pass 'arguments' to the jargon. The jargon adds levels of abstraction that let the users communicate faster, with less error, and higher precision.
For instance, the distinction between civil and criminal matters; or the distinction between malfeasance, misdemeanor, and felony.
Really? You think that lawyers (and physicians) use Latin/Greek words in order to confuse regular folks?
These fields are very old, and changing the meaning of something Mens Rea or lateral malleolus is going to require A) an exact drop-in replacement which will probably just as obscure, B) retraining of an entire set of people.
Our field is similar in that we have a mountain of jargon that's largely inaccessible to regular folks. The point of those words are not to converse with regular folks but to convey information to others in the field with as little ambiguity as possible.
Intuitive things come from prior experience. They are a kind of inertia that you just have to work with.
See economics for a great example of jargon-vernacular crossover.
Sure it can, which is why you gotta be double-careful naming things and not try to take metaphors too far. A jargon term needs to crisply identify the crux of the concept and not confuse with irrelevant or misleading details.
Agreed. I've never seen a vernacular term fill this role well.
If you need to learn the technical concepts either way to be effective, might as well give them a name that doesn't conflict with another definition most people know.
In economics, "cost" is a good example. This is a distinct concept from "price". "Comparative advantage" is another term in economics; this is perhaps not used in vernacular conversation, but I can tell you from personal experience that it certainly doesn't convey to most people the definition understood by someone with an education in economics -- the vernacular reading doesn't imply the jargon definition.
It seems to me that the difference is how the jargon is used. I imagine that someone without a CS background would quickly realize, when overhearing a conversation about binary trees, that the subject is something other than a type of flora.
I can tell you with confidence borne from frustrating experience that using economics jargon, such as that I mentioned above, with a lay audience gives the audience no such impression that the terms mean anything other than what they perceive them to mean.
The PDP-9, PDP-10, and PDP-18 have 18 bits registers. The world had not settled on 16/32/64 bits at all.
Even the intel 80286 far/fat pointers are 24 bits.
Admittedly I read that more than 30 years ago :-O
Alan Snyder's 1974 masters thesis [1] describes the Honeywell 6000 GCOS port in some detail. In 1977, there were three different ports of Unix underway – Interdata 7/32 port at Wollongong University in Australia, Interdata 8/32 port at Bell Labs, and IBM 370 mainframe port at Princeton University – and those three had C compilers too.
[0] https://www.bell-labs.com/usr/dmr/www/portpap.pdf
[1] https://apps.dtic.mil/dtic/tr/fulltext/u2/a010218.pdf (his actual thesis was submitted to MIT in 1974; this PDF is a 1975 republication of his thesis as an MIT Project MAC technical report)
The fact that “int” could be 16-bits on a PDP-11, 32 on an IBM 370, 36 on a PDP-10 or Honeywell 6000 - that was a real aid for portability in those days.
But nowadays, that’s really historical baggage that causes more problems than it solves, yet we are stuck with it. I think if one was designing C today, one would probably use something like i8,i16,i32,i64,u8,u16,u32,u64,f32,f64,etc instead.
When I write C code, I use stdint.h a lot. I think that’s the best option.
The real mistake in retrospect is that int and long are platform dependent. This is an amazing time sink when writing portable programs.
For some reason C programmers looked down on the exact width integer types for a long time.
The base types should have been exact width from the start, and the cool sounding names like int and long should have been typedefs.
In practice, I consider this a larger problem than the often cited NULL.
(b) Many important machines had word sizes that were not a multiple of 8.
I think you misunderstood. There's no explanatory comment. The "long" keyword is commented out, meaning that it was planned but not yet implemented.
...
init("int", 0);
init("char", 1);
init("float", 2);
init("double", 3);
/* init("long", 4); */
init("auto", 5);
init("extern", 6);
init("static", 7);
...for is missing too.
- Rewrite B compiler in B (generating threaded code)
- Extend B to a language Ritchie called NB (new B). This compiler generated PDP assembly There is no version of the NB compiler known to exist.
- Continue extending NB until it became the very early versions of C.
You can read the longer version of this history here:
Later, the language would be retargeted to the PDP-11 while on the PDP-7. Various changes, like byte-addressed rather than word-addressed memory, led to it morphing into C after it was moved. There was no clear line between B and C -- the language was self-hosting the whole time as it changed from B into C.
Mr. Ritchie wrote a history from his perspective published in 1993. I've mostly just summarized it above. It's available here: https://web.archive.org/web/20150611114355/https://www.bell-...
It has been done many times. The first assemblers were written directly in machine language. The first compilers were written in assembly. Many implementations of FORTRAN, ALGOL, COBOL, etc. As late as the 1970s it was not unknown to write a new high level language directly in machine code. Steve Woz's BASIC for the Apple II was "hand-assembled", as he put it.
So, taken literally, certainly not. People still do it today as a hobby or educational project.
But if we take the question in a looser sense? Yes kind of, at least in the UNIXish world. No one has implemented a serious high level systems language, except in a high level systems language, for a long time now. Rust, for example, was initially implemented in Ocaml. And normally that's what would be called the first Rust bootstrap. But OCaml was implemented in C, probably compiled by Clang or GCC. GCC was written in C and... so on.
Often one finds it does lead back to DMR at a PDP-7. On the other hand, I strongly suspect something like IBM's COBOL compiler for their mainframes (still supported and maintained today!) would not have such a heritage.
Very cool - so basically (haha) Woz's integer BASIC was bootstrapped by writing it in assembly language and translating it into machine code by hand.
I wonder if someone (Woz?) has written a 6502 assembler in integer BASIC to allow it to bootstrap itself?
[1] https://mitpress.mit.edu/sites/default/files/sicp/full-text/...
On the other hand, I still remember at least a few 6502 hex opcodes, not that I have any use for that information anymore. The instruction set is small enough that it doesn't surprise me that Woz would have it memorized.
So most weren’t exactly bootstrapped the way we are talking here.
LuaJIT is a prominent counterexample. The base interpreter in written in assembly for performance. It comes with its own tradeoffs, especially portability.
DynASM is actually really cool.
I almost forgot: apropos bootstrapping, because LuaJIT requires Lua to build, it includes a single-file, stripped down version of PUC Lua 5.1 which it can build to bootstrap itself if the host lacks a Lua interpreter.
A program using DynASM doesn't depend on Lua or a C compiler to generate assembly at runtime but you need both to build it.
http://pascal.hansotten.com/niklaus-wirth/cdc-6000-pascal-co...
"The method used to create the first Pascal compiler, described by U. Ammann in The Zurich Implementation, was to write the compiler in a subset of unrevised Pascal itself, then hand-translate the code to SCALLOP, a CDC specific language. The compiler was then bootstrapped, or compiled using the Pascal source code and the SCALLOP based compiler, to get a working compiler that was written in Pascal itself. Then, the compiler was extended to accept the full unrevised Pascal language."
Even if you're using a loose definition of "root" as in a parent language is any language involved in writing another language, I doubt there is one root. There are certainly have been languages and computer architectures (with assembly languages) that are independent islands.
https://bootstrappable.org/ https://bootstrapping.miraheze.org/wiki/Main_Page
In particular BOOTSTRA [1] looks really fun. I have also toyed with the idea of using MS-DOS 3.30 as a guaranteed-ubiquitous build environment.
It comes with a filesystem, a text editor (EDLIN.EXE), an object file linker (yup!), a debugger/assembler (DEBUG.EXE is an amazing tool), a programming language with decent string handling (GWBASIC.EXE), and a command interpreter with batch file scripting to glue it all together.
Not sure I would write the assembler as batch files though, that is really hardcore. :)
That entire build environment would fit on a single 360KB floppy.
https://github.com/fosslinux/live-bootstrap/blob/master/part...
Those binaries will boot on anything from an ancient IBM PC with an 8088 CPU all the way through to fairly recent x86 systems (if they still have legacy BIOS boot support).
It wouldn't be defensible to use those descriptors for Windows XP. In a word, it would be wrong. Calling DOS 3.30 "guaranteed-ubiquitous" is even wronger than that.
I'm not even sure it's the first C compiler written in C, though - it just says in the github description "the very first c compiler known to exist in the wild."
Regardless, if it's from 1972 it's a very early version.
This isn't obvious to me.
I just assumed that the first iteration was compiled by hand to bootstrap a minimum workable version. Then the language would be extended slightly, and that version would be compiled with the compiler v. n-1 until a full-feature compiler is made.
It makes sense to write a compiler in a different language, but given the era, I could see hand-compilation still being a thing.
https://github.com/mortdeus/legacy-cc/blob/master/prestruct/...
char waste[however-many-bytes-are-needed];
?
Especially on single-tasking systems, it doesn't matter how much memory you waste, because no other program is affected by it.
Or in other words: If you know you are guaranteed to have x KB of memory available and no other program can steal it, and you know that you don't need more than y KB, then why allocate up to y KB dynamically when you can just straight-up allocate y KB statically, which is less work?
9.34.2 Assembler Directives
The PDP-11 version of as has a few machine dependent
assembler directives.
.bss
Switch to the bss section.
...The feeling dissipated somewhat when our guide explained that it’s a treasure map
hshsiz 100;
hshlen 800; /* 8*hshsiz */
hshtab[800];
These were at file scope. I assume they default to int, but when was the demand for = added?IMO, it is close to Assembly:
* you reserve space and possibly set an original value with "hshsiz DB 100" or "paraml DB ?"
* assignment is a different business which involves a runtime instruction (MOV)
Hence the same difference (no = sign for the first, passive, compile-time operation; an = sign for the second, active, runtime operation) in this proto-C.
(Of course, this doesn't answer your question of "when" did the syntax of those 2 operation fuse :-) )
The full compiler is super compressed. Hundreds of lines per file. It would be great if there was an explanation of the general idea of the design somewhere.
Edit: Elsewhere in-thread, retrac posted a detailed summary of how C developed from B.