The original “Hello World”: B. W. Kernighan's intro to B (1973)
web.archive.org
web.archive.org
hello, world.
My guess is it refers to archy and mehitabel, very popular early 20th century free verse poetry by Don Marquis. His work was still very much in favor with intellectuals around the time of B’s creation. Or it might have been influenced by the poet e.e. cummings.Although the book was written in 1979, BCPL dates to 1967. There is a hello, world example on page 8. Wondering if it is likely that hello, world as a simple example program really began with B?
BCPL -> B -> C (perhaps) but there is a lot of overlap. C came out around 1973.
As you say: Hello World is quite old and probably predates these Johnny-come-lately, fly by nights like C.
http://progopedia.com/language/b/
>B is an interpreted programming language for mini-computers, a direct descendant of BCPL and ancestor of C.
http://www.catb.org/~esr/writings/taoup/html/c_evolution.htm...
>based on Ken Thompson's earlier B interpreter which had in turn been modeled on BCPL,
Iirc, and it’s been a while since I’ve read the history of Unix. B started as an interpreted language but evolved into C as types were added and it was switched to compiled to better interact with UNIX operating system internals (with Ritchie wanting to rewrite large portions of Unix in the new language.)
Here is a snippet from an interview of B's creator, Ken Thompson in 1989, explaining this:
"MSM: Did you develop B?
Thompson: I did B.
MSM: As a subset of BCPL
Thompson: It wasn't a subset. It was almost exactly the same. It was a interpreter instead of a compiler. It had two passes. One went into intermediate language and which one was the interpreter of the intermediate language. Dennis wrote a compiler for B, that worked out of the intermediate language."
~ https://www.princeton.edu/~hos/mike/transcripts/thompson.htm
As always on HN there is a certain amount of ... discussion about some of the finer points, generally about typing. I suspect that if you follow the refs in the WP article most of the usual bikeshedding here will resolve itself satisfactorily.
For me (50 year old bloke) I dimly recall C always being available https://en.wikipedia.org/wiki/C_(programming_language) - apparently 1972ish so I was 2 or 3 when C came out and replaced B which probably didn't do much ....
Oh look at this (wrt B): "However, it continues to see use on GCOS mainframes (as of 2014)"
C was needed because the byte-addressed PDP-11 was a bad fit for the word oriented B language. See https://archive.org/details/bstj57-6-1991 for details.
Although it appears their main business nowadays is a Windows-based application to manage maintenance requests for landlords, etc
> The second development concerned the language in which UNIX was written. By now it was becoming painfully obvious that having to rewrite the entire system for each new machine was no fun at all [0], so Thompson decided to rewrite UNIX in a high-level language of his own design, called B. B was a simplified form of BCPL (which itself was a simplified form of CPL, which, like PL/I, never worked). Due to weaknesses in B, primarily lack of structures, this attempt was not successful. Ritchie then designed a successor to B, (naturally) called C, and wrote an excellent compiler for it. Working together, Thompson and Ritchie rewrote UNIX in C. C was the right language at the right time and has dominated system programming ever since.
Tanenbaum doesn't say it, but it almost seems like B and C were designed for creating UNIX. I wonder to what extent the authors of B and C were designing the languages for creating UNIX.
[0] In one of the previous paragraphs, Tanenbaum mentioned that the first version of UNIX was written in assembly.
[1] Modern Operating Systems (ed. 4, p. 715)
This seems like a biggie.
> putstr(getstr(s)); putchar('*n');
What could possibly go wrong?
main( ) {
extrn a, b, c;
putchar(a); putchar(b); putchar(c); putchar('!*n');
}
a 'hell';
b 'o, w';
c 'orld';
B uses single quotes to denote a character (like C), and double quotes to denote a string. Each character is a word (as are all variables in B), which is 36 bits long, so it can hold 4 ASCII characters! A character literal with fewer than 4 characters is zero padded (as is, presumably, the left over 4-bit nibble). In fact an earlier snippet in that document just outputs "Hi!" because that way you only need one character: main( ) {
auto a;
a= 'hi!';
putchar(a);
putchar('*n' );
}Multi-character constants are possible, and have an implementation-defined value.
int fourcc = 'abcd';
(This may be supported by C++ also, I'm not sure. So that is to say, a multi-character constant in C++ perhaps doesn't have type char, but an implementation-defined type.)GCC on Ubuntu 18.04:
$ cat hello.c
#include <stdio.h>
int main(void)
{
int hello[] = { 'lleH', 'w ,o', 'dlro', '!', 0 };
puts((char *) hello);
return 0;
}
$ gcc hello.c -o hello
hello.c: In function ‘main’:
hello.c:5:19: warning: multi-character character constant [-Wmultichar]
int hello[] = { 'lleH', 'w ,o', 'dlro', '!', 0 };
^~~~~~
[ ... and similar errors ...]
$ ./hello
Hello, world!
It works if the program is compiled with g++ also, in spite of a single character constant like '!' being char.We have to write the characters backwards because of little endian. In the source code, the leftmost character is the most significant byte, and on the little-endian system, that means it goes to the higher address. The first character H must be the rightmost part of the character constant so that it ends up in the least significant byte which goes to the lowest address in memory.
Endianness wouldn't be an issue in B because it doesn't exist; there is no "char *" pointer accessing the data as individual characters. B could be implemented on big or little endian and the string handling would work right.
a "hello"; b "world'*;
v[2] "now is the time", "for all good men",
"to come to the aid of the party";I’m curious why octal fell out of style. Hexadecimal seems more useful in every way. Perhaps it relates to using 36-bit words?
Literally the only octal I use is for chmod.
Yep, before ASCII was standardized it was common for machines to be built with word-addressable memory and words that were multiples of six bits. Two octal digits easily represent a six-bit byte, just as two hexadecimal digits easily represent an eight-bit byte
Why was six bits chosen? The modern use of eight bits seems more natural to me, being a power of 2.
https://en.wikipedia.org/wiki/Baudot_code
(with control characters to shift to other character sets!).
Wikipedia says that the several six-bit character set standards for text were inspired by typewriters
https://en.wikipedia.org/wiki/BCD_(character_encoding)#Examp...
but they didn't represent lowercase (!). (2⁶ would have allowed you to represent lowercase but you would have to sacrifice a whole lot to do so -- as alphanumerics alone would use 62 positions, leaving you with maybe one position for a space and one punctuation mark, and no newline...)
Apparently EBCDIC derives from IBM's 6-bit BCD codes
https://en.wikipedia.org/wiki/EBCDIC
and is interesting because it uses 8 bits, successfully represents lowercase, and leaves a ton of (non-contiguous) positions unspecified.
Maybe our standards (no pun intended) are just shifting as we deal with more and more capable software, but I'd be inclined to say that seven bits "easily encode" standard English text, and six don't, on account of the lack of case distinction. (Although you could certainly choose to handle that with control characters, and I'm sure some 6-bit systems did so.)
Interestingly enough B talks about how the new computer can address a “char” and not just a whole word at a time.
main( ) {
extrn a, b, c;
putchar(a); putchar(b); putchar(c); putchar('!\*n');
}
a 'hell';
b 'o, w';
c 'orld'; main( ) {
auto c;
read:
c= getchar();
putchar(c);
if(c != '*n') goto read;
} /* loop if not a newline */
People forget the world of ubiquitous goto enabled control flow with languages like this and Dartmouth BASIC.(Good programming practice dictates using few labels, so later examples will strive to get rid of them.)
and then later...
The ability to replace single statements by complex ones at will is one thing that makes B much more pleasant to use than Fortran. Logic which requires several GOTO's and labels in Fortran can be done in a simple clear natural way using the compound statements of B.
c = c+'A'-'a';
Though they too learn... C fixed this problematic syntax: x =- 10
So they are not that different from me - just smarter.So it looks as though C switched from default static to default auto? I wonder if the programmers of the time sneered at the waste this added?
In C each variable and function has two attributes, type and storage class
I think strong typing with type inference hits the sweet spot of providing the safety of types, while also cutting down on boilerplate.
Nowadays we do byte addressing for everything, but in PDP-11 machine code a byte address and word address are different AFAIK, so char and int pointers would be incompatible.
Also I don't know much about B but given its age I seriously doubt that it's dynamically typed like Python. I suspect that it's more like assembly: you don't have types because most of everything is effectively an integer if you squint hard enough, and the way you decide to use the data lets the compiler/CPU know how to treat it.
For instance in this snippet from TFA:
v[10] 'hi!', 1, 2, 3, 0777;
You may think that `v` is clearly dynamically typed, since it contains both ints and a character string, but I actually think it's a lot simpler than that: a pointer to a string is effectively an int, so you can store a pointer in an int array no problem. You can still do it (mostly non-portably) in modern C, you'll just need a cast or two at most.Of course it means that the type info is not actually carried by the variable like in Python. If you write `'hi!' + 'oy'` in Python you get 'hi!oy'. If you write 'hi!' + 'oy' in B I suspect that you get a garbage pointer, if not straight up undefined behaviour.