The tyranny of the Hollerith punched card
pub.gajendra.net
pub.gajendra.net
Coincidentally (or perhaps ironically) the article itself is formatted for ~80 characters per line, not because it's bound by the Hollerith card, but because 80 characters is a good ergonomic width.
Maybe because it was written from a terminal?
And why does my terminal window have a baud rate? Why is it so slow?
In my grade 10 "Data Processing" class (when was the last time you heard someone use the term Data Processing?), we had to use IBM cards with a Sharpie (the card reader was thankfully an optical reader vs a punch card reader) to write programs in BASIC (when BASIC still used line numbers). When we had tests, we had to basically write our programs at our desks on cards, run to the card reader to run and check our results on a printout. It made for some hilariously tedious debugging cycles.
If my memory serves correctly, back in that era, having an 80 column card on your Apple II was something worth bragging about.
Hey, with punched cards you had: a) free confetti, b) gunfire soundtrack.
You were deprived and are in the Nile.
[1] https://www.python.org/dev/peps/pep-0008/#maximum-line-lengt...
The 100-character allowance is for code "maintained exclusively or primarily by a team that can reach agreement on this issue".
~80 characters is a good, readable target, but that's not what we have. For a long time, we've had a militant standardization on <81 characters, even in cases where the results are obviously worse than a long line. I imagine that all of us have read or written code containing a line break that blatantly decreased clarity.
So the question remains: why 80? 80-ish, sure, but why such a fixed standard? I don't know enough history to say if that's down to Hollerith cards, or the consequences of auto-formatters, or something else, but there's a more definite question than "what's a good width" here.
Because 80 is a very good approximation to 80-ish.
[1] http://www.righto.com/2016/05/inside-card-sorters-1920s-data...
I don't think that's true at all. While 80 columns becoming initially common for displays was no doubt a product of the H-card, the fact is that there was reasonably widespread support even fairly early in the pre-GUI PC era for alternatives -- both more and less, but probably most notably in use 132 columns -- for display, printing, etc., and they were used where they offered superior utility. But, outside of a few niches, 80 columns stuck even in the face of alternatives, and the reason it did so was ergonomics.
Once you have the number of rows, you want the number of columns to have a good balance. 80 columns gives you well proportioned characters, where other alternatives e.g. 132 do not. Add that to the historic precedence and you have the makings of a standard.
Teletypes were in common use on the first time-sharing systems, and those were only 72 columns wide. Consider yourself lucky to have those extra 8 characters.
http://www.seasip.info/VintagePC/Images/cga.png
So then you need one more row to separate lines (or live with the occasional joined characters, as CGA did with its 640x200 resolution).
> When the Type 80 sorter was introduced, standard AC power hadn't fully taken over and parts of the United States used DC or 25 Hertz AC.[9] Thus, the sorter needed to handle fifteen different line inputs including unusual ones such as 115V DC or 230V 25 Hertz AC.
For car width, I think that that's also a pretty reasonable width, since it's the perfect width to fit two people side-by-side comfortably, with a bit of room in the middle.
Now we only need to think of things that have developed for historical purposes, yet are not reasonable defaults.
Considering the numbers of programmers in the world, and how long they've been working with text, and how much annoyance text breaks / line wrapping / format flowing causes it's kind of surprising that there isn't better ways of wrapping lines without line breaks.
What's current best practice? Or is it just down to the project style guide?
There are more such rules which differ between standards (like Linux kernel style vs. PHP PSR-2 style), e.g. how to pad trailing lines of function call. Linux style says to align trailing lines at offset next to the opening brace. (This doesn't look good in case of long function name, though, but most of time readability of this way is very good.) K&R is referenced as an original of such style, just as everywhere in Linux coding style. Also Linux coding style is not just what is told in Documentation/CodingStyle - there are some unwritten tools which you'll encounter when you run scripts/checkpatch.pl or get reviews on your patch submission.
I actually was just contemplating ditching the 80-char limit the other day, at least for personal projects.
We've all got widescreen monitors now, right? So it makes sense that people are going well past 80 characters.
Sure we do, but some/many of those widescreen monitors are only eleven or thirteen inches diagonally, and can't reasonably fit two buffers of text 120-chars wide without shrinking the font to eye-straining sizes (yes, even on retina screens).
Plus, this only gets worse as you get older and your eyesight starts to diminish.
Yup, I'm at that point.
For the same reason, newspapers use multiple columns, even though the paper itself would allow for 200-character lines.
qry.setParameter("a", a);
qry.setParameter("b", b);
qry.setParameter("foo", foo);
qry.setParameter("bar", bar);
After the first "qry.setParameter" you're no longer reading anything but the parameters and your eyes are moving vertically.Or if you have something like this:
Query qry = getEntityManager().createNamedQuery("User.findUserBySomething", Long.class);
You're only reading the "User.findUserBySomething". The rest of it is mainly fluff.It's also worth noting that newspapers traditionally used columns because they allow individual articles to be typeset and then divided into pieces later during composition. Not because they offer some kind of readability advantage.
I think on a mental level, the problem is that when the lines are this short, the grammar is extremely broken up, so you need to "cache" 4-5 lines of context to understand anything. When it's 80 per line, lines are parsed in a more independent fashion.
I think 70-100 is perfect, depending on other factors.
Screenshot of what I mean: http://www.trbimg.com/img-5667d3f6/turbine/la-la-na-menace02...
[1]: https://tkware.info/2016/05/08/a-case-for-increasing-pep8-li...
There are contexts when Python doesn't force you to, such as the example the article shows as "ways [that] actually make our code less obvious and less readable":
err("HTTP 403 Forbidden -"
" Does your bitbucket user have rights to the repo?\n")
But if you indent it properly, I don't think it's less readable than the extra-long-line-with-scrolling: err("HTTP 403 Forbidden -"
" Does your bitbucket user have rights to the repo?\n")
The other example, build_url = "{url}/pipelines/{pipeline}/jobs/{jobname}/builds/{buildname}".format(
url=os.environ['ATC_EXTERNAL_URL'],
pipeline=os.environ['BUILD_PIPELINE_NAME'],
jobname=os.environ['BUILD_JOB_NAME'],
buildname=os.environ['BUILD_NAME'],
)
could perhaps be written as urltmpl = "{url}/pipelines/{pipeline}/jobs/{jobname}/builds/{buildname}"
build_url = urltmpl.format(
...
or like this: build_url = "/".join((os.environ['ATC_EXTERNAL_URL'],
"pipelines", os.environ['BUILD_PIPELINE_NAME'],
"jobs", os.environ['BUILD_JOB_NAME'],
"builds", os.environ['BUILD_NAME']))(Coincidentally, this is why the separator there is a dash, rather than the more grammatically correct comma)
And I thank you for the other examples of ways to handle the URL, but again we've moved into making changes for very minor benefit just to please the analyzer, rather than the human reading the code.
I used to work at a dysfunctional startup which could not agree on any sort of coding style conventions, and one of the other devs would just zoom his Visual Studio window fullscreen and type, type, type until he bounced off the far edge, no sense of column width at all. His code was excruciatingly terrible for a lot of reasons, but at least it was easy to spot since it looked just as messy as its structure.
How do you know the width of a standard human isn't determined by 2000 years of people choosing partners that would fit beside them on a cart comfortably.
It gets tiring listening to devs rage about the 80-char limit, then go on to produce code that's basically unreadable because it scrolls way off the screen.
Citation needed.
Hacker News keeps clean discourse by downvoting comments that are not relevant or do not stimulate debate.
I'm also not convinced that 80 chars "happens to line up quite well with the human eyes ability to track horizontal lines of text". It's claimed often, but in the absence of any actual studies, the best you can really say on this is that it is your subjective experience. Mine is different - I find 100-120 columns per line to be eminently more readable.
Until the aforementioned study determines what the average peak readability length of _code_ (which I think is likely to be different from regular text, because the semantic units are different between the two), we'll just have to agree to disagree.
http://infolab.stanford.edu/pub/voy/museum/pictures/display/...
So, given that, perhaps 80 characters isn't so tyrannical...
class A
{
A::a()
{
if (expr) {
do1();
} else {
do2();
}
}
}And, believe me or not my space-loving friends, tab-indented codebases still exist.
Code is different in several ways. Firstly, code is generally shown in monospace typeface. Secondly, breaking lines too early can be even worse for code than for prose. Thirdly, even with a long line limit, code will generally have lots of short lines anyway, and this helps the eye find the next line when reading. The long lines (as long as they aren't all long) stick out from the rest of the code text, making it easier to read.
Second, we code very differently then we used to - several layers of indentation are common now but were not when the limit was seen as reasonable.
Except for Linux Kernel people (who actually have to work and debug stuff in the builtin 80x24 screen) there is no excuse for not using something bigger
Having to break a line because it's 81 characters and your linter is going to complain is completely and utterly ridiculous
It's 2016. How long since the first IBM PC again?
Yes, it's a tyranny. I don't care if the average is around 80 columns, I care about having a hard limit
My current terminal is 110 chars wide, and that's because I use a bigger font.
The same kind of "cruft" you see in systems and designs that have been around for a long time (Why is it the "Referer" header in HTTP? Why do we still call it Ajax/XHR when XML is rarely used?) is seen in many aspects of life.
Etymology is choke full of them. Consider the word, "consider" - its roots may have meant something like "consult the stars" at one point in time. (I think I first read about this in Sagan's Cosmos)
(then again, maybe the original keyboard maker put them there for ergonomic reasons)
80 columns is a pretty good length for lines of text.
I see no tyranny here.
When compatibility is a burden, people do work hard to break it.
Anyway, there's a "(2012)" missing on the title, and even that makes the article kinda late to the discussion. By that time, most people already didn't care about the 80 lines limits.
When the card reader jammed, it destroyed the card. We had to retrieve the bits of card, and tediously re-create the card on an IBM 026 keypunch, one column at a time.
One time a card got caught between two pinch rollers, and smoke started billowing from the machine.
IBM sold the "Multi-function Card Machine", or "MFCM". Customers called it the "Mother Fletcher Card Masher", except they had another word for the "F".
The computer's card punch could only do 200 cards/min.
In humid weather, cards would swell and jam the reader's picker knife.
I was so glad when we switched to mag. tape.
I believe you're allowed to say "Fuck" and it's various declensions on HN now!
I'm always particularly intrigued by people who write things like "Sh#t" or "F#ck" because it's unclear who they are censoring for, and indeed why. If it was simple prudishness, I assume they would either use some twee alternative or mis-spelling, but simply eliding one letter suggests they themselves have no problem with expletives. It certainly doesn't protect viewers from seeing the taboo word, as a string of punctuation like "#&$!" might, so that isn't the purpose. Wildly off-topic, so apologies...
If you actually use long/descriptive variable names, you can mentally process much wider text quickly, invalidating the research the "80 character optimization" was based on.
Personally I use a 145 character width, because that fits nicely on my portrait-oriented monitors, with room in both left and right margin for the editor to highlight/mark. 80 character wide limits on a modern project is like an adult riding a "Big Wheel".
foo = bar(baz, qux)
myClassInstance.someWritableVar = lib.someFunction(param1, MyOtherClass.constants.FOO)
The second line is well over 80 chars, but may be as understandable as the first. const uint64_t *p=bsearch(&ref->u,
df->file_offsets,
df->num_file_offsets,
sizeof df->file_offsets[0],
&CompareU64s);
That's from emacs. It's a bit hit or miss whether editors do something nice like the above, or just indent the arguments by one stop, but other options seem to be fairly rare.I generally have one function call per line for ease of debugging. Nothing worse than having to do a fiddly step in/step out dance when you're trying to think about what's going on. But it's also good for keeping on top of line lengths too.
Which is exactly why all such paradigms are _awful_.
Do you have any defense of that? I'd be interested to read it.
I think every article on coding advice I've read for the last decade favors long and descriptive function/variable names.
But long variable names?!? There is no place for such an abomination under this sun.
You mean field names? A paradigm where you have such a thing is _awful_. Seriously. And this is exactly one of the reasons why it is awful and why it leads to an unreadable and unmaintainable code.
And no, I don't believe you. Making your variable names more descriptive does not make your code less readable. Even when it's not strictly necessary, it's not harmful to comprehension UNLESS you refuse to move from a completely outdated line width.
Some people argue that it takes longer to type and edit. But as we all know, we spend far more time reading than writing code, and modern editors (like vim!) have solved this non-problem anyways.
And, yes, OOP is a filth.
Think about it: there's a reason why human languages are filled with pronouns, because if you use the full name for everything, communication will soon become tedious and it will become actually harder to understand.
Any desire for a long and descriptive name should be tempered with a matching desire for conciseness. Otherwise you end up with names like MaybeUpdateDisplayParameterListForValidation. Throw twenty of these names onto a screen, and I have no idea what the hell is going on: I can't even figure out which names are the same at a glance.
The "break" and "span" terms were familiar from a time when more people knew the Snobol language, which has two frequently useful pattern matching operators: BREAK and SPAN. Snobol's BREAK(S) pattern matching operator matches the input up to but not including the single-character match for any of the characters in set S. The set S delimits or "breaks" the sequence. SPAN(S) matches a sequence of one or more characters from the set S.
strpbrk tries to fit "string" "pointer" and "break" into a symbol that is different in the first six characters. It was once a common linker limitation that only the first six characters of external symbols were stored.
Actually the C function which corresponds to the concept of BREAK is strcspn (complemented span), because this gives the (length of) the range characters up to the first match in the set. That is to say, strcspn could have been called strbrk! Then we would have had strspn and strbrk as a complementing pair. In any case, the strpbrk function points one character past the substring indicated by this function; giving a pointer to the breaking character. I think, the following equivalences hold:
strcspn(str, set) <--> strpbrk(str, set) - str;
str + strcspn(str, set) <--> strpbrk(str, set);
which further supports strbrk as a good name for strcspn.Trivia: break and span appear as functions in the Scheme SRFI 1, by Olin Shrivers [1998]. I think these correspond to the take-while and drop-while in Clojure and imitations thereof like the Emacs Lisp dash library.
But there's a reason that this sort of name has been left behind in modern APIs.
Short names are not left behind in core languages. For instance, a function that gives the length of a list, string or other sequence is often called length or len, and not length_of_sequence or whatever.
Arc, in which HN is programmed, has reduced "lambda" to "fn". "fn" is the same sort of shortening as using "pbrk" for "pointer to break".
Ruby has shortened "print" to "p".
For the basic core of a language, shortening names is good. When you're reading code, the short names by their very brevity tell you "I'm a thing in the core language, and not some external API to some add-on lib", which also has connotations of "I might be useful in many contexts; it may be worth it to learn about me and remember me".
You do have to admit the slight irony of the shortened string-handling function names in C, given this fact.
A couple 14 char class names and some indentation and suddenly everything is a formatting contortion, even if the average content length of a line is well under 80, occasional outliers and moderate indentation (at least 3 or 4 indentation levels are common in object oriented languages from class declaration to control flow) and it's well worth it to use 120 char lines.
(And if you write Java you probably need even more space to handle stuff since class names are often a ridiculous number of words long...)
Outside of programming languages, I also think the 80 char limit is terrible for readable html or html templates.
http://vt100.net/docs/tp83/chapter5.html
> 132 column × 14 lines (VT100, VT101, VT125) display
> Allows you a larger format for detailed or spread sheet work so you can preview reports prior to printing. (Note: The advanced video option, which is standard with the VT102 and VT131, provides 132 column × 24 lines display.)
"10 pixels wide for 80 columns, 6 pixels wide for 132 columns."
If their baseline was 800 pixels (80 cols * 10 pix/col) then 132 is the largest number of columns fitting into 800 pixels at 6 pix/col.If you're in the bay area, you should go see the IBM 1401 demo; it's cool to see a punched card computer in operation. Demos are Wednesdays and Saturdays; schedule is http://www.computerhistory.org/hours/
They've got contact info on their site: http://www.computerhistory.org/contact/
It may be worth it to send an e-mail to see what you could work out.
Nowadays, most people are using graphic displays, at a very high resolution, with beautiful, mathematically described fonts, which can be resized at will. So why are we restricting ourselves to an arbitrary number?
80 seems too low, no matter if you are writing javascript callbacks or writing Scheme code. Or doing python with a 4 space indentation. But it may be too much for older folks.
I feel that the reason we have computers is to deal with this kind of stuff, so why can't we do whatever we want and have the computer display the way we like?
> But breaking the statement arbitrarily at some point
> seems more like saying that you have to end a sentence
> at the end of each line, whether the sentence is finished
> or not because, you know, that's the way it's always been
> done. (I know it's not a perfect analogy.)
Isn't it more like line-breaking sentences at the edge of a paper page, which we do, and which is perfectly reasonable because of the page's fixed size?Sure, you can design your document for a different paper size, but then you have to reconfigure the printer for that size, make sure you have the right paper size loaded in the printer (and change it back afterwards!), get new envelopes in the new size, get new binders...
Similarly, if I read a program wider than 80 characters, I have to resize my terminal window and my Emacs window. When I then switch to working on other programs in the same windows, those programs then tend to be infected by wider-than-80-character lines, unless I remember to resize things back to normal, which in turn perpetuates the problem. So either the applications would have to constantly resize themselves when I switch between files (which seems really annoying) or we could just decide to stick to the standard 80 columns which, as many comments here have pointed out, is a reasonable choice to standardize on even if it's arbitrary.
Should be an option to auto resize or auto wrap
> When I then switch to working on other programs in the same windows, those programs then tend to be infected by wider-than-80-character lines unless I remember to resize things back to normal, which in turn perpetuates the problem.
Why would a program care about what size the window was when you ran it in the past?
> So either the applications would have to constantly resize themselves when I switch between files (which seems really annoying)
Isn't that exactly what a GUI is for?
http://www.throwcase.com/2014/12/21/that-five-monkeys-and-a-...
In one of my early jobs, we used binary format to store data in Hollerith cards. Twelve rows were divided into three groups of four bits, giving us three digits per column, or 240 digits per 80-column card.
The cards looked like lace doilies.
The first railroads in the US and UK standardized fairly quickly on 4'8½", but that's not the only size (even today!). The Great Western Railway, for example, was infamously at a 7' gauge. Even in the US, most railroads in the south converged on a 5' gauge until they shifted to 4'9" (not 4'8½"!) on May 31, 1886. The gauge differences in very early (1820s and 1830s) railways were just as diverse as the coal tramways that proceeded them--which means you basically have every value between 4'6" and 5' as the actual track gauge.
Of course, that belies the fact that your track gauge does not determine what you can stick on a train. That's the loading gauge, and these are nowhere near as consistent. The UIC standard loading gauge in Europe is 3.15 m wide, while the standard US loading gauge is 3.25 wide--and they both share the 4'8½" track gauge. (Note: the actual adherence to these standards, particularly in 100+ year-old railways, is a different matter.)
That gentleman just made me make my powershell terminal choke.
I can understand the benefits, though, to having sane rules around "how much can go in one line", I just think the overall width of a line is of way lower importance than breaking things up into logical parts -- especially in C# in code that avoids LINQ statements[1] where method chaining can result in some extremely wide, difficult to parse, code. I write almost all of my software on a Widescreen 1080P display where vertical real-estate is vastly lacking and horizontal real-estate is greatly underutilized. I use a pragmatic approach over artificial hard barriers and the assumption that I'll be using my display to look at code the way I do 99% of the time - one or two code files on screen at a time (but almost always a single code file).
My code still tends to mostly prefer vertical orientation. I use these simple rules:
(1) Avoid same-line chaining (Even "ToList" or "ToArray" gets its own line). Line up the dots of the calls with the previous line's dot to make it easy to see it's a part of a chain. This results in "wider" code but makes that code able to be parsed quickly while not wasting vertical space.
(2) Avoid more than one statement per line (conditionals excluded where readability requires -- doing if (isThis && isThat && string1==string2 && string3==string4) would put the two string comparisons on two separate lines usually).
(3) Omit (yes, omit) braces for single statements that don't have indented single statements below them. This gets me yelled at sometimes, but my IDE enforces code formatting for me, so the drawbacks of "accidentally having a statement appear to apply to a conditional that it doesn't" don't happen and with the lack of newlines-for-newlines'-sake in my code, code that's indented for branched statements is easily discernible from code that has a newline due to chaining or some other reason.
(4) Break the rules if readability is improved and avoid religious wars.
At the end of the day, I can live with nearly any code formatting ruleset. I read code more than I write it and I end up reading code that follows a large variety of rules so I try not to be a pain in the ass about it. If I haven't been the most substantial contributor to a project, I follow the other person's rules. If I have, I hope that they'll follow mine. The result has been few, if any, arguments over "tabs vs. spaces" theology.
[1] I have never liked the LINQ statement format as I find it hard to follow "What's Being Done" so I have, with rare exception, stuck with method chains.
I've tried rotating my monitors so they are long in the vertical direction, which is great for reading, but doesn't account for the space needed on either side of a document for supporting information, or what have you.
In my basement office, I have both of my displays set at a vertical orientation. When I took a new job, which has me working full time out of my home, I vowed to get used to using my laptop's display (and purchased a laptop based almost entirely on the display characteristics) because I wanted to be able to code wherever I was. This resulted in my writing an extension for Visual Studio to help identify certain code blocks I often scroll around aimlessly for and using parts of the IDE that I hadn't often used due to being able to see even terribly written classes all on one screen.
With a few adaptations, which included relaxing my style rules, I'm as productive on my 1080P display as I am with those two monitors. Some things are more difficult, but its made up for by being able to get up and work outside, up the road or on the beach up north, when I run into a wall with a coding problem. I require fewer breaks because I work feels less like work when you can change your environment on a whim rather than adapt to it.
Frankly, it sounds like numerology, to me.
"Hey! Both these numbers are 80, and they describe closely related dimensions. No way that's a coincidence! The newer one must have been specifically chosen to match the older one."
While it could of course be true, until I see reasonable proof, I'm going to assume that both values were the outcome of fitting a data representation (holes or characters) into a size that was effectively fixed for other reasons (standard card size, commonly available paper and CRT widths).
Especially since IBM later put out 96 column cards.