The most surprising Unix programs
minnie.tuhs.org
minnie.tuhs.org
paste allowed me to interleave to streams or to split out a single stream into two columns. I'd been writing custom scripting monstrosities before I discovered paste:
$ paste <( echo -e 'foo\nbar' ) <( echo -e 'baz\nqux' )
foo baz
bar qux
$ echo -e 'foo\nbar\nbaz\nqux' | paste - -
foo bar
baz qux
I wonder what other unix gems I've been missing... join -t $'\t' file1 file2 (BASH only, I think.)
join -t '<CTRL+v><Tab>' file1 file2
join -t "`echo '\t'`" file1 file2
Why 'join -t '\t' file1 file2' is apparently beyond the pale has me mystified.$ join -t '\t' join: illegal tab character specification
* John A. Kunze's jot and rs
* John Kerl's mlr ("Miller")
* jq
* join and comm, as mentioned
* fmt
* ex
And, given what you just wrote:
* printf
seq[0]
EDIT: I didn't see that you had already included 'jot' so removed it from my original reply.
0 - https://www.freebsd.org/cgi/man.cgi?query=seq&apropos=0&sekt...
20+ years in UNIX, never encountered it. Part of GNU coreutils.
You should debug for it to not happen though but probably better than seeing your system choke with dozens of instances even slowing the jobs more.
There's also "run-one" that allows you to control which job should survive instead of just killing the old one.
http://manpages.ubuntu.com/manpages/trusty/man1/run-one.1.ht...
(Of course, this can be surprising in a different way when you expect a new job to start, but it does not because a previous one is still lingering. Not saying it's better or worse, just pointing out for those who don't know.)
Writer's Workbench was indeed a marvel of 1970's limited-space engineering. You can see it for yourself [1]: the generic part-of-speech rules are in end.l, the exceptions in edict.c and ydict.c, and the part-of-speech disambiguator in pscan.c. Such compact, rule-based NLP has fallen out of favor these days but (shameless plug alert!) Writer's Workbench inspired my 2018 IOCCC entry that highlights passive constructions in English texts [2].
[1] https://github.com/dspinellis/unix-history-repo/tree/BSD-4_1...
Do you know where I could be able to find the source for all of these?
I'd be interested to "revive" these utils, possibly rewriting as python or bash for easy hacking. I have some basic scrips for that, and they are already proving to be useful even though they simply call grep https://github.com/ivanistheone/writing_scripts
The files are in cmd/wwb.
Other tarballs may have more Writer's Workbench code but I haven't looked at them.
[0] https://www.tuhs.org/Archive/Distributions/Research/Dan_Cros...
https://svnweb.freebsd.org/base/head/sbin/init/init.c?view=m...
Since the names were mostly Indian, we did not even have a standard database of names to test against.
What we did was the following: go through the entire database of all applications, and build a trigram frequency table. Then, using that trigram table, do a second pass over the database of names to find names with anomalous trigrams - if the percentage of trigram frequency anomaly in a name was too high (if the name was long enough), or the absolute number of trigrams in name was too high (if the name was short), we flagged the application and examined it manually. Using this alone, we were able to filter out a large number of dummy application forms.
Of course, it is not a comprehensive tool since what forms a valid name is very vague, but I think this kind of a tool is useful and culture-neutral.
Why not send them an email to verify? Or use captcha tech developed by big companies that actually has some science behind it.
The only reliable means of communications back to students is by a government approved website, or newspapers, or official media. (The process also has to stand up in court in case some student says that (s)he did not get the communication, and newspaper ads are a documentable evidence of communications on a specified date.)
(BTW, the forms do have captchas, the spurious forms are manually filled in by mischievous/malicious/curious applicants.)
Yes, they'll catch some extreme cases, but anything that looks like the normal bad case will go by unnoticed generally more so than if the human is doing all of the review work.
For those who don't know, Dough is the guy that invented pipes.
> Originators of nearly half the list--pascal, struct, parts, eqn--were women, well beyond women's demographic share of computer science.
An interesting statistic would be the gender ratio for people who published papers.
My source is my own family history and it makes me annoyed at the anachronistic assumptions and framework that people shoehorn historical tidbits into when discussing this topic.
By the way, how can you have 50% in the 60s and then a peak at 37% later?
I suspect that shifting narrative may tell a chapter of the story of how women were gradually pushed out of our industry.
Although the programmers might have looked down on the operators as grunts, these people were in the privileged position of actually getting to interact with the machine directly, which is something.
You're pretty much right on re their exodus. Thompson's book, which I mentioned in my earlier comment, has a chapter called 'The ENIAC Girls Vanish'.
My life experience corroborates this.
When I was in school, girls were taught to type, and boys weren't. Because of that, many of the girls from my school went into computing, while the boys went to more "manly" pursuits. It's also why Catholic nuns were over-represented in early computing.
Would you mind expanding on that?
Catholic nuns have done much for the business world that is overlooked.
From memory, so no citations:
The first female CEOs were Catholic nuns — in the early 1900's, a time when women in the American workplace was uncommon, Catholic nuns founded over 800 hospitals in the United States. They also ran schools, colleges, and universities.
The concept of just-in-time delivery was invented by nuns running those hospitals.
The first health insurance company was founded in Missouri by nuns to care for railroad workers.
There are others. I once ran across a list of them on the intarwebs, but these are the ones that stuck with me.
https://www.math.upenn.edu/~wilf/AeqB.html
Dad - a history major turned COBOL programmer and later, project manager - found his way into his prior company (and then stayed for 32 years) by this route.
He met Mom there too - also not uncommon for the time - who was a math major who decided that teaching wasn't the right fit for her.
This sounds really weird to modern ears but it's literally true. Computer was a job before it was a machine.
> pascal
>
> The syntax diagnostics from the compiler made by Sue Graham's group at
> Berkeley were the mmost helpful I have ever seen--and they were generated
> automatically. At a syntax error the compiler would suggest a token that
> could be inserted that would allow parsing to proceed further. No attempt
> was made to explain what was wrong. The compiler taught me Pascal in
> an evening, with no manual at hand.Most unixes have one, although the format differs.
https://unix.stackexchange.com/questions/164555/how-to-empha...
I am curious about this one, though, has anyone used it?
> The syntax diagnostics from the compiler made by Sue Graham's group at Berkeley were the mmost helpful I have ever seen--and they were generated automatically. At a syntax error the compiler would suggest a token that could be inserted that would allow parsing to proceed further. No attempt was made to explain what was wrong.
On the surface it sounds a lot like it would produce error messages like “expected ‘;’” that most beginner programmers come to hate: was it any better than this, or was that the extent of its intelligence and everything else at the time was even worse?
Do people really come to hate these? I'd expect the opposite -- that people would start off hating messages like "expected ';'", but fairly quickly become accustomed to what they almost always mean.
As long as you can look at the message and have a good idea of what's wrong, it's not a bad message.
> I’m genuinely curious to hear if there’s any strategies on improving these or work done in this area.
It's just a ton of work. You look at what the compiler produces, look at what information you can have at that point, see if you can craft something better, and then ship. And then look at the next error.
That's a fantastic attitude and I really appreciate that someone is working towards that goal, thanks.
> It's just a ton of work. You look at what the compiler produces, look at what information you can have at that point, see if you can craft something better, and then ship. And then look at the next error.
Exactly. That person is a saint.
On the other hand, a missing semicolon in Microsoft C would often give a litany of unrelated errors.
Compilers have gotten significantly better in the past couple years. But in the browser I'm using to write this comment is even worse than "unexpected x", since it gives me the type and not even the token:
var x = 1 console.log(x + 1)
SyntaxError: unexpected token: identifierI used an Algol compiler that had messages such as:
Semicolon missing after end (inserted)
Undeclared identifier ‘foo’ (assumed integer)
Both of these hugely improved the compiler output, as far fewer utterly useless error messages would be produced (yes, I know I didn’t declare ‘foo’. You told me so the previous 12 times I used it)Parsing valid programs is easy, so are bailing out or going into the woods when encountering invalid syntax. Producing meaningful error messages for line 100 after having seen errors on lines 13, 42 and 78 can be fairly hard.
Great learning tool, but also the ability to copy-pasta from terabytes of code ... (scurries off to do some analysis).
I think it's interesting that McIlroy was able to learn Pascal from it!
This sounds like something from the same family as hyperloglog
Wikipedia traces that back to the Flajolet–Martin algorithm in 1984. When would typo have been written?
If you count 5 "abc" and 5 "xyz" in a counting bloom filter, it will always say you had 10 events, but might say they were 10 of the same event.
If you count the same in Morris's structure, it will never confuse the two different sets, but might say one occurred 4 times and the other 8.
Of course, that means you can combine the two, for the benefits and downsides of both - storing very high (and inaccurate) counts of very sparse (and maybe misattributed) event sets.
Sounds like it's just doing something like replacing `counter++` with `if(rand() % counter == 0) counter++`, so that the counter will increase slower and slower the larger it gets.
Unix V5 was released mid '70s, but as others have pointed out, counting is different from count distinct.
Super powerful and saved me hours of work.
This may seem obvious, but there are many tiny ways that sorts can differ between locales, operating systems and programs (e.g. Excel), especially when dealing with Unicode. It may look the same 99% of the time, and you may not realize until later that you’ve accidentally filtered out values.
comm <(sort fileA.txt) <(sort fileB.txt)Plus diff was in part written by the author of the linked content :).
sudo apt-get install tkdiffIt allows you to create pipes as files! So you can do:
`echo 'hello world' > mypipe` on one terminal and `cat < mypipe` on another!
Very neat, I'm sure I'll find uses for it in the future.
And if that does indeed work, that's pretty cool.
https://dl.acm.org/doi/10.1145/2911981
Provides a good overview of how it works and perms website has more information.
What’s cool about the computable reals implementation is you can increase the precision after the fact and it will recalculate up to that precision. Basically it memoizes the steps of the calculation and how they affect the precision.
```
( ) (@@) ( ) (@) () @@ O @ O @ O
(@@@)
( )
(@@@@)
( )
==== ________ ___________
_D _| |_______/ \__I_I_____===__|_________|
|(_)--- | H\________/ | | =|___ ___| _________________
/ | | H | | | | ||_| |_|| _| \_____A
| | | H |__--------------------| [___] | =| |
| ________|___H__/__|_____/[][]~\_______| | -| |
|/ | |-----------I_____I [][] [] D |=======|____|________________________|_
__/ =| o |=-O=====O=====O=====O \ ____Y___________|__|__________________________|_
|/-=|___|= || || || |_____/~\___/ |_D__D__D_| |_D__D__D_|
\_/ \__/ \__/ \__/ \__/ \_/ \_/ \_/ \_/ \_/
``` $ cowsay "hey dude"
__________
< hey dude >
----------
\ ^__^
\ (oo)\_______
(__)\ )\/\
||----w |
|| ||fortune | cowsay
And voilà, you have a little quote running in your cow friend whenever you open up your terminal.
Also: Does anyone have any more good fortune files? I only have the fortunes that came preinstalled on Ubuntu but would love to have more.
We could've had prettier et al instead of style linters 40(+?) years ago. :(
The original Ratfor brings FORTRAN 66 nearly up to the level of a respectable programming language.
It turns this:
if (a > b) {
max = a
} else {
max = b
}
Into this: IF(.NOT.(A.GT.B))GOTO 1
MAX = A
GOTO 2
1 CONTINUE
MAX = B
2 CONTINUE
... with proper columnization, of course.Going the opposite direction is pretty miraculous to me.
Ratfiv is the follow-on, which did the same to FORTRAN 77. However, FORTRAN 77 had control structures beyond the conditional GOTO, so Ratfiv was somewhat less necessary.
FORTRAN 77 would look like this:
IF (A .GT. B) THEN
MAX = A
ELSE
MAX = B
ENDIF
https://en.wikipedia.org/wiki/Ratfiv[0] https://www.goodreads.com/book/show/515603.Software_Tools
That has almost nothing to do with the spaces-and-braces nitpicking of prettier/gofmt etc.
When I started programming in the 90s, "spaces-and-braces" checking - as you say, nitpicking - was basically all we had, along with limited automatic tools to fix them (all more or less as good as `M-x indent-region`). If you were lucky and in a widely-used language you could cobble together compiler warnings, lint, and a few other tools to also get warnings about legacy interfaces (gets), dangerous practices (ignoring error codes), and unusual structure (shadowed variables, loop conditions that seemed impossible). Today we finally have considerably better tools that don't just check if you match a style guide but do a full reformat (not nitpicking, but doing it for you) and linters that can enforce 'deeper' structural demands, sometimes with automatic fixes.
But 40 years ago we had tools to completely restructure programs to a normalized form, and the practical experience to know programmers found this form preferable! And like so many things in our field, 10-20 years later we had to rediscover it, painfully, all over again. Probably because today's programmers think source-to-source Fortran/Ratfor translation has "almost nothing to do" with the challenges facing them today.
By doing this I really read the code, and really get to understand what the previous programmer was doing.
Also something I took from Asimov's foundation series, code that doesn't look right, doesn't run right.
I know, not really; compiler gives no fucks, but I'm not a compiler however, and GCC error messages (Clang too!) are still about as useful as a hot bikini wax is to a walrus.
This was one of the features that made me really fall in love with Emacs back in the day! I could set it up to force my style requirements, then even yanked (pasted) code would be proper (mostly) and I couldn't fat finger my code to death.
My only emacs complaint is lisp. I get it, I just don't like it. I'll take fortran 77 over lisp any day (not 66 or before, tho. I'm not that crazy). So, sorry mr(s) moar lisp.
histogram - simply counts each occurrence of a line and then outputs from highest to lowest. I've implemented this program in several different languages for learning purposes. There are practical tricks that one can apply, such as hashing any line longer than the hash itself.
unique - like uniq but doesn't need to have sorted input! again, one can simply hash very long lines to save memory.
datetimes - looks for numbers that might be dates (seconds or milliseconds in certain reasonable ranges) and adds the human readable version of the date as comments to the end of the line they appear in. This is probably my most used script (I work with protocol buffers that often store dates as int64s).
human - reformats numbers into either powers of 2 or powers of 10. inspired obviously by the -h and -H flags from df.
I'm sure I have a few more but if I can't remember them from the top of my head, then they clearly aren't quite as generally useful.
Anyone else have some useful scripts like these?
I understand there's vim plugins for this, but, ehh.
* http://jdebp.uk./Softwares/nosh/guide/commands/console-flat-...
Is this much different than `alias histogram="sort $1 | uniq -c | uniq -nr"`
Sidenote: I started https://github.com/jldugger/moarutils as a means of publishing and sharing these, but it turns out I don't even have a lot of dumb ideas. Will probably end up bookmarking this HN post for "later."
Stuff like this really makes me love what the pioneers of CS did in the past. In the past, they were counting every byte and every register while nowadays, programmers make things without considering the impact it will have on the HW.
I wonder if compilers could do this today? If you can bound values for floating point operations, you might be able to replace them with fixed point equivalents and get a big speedup. You might also be able to replace them with ints or smaller floats if you can detect the result is rounded to an int.
CPU's also have the possibility to do this since they know (some of) the actual values at runtime, and could take shortcuts with floating point calculation in places where not needed for the result.
This could make sense for SIMD however, but then the problem is getting the array data in the right format before the computation — if you’re converting from float to int and back within the loop, it destroys any performance gain.
Perhaps a good example of that is video encoding, which is mostly fixed point, despite it looking like a pretty close fit for floating point maths.
Video encoding is a bit of a special case though because the common algorithms are carefully designed for hardware acceleration. For most rendering, it doesn’t make sense to go out of your way to avoid the FPU.
https://dl.acm.org/doi/10.1145/2911981
Is a nice overview.
With regard to roff in general, when I got into Linux-based typesetting around the turn of the millennium, that was already seen as antiquated tech, superseded by LaTeX which was undergoing a frenzy of development and improvement around that time. So, anyone under the age of 30 will probably be hearing of such *roff stuff for the first time (and sadly even familiarity with LaTeX has waned).
I'm also using TeX/LaTex, but it's still a programming language whereas roff/eqn etc are non-Turing DSLs and renderers for particular narrow purposes. I get your point, but saying these are "antiquated" is like saying HTML is obsoleted by JavaScript.
My (very recent) university education was unfortunately quite light on UNIX folklore, but this was converted in our formal automata course as we traversed the Chomsky hierarchy.
I know a number of projects that generate their roff by using pandoc. They don't actually know, or have the inclination to learn, exactly how g/roff works.
I think you mean "Thompson NFA construction" and "NFA->DFA."
Regardless though, this is not what the OP is pointing out. 'egrep' (or just GNU grep these days) is doing something more clever (emphasis mine):
> Al Aho expected his deterministic regular-expression recognizer would beat Ken's classic nondeterministic recognizer. Unfortunately, for single-shot use on complex regular expressions, Ken's could finish while egrep was still busy building a deterministic automaton. To finally gain the prize, Al sidestepped the curse of the automaton's exponentially big state table by inventing a way to build on the fly only the table entries that are actually visited during recognition.
Russ Cox talks about this a bit in part 3 of his articles on regex matching[1]. Its implementation in RE2 is here: https://github.com/google/re2/blob/master/re2/dfa.cc
Yep, only noticed it later, then left it in to see who's paying attention :)
It was amazing that you could write pretty complex math with just a few special literals like “sup” and “sum”, and the braces. It turns out that compositionality is so strong in math that it’s most of what you need. This isn’t obvious until you try it!
Paired with a LaserWriter (vintage 1986, say), and troff, you could get almost book quality typesetting.
Later on, TeX got the details of math much better, but the basic language was the same.
...there goes my weekend.
[0] https://github.com/dspinellis/unix-history-repo/blob/Researc...
[1] https://github.com/dspinellis/unix-history-repo/blob/Researc...
% uname -a
FreeBSD skyrocket 9.3-RELEASE FreeBSD 9.3-RELEASE #1: Fri Nov 27 20:28:19 UTC 2015
Earlier filesystems were trying much more to be like databases.
Here is a paper from Bell Labs
Read manpage before trying it.
Once after blowing up an in production database server during the day, I suffered the unfortunate difficulty of having to explain why running a command "killall" on a critical server that killed everything was an innocent mistake and that I didn't have any reason to expect it to kill everything.
It's extremely difficult to not sound like a moron when explaining that you didn't expect "killall" to "kill all".
The peoples' names were more recognizable.
all are so exciting!!