What you need may be “pipeline +Unix commands” only
nanxiao.me
nanxiao.me
Eventually, they'll become the ones that decide the fate of software engineers (by being hiring managers, etc.) and we'll see more and more monstrosity like the article portraits, instead of cleverly using UNIX tools where applicable.
There's so many things that the software world is doing wrong that I am surprised that even at this inefficacy, it's such a viable and well-paid profession. It's almost as if we are creating insanely complex solutions that in turn require a large amount of developers to support them, whereas we could have chosen a much more practical solution which is self-sustaining.
I think the problem is scale. Back in the day (before I was born), very few people were programmers and the resources they could use were limited. This means they didn’t need insanely complex solutions because they already needed complex solutions just to make it work on the limited hardware. People were trying to solve problems with computers. Nowaday you take a problem that could be solved by an microcontroller with three buttons and make it a cloud app with web server, web interface and all kind of other things like containers.
We donlt really tend to ask the question what a good solution would look like. Often it is the case that you just use the technology the developer wants to learn
These crazy complex solutions also look a lot more difficult on the slides than the simple 3 layer architecture that’s often shown to me.
With something that simple, and requiring so few people, how am I ever going to convince my clients to pay me multiple millions of dollars for it.
If you've jumped straight into programming, you'll probably consider any of those problems as a nail to your C/JS/Java/Python Hammer.
I was lucky to be initiated to the GNU / UNIX toolset by operation folks when doing tech support in a SAAS biz. We were dealing with a lot of text files and it didn't feel right to offload whatever my problem was to them, so I started learning from them and the scripts they wrote.
Mostly trivial stuff, such as in a directory, only select the relevant files, iterate through them to sum up something or find exactly the bit of information you need.
This allowed me to decrease time spent looking for the answer to my current case or colleague's case.
Of course, I then built those functionalities into a .bashrc (or was it bash_profile?) script to help my fellow support folks finding the answer on their own..
Turns out, most people don't want to have anything to do with a command prompt, even if the hard part has been done for you. That's been a pretty good lesson BTW.
Fast forward a few years later to now and I'm realising a good amount of stuff I've been writing lately could potentially be GNU'd instead of writing a Python script.
So that's how it goes I guess.
Well, the joke is on them for choosing to only stick to the GUI.
That's... very bizarre. Amongst the engineers I work with it's always an "oh cool, you know how to use bash/terminal" rather than "weird, why are you using terminal".
Me too. I'm thinking of doing a presentation on the Dunning-Kruger effect to see if it sparks some introspection in my colleagues.
My experience has been that some folks are resistant to the command line (but I wouldn't say most). This is too bad because I feel like it's a crucial part of development. I even wrote a post about it: https://letterstoanewdeveloper.com/2019/02/04/learn-the-comm...
So obviously, whenever I started doing trainings on how to use the cli / the toolkit, I could see what is best described as mild panic. Which I get, CLI isn't inviting at all, it's pretty daunting, there's no real emphasis on safety (as in not breaking anything), it's far from being easy to reason with when you're used to the Windows / GUI world.
The lesson was not so much about the prompt and more about understanding your target audience and catering to them really.
import pandas as pd
data = pd.read_csv(filename)
print(data.sum())
and have the same result, I'm going to do the one that is faster to write, fewer characters, and lets me understand what's going on.And don't get me wrong, I've written some gnarly pipelined bash before, although I'm by no means an expert, but that doesn't mean its always the right tool for the job.
To be fair, what the matrix libraries do is provide readability and clarity.
Show a programmer unfamiliar with awk your statement and they're going to be spending quite a while parsing it.
Show a programmer unfamiliar with numpy/pandas some matrix multiplication and they may understand it intuitively for the most part without even having to look up references.
edit: I'm actually an awk noob myself but after rereading your code for a second, it makes quite a bit of sense. "For all rows in the column, take all digits 0-9 and sum them". So perhaps not much is gained by the library
The company I recently started with is really big on Splunk.
The fact that they're proud enough of coming up with the tagline "Taking the sh out of IT" to print it on branded t-shirts featured in their training material was a hint that I wouldn't be a huge fan of the product, personally.
Abstracting things away is great, but something about IT pros being proud of avoiding the command line rubs me entirely the wrong way.
Also, I'd like to point out that in both this blog post and the Taco Bell one, the "UNIX way" is being compared against a straw man example, not a real example of over-engineering. Neither post actually provides any evidence of any real inefficiency. Both authors are just trying to explain and improve their own thinking about programming, not trying to cast judgement on a generation.
I think that by itself isn't a problem, but fading right along with it is the capacity to decompose and structure the problem domain.
Even if one ends up writing a solution in a different language for whatever reasons, starting out by mapping the problem with UNIX command line tools will result in a better understanding of the problem; an understanding that is language agnostic and can be transferred to any preferred method of implementation.
Maybe it’s Python instead of Perl but that is about the most-significant change.
I think git has a pretty significant amount to do with the persistence of command-line tool use among the younger set -- it's just so much more efficient / easier to use from the terminal.
People said the exact same thing in the 1990s, too. "Enterprise"-managed projects, using tech stacks like Win32 or Java, have generally tended to produce large, unwieldly monoliths.
Societies are large and full of random fluxes and waves.. right now it might be the time for Wirth 17 pages long solutions .. but McIlroy one liner will come back.
It's not so much about "tiny" but about no fuss, low complexity, no-nonsense, etc. For example, Forth may be tiny, but it's unusable for most applications. Assembler languages are tiny, too.
The problem with this type of thinking is that "composing tiny bits" is obviously good to anyone. BUT there are vast differences between composing unix commands and composing lisp functions (and composing tiny JS libraries in NPM etc.) and some of these are good, while some are not.
It's actually very difficult to properly create a system out of composable, re-usable components. I would actually say that Unix pipes are a particularly bad model of how to do that (as are NPM micro-libraries). No well-defined interfaces, arcane naming etc. - they are all reasons why most modern versions of these tools have grown and grown.
> right now it might be the time for Wirth 17 pages long solutions .. but McIlroy one liner will come back.
In case someone is interested in &ollowing up this reference, it’s Knuth’s 17 page long solution.
For what it's worth, I'm a first-year computer science major and the class I'm currently taking is very much focused on the "art of Unix." We've been doing shell scripting and regular expressions and the like. I quite enjoy it.
When all you've got is a computer and your own hands, you use the Unix way.
If you have a team (or the money to hire a team) of 20 developers, then you start looking for a problem that you can solve with that hammer.
So I wouldn't worry to much. All other small companies who have people who don't know better, don't have too much technology inside anyway. Or at least often enough.
For example, if you want to do uniq without sorting the input, that's:
awk '{ if (!($0 in seen)) print $0; seen[$0] = 1; }'
This works best if the number of unique lines is small, either because the input is small, or because it is highly repetitive. Made-up example, finding all the file extensions used in a directory tree: find /usr/lib -type f | sed -rn 's/^.*\.([^/]*)$/\1/p' | awk '{ if (!($0 in seen)) print $0; seen[$0] = 1; }'
That script is easily tweaked, eg to uniquify by a part of the string. Say you have a log file formatted like this: 2019-03-03T12:38:16Z hob: turned to 75%
2019-03-03T12:38:17Z frying_pan: moved to hob
2019-03-03T12:38:19Z frying_pan: added butter
2019-03-03T12:38:22Z batter: mixed
2019-03-03T12:38:27Z batter: poured in pan
2019-03-03T12:38:28Z frying_pan: tilted around
2019-03-03T12:39:09Z frying_pan: FLIPPED
2019-03-03T12:39:41Z frying_pan: FLIPPED
2019-03-03T12:39:46Z frying_pan: pancake removed
If you want to see the first entry for each subsystem: awk '{ if (!($2 in seen)) print $0; seen[$2] = 1; }'
Or the last (although this won't preserve input order): awk '{ seen[$2] = $0; } END { for (k in seen) print seen[k]; }'
I don't think there's another simple tool in the unix toolkit that lets you do things like this. You could probably do it with sed, but it would involve some nightmarish abuse of the hold space as a database. awk '{ if (!($2 in seen)) print $0; seen[$2] = 1; }'
You can even shorten this a bit! "awk '!seen[$2]++'" does the same thing -- awk will print the whole line when it's provided a truthy value. It's definitely more code-golfy than being explicit about what's actually going on thoughhttps://www.ibm.com/developerworks/library/l-awk1/index.html
Although I admit that the key argument for Unix tools is that they don’t get updated. That sounds awful, but think about it, once it works, it works everywhere, no matters OS type, version or packages installed. That is something experienced programmers always want from their solutions.
import sys
for line in sys.stdin:
fields = line.split()
# now you can do your logic
If you want to use regular expressions, that's another import.Python also doesn't play well with others in a pipeline. You can use python -c, but you can't use newlines inside the argument (AFAICT), so you're very limited in what you can do.
I still wouldn't touch Perl with a bargepole, though. Sorry not sorry.
https://github.com/thisredone/rb
Probably not as fast as many of the individual unix tools it could replace, but does look like a great way to leverage one's knowledge of ruby.
the key argument for *nix tools is that they do one thing and only one thing extremely well. at a meta level these tools are units of functionality and you’re actually doing functional programming, on the command line, without realizing it.
Kind of simpler to learn one language that does lots of things well!
each tool you listed (with the exception of awk) does one thing and does it extremely well.
my goal is not to use a language to solve all possible variations on problems i have. my goal is to solve the problem.
another interesting side effect is that a lot of times this is super compact and good enough. when it’s not you can go to a programming language
sed is for editing streams
Perl can, since it borrowed a fair amount of awk. It's also almost as commonly already installed. The one liner equivalents to what you showed are pretty similar, for example: https://news.ycombinator.com/item?id=19294575
Though, I concede it falls outside the realm of "simple tool".
> However, since joining this discussion, a lot of Unix supportershave sent me examples of stuff to “prove” how powerful Unix is.These examples have certainly been enough to refresh my memory:they all do something trivial or useless, and they all do so in a veryarcane manner.
Do you alias these awk commands on all machines you work on, or other way put, I did not find a nice way to keep my custom aliases 'in sync' over different machines, perhaps you have some recommendation or workflow that is really sweet?
TIA!
I have a script on my path called huniq ('hash uniq') that contains that awk program. I prefer scripts to aliases because they play better with xargs and so on.
I have a Mercurial repository full of little scripts like this, and other handy things, which lives on Bitbucket, and which i clone on machines i do a lot of work on. In principle, whenever i make changes to it i should commit and push them, then pull them down on other machines, but i'm pretty slack about it. It still helps more than not having anything, though.
> awk '{ if (!($0 in seen)) print $0; seen[$0] = 1; }'
fits right in.
perl -ne 'print unless $SEEN{$_}++'
Then I decided to run it on larger dataset (because I needed too). Like week of logs, not a day of logs.
While it was running, I wrote rust CLI, which was working like `cat /*.log | logparser` and did one day in 12 seconds, and a week in a two minutes.
And I gave up waiting on awk, btw. It is not always better to use command line. If you have gigabytes or tens of gigabytes of data, it would be easier to write some cli tool to help you out.
Also, it was much easier to put significantly more complex logic into it because of type checking, and, you know, being actual high level programming language, not hack&slash awk script.
EDIT: Looking back on my "one liner" vs "rust cli" I would not be able to make meaningful adjustments to one liner comprehension. It is, to my sorrow, write-only thing.
AWK scripts tend to be very readable (much more so than e.g. sed) as long as they stick to the "stateful filters" use-case as https://news.ycombinator.com/item?id=19294195 calls it, but yes they have their limits.
If speed is a concern, you may want to try using mawk instead of GNU awk/gawk. I've had 4x speedups with mawk.
For example, I was parsing logs. They had entries urlencoded JSON, one document per line, each could be invalid, I had to extract 'id' field, and count number of entries with same IDs, and number of entries with same IDs and with special marker. Then take only entries with 10000+ results.
You can totally write cat/grep/awk/sort/head. But then I wanted to add a field, and it was hard to edit. Rust solved my problem, and I had pleasant time writing code, not editing foot long line of untyped code.
I tried to look it up and couldn't find, sorry. It was more than half a year ago.
* If it requires aggregation and it's small, use cli tools.
* If this is data you're using over and over again then load it in the database and then do the cleaning, ELT.
* If it's 2tb of data and under, still use bzip2, get splittable streams and pass it to gnu parallel.
* If it requires massive aggregations or windows, use spark|flink|bleam.
* If you need to repeatedly process the same giant dataset use spark|flink|bleam.
* If the data is highly structured and you mainly need aggregations and filtering on a few columns use columnar DBs.
I've been using Dlang with ldc a lot because of how fast its compile time regex is, and its built in json support. Python3+pandas is also a good choice if you don't want to use awk.
Sort is good for aggregations that fit on disk (TBs these days, I guess)
Perl does well too if the output fits in a hashtable in DRAM, so 10’s (or maybe 100’s?) of GBs
While I think it's important to make that argument, the posted article and the one it refers to lack some guidance on how to reach "command line mastery". I recently came across this great resource here on HN:
https://github.com/jlevy/the-art-of-command-line
It gives great overview of the toolbox you have on the command line. Equipped with `man` you're ready to optimize your everyday work. And always remember to write everything down and ask yourself WHY something works the way it works. The interface of the standard tools is thought out very well. Getting comfortable with this mindset pays off.
https://adamdrake.com/command-line-tools-can-be-235x-faster-...
If you want to do the parsing in Python instead of awk, just make a tiny script that reads from stdin and writes to stdout - that way you can put it between xargs or parallel and whatever else is in the pipeline.
The parallelization is a separate concern, so it doesn't need to be mixed in with the parsing (or whatever) concern. The downloading is a separate concern; use wget or requests in a Python script or whatever, it doesn't need to be mingled with the parsing.
Unix commands are great up to a few GBs of data, Excel is even better if you're dealing with less than a few tens of MBs. But to deal with Terabytes of data quickly and efficiently, these tools totally break down.
It's an important point to remember that a lot of things involved in human society have not exploded in size or complexity in the last 30 years.
Many data sets are basically proportional to the human population (health records, criminal records, property records etc), and these have been measured in the millions for 30+ years. In the same time the compute power of a single script has moved from the millions into the billions.
It's important, because if a government needs to, say, calculate something involving "every building in the country", or "everybody with a criminal record" they need to understand that this task, in 2019, can in fact be done by a single programmer parsing flat text files on their MBP, and does not need a new department.
This is a bit like Grace Hopper always pointing out the difference between a microsecond and a nanosecond - https://www.youtube.com/watch?v=JEpsKnWZrJ8
Awesome point well put.
Scaling up to 400~500GB of logs with awk and parallel has not been a problem for me. I dont think TB will be particularly hard, especially if one is reasonably proficient with the tools. Of course, if one has the mental bias that one has to throw hadoop or spark at it, thats a significant obstacle right there. Upton Sinclair effect also plays a role -- It is difficult to get a man to understand something, when his salary depends upon his not understanding it.
Of course at some scale simple unix tools become impractical, but usually people reach for flashier tools even at scales where unix tools will suffice.
I doubt this is the case. If you're working on even a medium sized team, you'll see this level of data, whether it's internal server logs, publicly-sourced data, or a variety of other applications. Almost by definition, any company that runs a cluster is probably dealing with TBs of data (otherwise they could probably run it on one machine). I don't know the percentage of developers that deal with this, but I do know that it's pretty common.
This doesn't follow:
1) One reason to run a cluster is to use flaky commodity hardware instead of high-reliability specialized hardware. Note the transition from specialized hardware and computer rooms of the '80s and '90s to cloud computing on preemptible instances.
2) They might just be straight-up wrong / misguided. http://www.frankmcsherry.org/graph/scalability/cost/2015/01/... As a professional developer I've seen tons of distributed systems that could be replaced by a well-designed non-distributed system, but ELK and Mongo and etcd and Kafka and more generally cloud servers are easy off-the-shelf tools.
You can get a lot of milage out of parallel xargs and other of the shelf tools.
If you're literally processing TB or PB of data, you'll want to parallelise. Though shell tools do this amazingly well.
80% of bioinformatics.
This is a frustrating viewpoint, and it feels to me like the same poisonous worldview as "If your company isn't growing 10% month-over-month, it's not worthwhile and we won't invest in it." There are plenty of meaningful and useful things you can do for the world at a static size, and doing continued interesting work with a dataset of the same size doesn't mean you're not part of the "real world."
Or they don’t. Not all data are cloud-scale aggregations.
Also most people don't deal with terabyte sized data sets
Isn't it logical to have historical summary info and then the full log for say the past month?
What if you want to change what's contained in your summary?
https://www.gnu.org/software/coreutils/manual/html_node/inde...
I suggest changing the link to: http://widgetsandshit.com/teddziuba/2010/10/taco-bell-progra...
That’s cringe-worthy...
No flexing about how you use X needed.
Anyway, it's poor form to put gender into the mix of what good coding looks like. Don't man up, do it hard core or bravely. Don't pussy out, chicken out. Don't try something ballsy, try something gutsy.
I've been freelancing for a long time but never automated invoicing people up until recently.
So I combined grep, cut, paste and bc to parse a work log file to get how many hours I worked on that project, what amount I am owed and how many days I worked. I can run these analytics by just passing in the log file, a YYYY/MM date (this month's numbers), YYYY date (yearly numbers) or no date (lifetime).
Long story short, the working prototype of it was 4 lines of Bash and took about 10 minutes to make.
Now I never have to manually go through these work log files again and add up invoice amounts (which I always counted up manually 3 times in a row to avoid mistakes). If you're sending a bunch of invoices a month, this actually took kind of a long time and was always error prone.
We spent about a week writing a program in QuickBASIC that successfully parsed the files and performed the transformation.
Some years later, I realized that this would be a one-line awk script which I could now write in 20-30 seconds. (Probably someone comfortable with Excel could also perform the transformation in 20-30 seconds, although it might not scale as well to larger files.)
Bottom line: when facing an engineering problem, start with the simplest, fastest to implement solution, and build complexity as necessary. The simple solution suffices most of the time.
A good GPU can do >1bn calculations per frame.
You need to code more defensively with them. For example, it is rare, but every so often a newline will be fail to be emitted.
kinda\n
likethis\n
\n
example\n
There are many other gotchas, but that one is a doozy because if you're using, say, tab delimited data and cut you'll miss a line. It's one of the reasons I use line delimited JSON if at all possible.Also, this constant re-parsing of text does mean your string validation needs to be more paranoid. For example, some JSON parsers parse curly quotes as normal programming quotes. Horrible practice, I know, but it could have been avoided. Also, it's easy to accidentally do shit like this when you're in a rush. Some string matching tools handle the character matching of the different ways of creating, say, "ë" will also make matching quotes more relaxed.
Anyway, all of this to say that I 100% agree with the posted and linked articles, but each method has its own security considerations and software folk should be aware of them before starting.
We should be composing tools with multiple typed io stream paths in GUIs (or TUIs I suppose), leveraging two or even three dimensional layouts. All our interfaces should be composed this way, allowing us to take them apart and modify them at will to fit our workflow.
But that never happened. We never made a better hammer, we just try to squeeze all our problems into ASCII-processing nails instead.
It's to the point where I've been toying around with creating my own shell and faking typed IO streams via Postgres+DSL. It's tricky though. Sometimes I want pub-sub, other times I want event stream. Sometimes I want crash-on-failure, other times I don't. There is this problem in software that I can't really word precisely, but the closest I can come is "do it like this, except these cases here, except-except those cases there" and these things kinda keep stacking up until you have a program that has too much knowledge baked into it.
Take, for example, emoji TLDs. Because emojis aren't consistent across platforms they can get coerced into different types. I didn't know that when I bought and used a couple emoji domains. When someone tried to click on a link in Android and was met with a 404, I was so confused. I wasn't even seeing the request come into nginx!
After I figured it out, I realized that emoji domains won't work. The underlying assumption of TLDs is that there is one, and only one, way of encoding something and that these things aren't coerced. That assumption is wrong.
Try multiple interpreters with timings.
Gawk's profiler can be invaluable.
I think this statement is wrong. The popular meaning of the hype term “big data” can not be easily changed.
Rather, awk, sed and other tools that can read from stdin and write to stdout are great tools for “big data” and often more efficient and suitable than larger and more hyped systems.
These days, these are called big data. No it isn't...
I know bash, and know a lot of basic commands, but I'm not familiar with some more advanced things. I don't know awk or sed for example.
So I end up writing a lot of bash scripts. A lot of small python scripts. Stupid scripts. Over time, you build up a library of tricks. You read threads like this where you pick up new tricks. It isn't something you'll learn once. Some of the tools take years to really get the feel of.
if your data set can be disposed by an
awk script, it should not be called “big data”.
Why not? I don't see how awk is limited to a certain amount of data.The funny thing about "big data" in my experience, is just how small it actually becomes when you start using the right tools. And yet so much energy goes into just getting the wrong tools to do more...
Rings way too true for me atm.
My current workplace is currently struggling, because one of our application stores something like a combined 300G of analytics data in the database with the application data. Modifying the table causes hours of downtime because everyone claims that backwards compatible db changes are too hard. And everyone is scared because with more users there's "so much more analytics data" incoming. Yes, with 300Gb across 3-4 years.
And I'm just wondering why it's not an option to just move all of that into one decently sized mysql/postgres instance. Give it SSDs, 30 - 60Gb of ram for the hot dataset (1-2 month) and it'll just solve our problems. But apparently, "that's too hard to do and takes too much time" without further reason.
Add in images and the rest of the Web payload (800 KiB per page), and that swells to Petabyte range. But the actual scale of the user-entered text is stunningly small.
https://twitter.com/devops_borat/status/288698056470315008?l...
That’s not really ‘big data’ in my opinion.
I've been programming for 40 years and using unix/linux since the 80's and in this little one-line script, I discovered two things that one can do with the appropriate arguments that I've never known. YMMV.
There are 8.1 million communities in total, and thanks to some friendly assistance, I'd identified slihtly more than 100,000 with both 100 or more members, and visible activity within the preceeding 31 days, as of early 2019.
The task of Web scraping those 100k communities, parsing HTML to a set of characteristics of interest, and reducing that to a delimited dataset of about 16 MB, was all done via shell tools, and on very modest equipment.
Most surprising was that parsing the HTML (using the HTML-XML utilities: http://www.w3.org/Tools/HTML-XML-utils/README) took longer than downloading the data.
Creating the datafile was done with gawk, and most analysis subsequently in R, though quick-and-dirty summaries and queries can be run in gawk.
Performance: downloading (curl): 16 hours, parsing (hxextract & hxselect) 48 hours, dataset preparation (gawk): 2 minutes, analysis (gawk / R), a few seconds for simple outputs.
The parsing step is painfully long, the rest quite tractable.
I'm planning on posting the data, probably to https://social.antefriguserat.de/ and will include procssing scripts.
This is the fetch-script, which saves both the HTML and HEAD responses:
#!/bin/bash
sample_file=$1
comm_path='community-pages'
base_url='https://plus.google.com/communities'
i=0
time sed -e 's,^.*/,,' $sample_file |
while read commid;
do
i=$((i+1))
echo -e "\n>>> $i $commid <<<" 1>&2;
url="${base_url}/${commid}"
commfile="${comm_path}/${commid}.html"
commhead="${comm_path}/${commid}.head"
echo "curl -s -o '${commfile}' -D '${commhead}' '${url}'"
done
The sample file is simply a list of G+ community IDs or URLs, e.g.: 100000056330101053659
100000310247038604843
100000355641542704509
100000408644688836681
100000537266485621548
100000813948204546252
100001055751908082772
100001158162744298957
100001173291703462139
100001193552641351693That's why the script echoed the curl commant rather than run it directly. It fed xargs.
The other problem was failed or errored (non 3xx/4xx, or incomplete HTML -- no "</html>" tag found) responses. There was no runtime detection of these. Instead, I checked for those on completion of the first run and re-pulled those in a few minutes, a few thousand from the whole run, most of which ended up being 4xx/3xx ultimately.
HA! This has almost been my line for years regarding Mexican food. What I like to say is: it’s amazing how every possible permutation of 8 ingredients has been named. BTW I love Mexican food, lived in Mexico.
> The post mentions a scenario which you may consider to use Hadoop to solve but actually xargs may be a simpler and better choice.
I do feel like there’s a corollary to Knuth’s “premature optimization” quote regarding web scaling; premature scaling and using tools much bigger than necessary for the job at hand is pretty common.
I run 'motion' on my linux desktop at home to serve as a security camera when no one is home. For months I've been manually starting and stopping it, figuring I needed to setup an IoT system if I wanted to automate things. i.e. IFTTT on our phones, an MQTT server in the cloud, etc. Then I realized - I just need to start the camera when all of our phones are off the LAN. It took about 15 minutes to setup, and now I never have to worry about forgetting to stop or start the camera.
Good enough programming is something you code that provides something people want -- and you never look at the code again. Find the problem, solve the problem, walk away from the problem. That's not sexy. It's not going to get you an article to write for a famous magazine, but it's good enough.
We have lost sight of "good enough" in programming, and without some kind of guardrails, we end up doing stuff we like or stuff that sounds good to other programmers. For instance, while I love cloud computing, I'm seeing "how-to" articles written about setting up a VPC for doing something like playing checkers. Yes, it was an oversimplified article, and you have to write that way, but without wisdom, how is the reader supposed to know that? What criteria do they use to determine whether it's a co-lo server, a lambda, or a world-wide distributed cloud?
We're going like gangbusters selling programmers and companies on all kinds of new and complex ways of doing things. They like it. We like it. But it is in anybody's best interest over the long run?
Recently I rewrote a pet project for the third time. First time it was C#, SQL Server, and an ORM. Then it was F#, MySQL, and linux. The last time it was pure FP in F# and microservices.
Some of you may know where this is going.
Just as I finished writing the app in a real microservices format, I realized. Holy cow! This whole thing was just a few Unix commands and some pipes.
My thinking went from all kinds of concerns about transactions and ORM-fun to just some nix stuff in a small script. The problem stayed the same.
Something else happened too. At each step, I did less and less maintenance. The last rewrite has had no maintenance required at all. In my spare time, I'm going to do the nix one using no servers at all on a static SPA. In a very meaningful way, there's no app, there's no server, and there's nothing to maintain. Yet I still get the functionality I need. And I never maintain it.
Of course that's not possible for every app, but the key thing I learned wasn't the magic of serverless static SPAs or the joys of unix. It was that I didn't know whether or not it was possible or not until I did it. By thinking in a pure FP fashion and deploying in true microservices, the rest just "fell out" of the work. At first I was actually thinking in a way that would have only led to more and more complexity and maintenance requirements.
My belief is that we get our thinking right first, use code budgets, and try for a simple unix solution. If it doesn't work, why? At least then we've made an effort to be good enough programmers. That beats most everybody else.
Edit: Can I can sign up somewhere to get a heads up when your book is available? Would be appreciated!
EDIT: Lex/Yacc that is some faster parser generator, I'm not too knowledgeable on that.
edit: same thing can be done with Python/Perl
Do you seriously think non-whitespace separated structured records is a novel idea which the simpler times of Unix didn't have to deal with? Have you looked at the passwd file? /rant
The fact that there are some standard tools available doesn't mean you are limited to that.
If you have a CSV file with spaces and newlines use cvskit or a small python script importing the relevant library. If you have to parse JSON file you can use jq to pick the relevant fields regardless of how the document is formatted. You can even process binary data as long as the file format is understood by the tool.
The choices are limited and when you go there you know there's a good chance that you will end up wondering why what you got caused so much pain.
Pragmatism almost always loses to CV-padding and office politics.
The post makes a good point that I fully agree with, just doesn't explain it well enough.
For example the initial implementation of AWK was in 1977 [1], a few years before GNU even existed [2], so it _is_ a Unix tool.
[1] https://en.wikipedia.org/wiki/AWK#History [2] https://en.wikipedia.org/wiki/GNU#History
So while GNU is obviously an important project, it would not be correct to say "xargs" is not a Unix tool.
GNU awk is one of the popular awk implementation (usually referred as gawk). I personnally prefer mawk. awk is not GNU.