Perl first commit: a “replacement” for Awk and sed
github.com
github.com
Perl in one stroke collapsed the programming of C, text manipulation, the capabilities of all of the Unix utilities, and data structures into one system. For anything which wasn't subsumed into the monolith of Perl, you could easy access via backticks. It was very friendly in dealing with text streams, and that's what those call-outs in those back ticks spoke.
Yes, awk and sed were replaced by Perl, but more importantly, the unmaintainable nightmare that glued all of it together was wiped out.
Adding to that some popular examples are here [1][2]
[1] - https://www.commandlinefu.com/commands/matching/awk/YXdr/sor...
[2] - https://www.commandlinefu.com/commands/matching/sed/c2Vk/sor...
I still use awk and sed semi regularly. I haven’t used Perl in over a decade.
They've become much more useful because of the influence of things like perl.
To see what I mean, here's SunOS 4.1.1's man page for awk and sed as grabbed from https://github.com/ambiamber/Run-Sun3-SunOS-4.1.1/blob/main/... and ran through groffer(1) . These are from 1989 and 1987 respectively. The core stuff is there but not much else.
I was incorrect about 30 years. It's actually 35.
You were the one that that said POSIX awk to begin with, I was using your terms. I understand what you meant, why don't you?
As far as shitting on the GMU tools, I don't think I've seen someone do that in decades, especially just referencing the bottom of the man page like that.
As far as speed, I don't know how you're using these tools, but if gawk is your bottleneck and not i/o, there's probably smarter ways to do things.
This is not a productive conversation. You can live life however you want and if you're happy than I'm happy. We are definitely at an impasse. I'm off to bed.
POSIX was initially a bunch of the Unix vendors getting together in the 80s seeing a rise of incompatibility saying "this shit is crazy let's find something we can agree on" and then the conflicts of trying to still have a differentiating But Also compatible product. They knew that if they cannibalized themselves, IBM Novell and DEC were there to snatch up the customers
As a result, POSIX was intentionally kneecapped by all the players but not enough to make it completely worthless.
By the 2010s I decided to simply ignore it and if something breaks on FreeBSD or whatever, make the decision then whether to support it or not. Mostly it's platform testing at the top and then subbing the gnu version of things and bailing out if it's not there.
Same. Awk and Sed are delightful little tools that have aged exceptionally well.
Other than that, I'm completely shocked that somebody would praise sed or awk in well into the 21st century. Turing-complete languages barely capable of properly solving any problem that has any sort of algorithmic complexity, let alone one that requires proper data types, variables or interacting with any sort of system. Really, you can do that in Perl (even in Raku) without your couple-of-liner horribly breaking down once you need some custom logic to it. Oh, and it will even resemble programming languages, not some hieroglyphs left behind by aliens.
I'm not in this space, but the similarity in names made me wonder if these would really have started around the same time, with an unclear "dependency order", so I reached for a search engine, and according to Wikipedia, CPAN (1993) is based on CTAN (1992).
There isn't "some question", there is a definite answer, namely in "Programming Perl" and I quote the 4th edition, page 629:
History
Toward the end of 1993, Tim Bunce, Jarkko Hietaniemi, and Andreas König set up the perl-packrats mailing list to discuss the idea of an archive for all the Perl 4 stuff floating around the Internet. Perl 5 development had started that year, and one of its main features would be an extensible module system that would allow people to extend the language without changing perl. Jared Rhine suggested he idea of a central repository, but nothing much happened. His idea had come from CTAN (http://www.ctan.org), the Comprehensive TeX Archive Network.
Edit: fix markup
I may not completely understand what you're describing, but with perlbrew managing different Perl versions, and various CPAN clients installing modules in versioned libraries - as well as the test/smoke servers that test CPAN on tons of versions of Perl for Perl CPAN authors automatically, I don't know if I've come across what you're describing.
What's still a pita is when there's an outside dependency from the Perl ecosystem that breaks a module, but I'm not sure that's Perl's fault.
This doesn't sound that horrific to me. It's the classic Unix approach of building small tools that do one thing well, and composing them in novel ways to solve problems. For any problem that can't be solved this way you write another small tool using your programming language of choice. Rinse and repeat.
But occasionally Unix attracts users and programmers who reject this approach, and who prefer building a monolithic tool, or in the case of Larry Wall, new programming languages. To be clear, I'm a fan of Perl and think it has its place, especially in the era it came out. It inspired many modern languages, and its impact is undeniable.
Personally, I find solutions you refer to as "unmaintainable nightmare" to be simple and elegant, if used correctly. No, you probably shouldn't abuse shell scripts to build a complex system, and beyond a certain level of complexity, a programming language is the better tool. But for most simple processing pipelines, Unix tools are perfectly capable and can be used to build maintainable solutions.
The classic Knuth-Mcllroy bout[1] comes to mind. Would you rather maintain Knuth's solution or Mcllroy's?
I have, and mentioned one lower down in the comments. Unix philosophy was great but does not scale well in terms of maintainability or efficiency. Invoking processes over and over again loops is godawful slow. And the horror of complicated shell scripts is legendary.
But this doesn't make this approach inherently wrong or obsolete. The programmer is wrong for trying to use the tools beyond their capabilities. Where that line is drawn is subjective, as is the concept of maintainability, but if you feel that you're struggling to accomplish something, and that it's becoming a chore to maintain, the path forward is choosing a more capable tool, like a programming language.
Perl was invented because the gap from shell to more capable languages was (and is) really big. Languages like Python and Ruby didn’t exist yet, and Perl had a really, really strong sweet spot in text processing.
Still does.
Ubiquity, speed, and conciseness.
Perl is usually installed by default on Linux and Unix systems. Ruby might be there, it depends.
Perl is faster than Ruby. Ruby has been one of the slower scripting languages. But Ruby has been working on performance improvements in the past few releases. I have not seen any benchmarks of the current Perl versus the current Ruby, so this may have changed.
Perl is more concise than Ruby allowing more functionality for less code.
The JVM would like a word. (It has slow startup, but can be very fast at runtime due to JIT optimization and cacheing.)
Scripts should start and run quickly.
Ruby has historically been fairly slow which is why Ruby 3 focused heavily on performance. It has been improving a lot, but I have not seen any benchmarks against other programming languages.
This results in weird behavior, such as writing a groovy (Java?) script for Jenkins to execute bazel in order to build a go binary that runs the very same commands in an exec.Command() construct. Or people who download and import pandas to grab the third field in a csv file.
During the course of learning, I've naturally written code in bash that should have been written in another language. I replaced if statements with case because they turned out to be more performant. It's a great learning experience and why I got into python and go.
IMO we should use the right tool for the job. Sometimes that tool is a combination of unix utilities that you can put in a shell script for easier maintenance. It's just procedural execution of (usually very efficient) binaries, akin to a jenkins script or gitlab pipeline. Just mind the exceptions and use exit codes.
* often times, it's not just the third column I want. Sometimes it becomes "third column unless the first column is 'b' then instead grab the fourth column". Having a good data representation makes sure that I'm not mixing logic code with representation parsing code
* I don't have to care about CSV parsing edge cases. Escaped comma? Quotes? I don't care, the library will either handle it or throw an explicit error. With custom parsing code, instead of an error, I'll get some mangled result in the middle of the file that I won't even catch / notice until later down the line
* when working with CSVs, in my area (ML / scientific compute), Python is often the right context to be in.
It's not a lack of being "taught" shell scripts. It's the fact that shell programming constructs aren't well documented, your "standard library" is basically dependent on whatever binaries happen to be available on the filesystem, error handling is almost non-existent, etc.
It's very easy to write a bad shell script that "solves" a problem as long as a bunch of assumptions aren't violated. In my experience, senior software engineers are extremely averse to hidden assumptions and very concerned with reliability of the systems they build.
Of course this can happen with any language, especially as it ages and adds complexity.
Everything was needlessly hard because these tools were not built for that. Easy to talk about philosophy and the "classic Unix approach" if you don't have to build modern applications this way.
You hit it on the head with the slowness of loops when the body comprises a series of program invocations. The horror really seeps in when you realize the original author wasn't stopped by the lack of data structures: they could get around that with some creative variable names.
Other backronyms too.
Whippersnappers! :D
The first big iron I had the luck to work with was an IBM 3090 , essentially a gift from IBM, it handled the university entrance exams of the entire country of some ten million people and it had 64 MB of RAM. (It was also the first computer in Hungary permanently connected to the Internet via a leased line to Austria so it had an Austrian IP address. Hungary didn't have its IP region for two more years.)
I think the first machine with 128MB was a VAX 6510 a year or two later at another university. A little bit later, in 1994, CERN had gifted a VAX 9000 with an astounding 256MB of RAM.
To compare, the first server I installed Linux on had a grand total of 4MB RAM -- and that was one of the largest computers a small department at the university had.
It would be a long, long time before "128MB" and "mine" entered the same sentence.
4MB?!?
My first encounter with IBM kit was a, er, darn I'm not sure cuz I'm getting old, but I think it was a 4300? Not big iron in some senses, but still with a box that was something like 6-8 feet long iirc and definitely several feet wide and high. (And a bank of about 6-8 tape decks, each as tall as me, and two disk units, each the size of a washing machine, and so on.)
Its RAM? A massive 1 MB.
That IBM kit was the heart of the super new expensive upgrade in 1980 that cost something like 5-10 million pounds iirc to build, including a brand new building to house it and a team of programmers.
The older setup, which is where I was until its last days, was an ICL system that was expanded at the end of its life to a whopping 48KB -- yes, KB -- of RAM.
And that kit ran all the systems, internal (payroll, accounting, etc., etc.) and external (sales etc.) for the largest car dealership in the UK.
128MB? 4MB? Even 1MB? That was an unimaginably insanely large amount of RAM!
(Yes, it was very weird to be working with this physically enormous setup, and dealing with keeping it all cool enough not to halt for a half hour or so, through super human efforts when the A/C broke down, when the likes of PETs, Sinclair Z80s, and Acorn Atoms were a thing...)
Ha yes the aforementioned IBM 3090 was so big for installation they removed the roof of the building it was living in, craned it in place and put the roof the back. Bringing it up the elevator or stairs was impossible.
Much later, in the second half of the 90s, I remember the four of us carrying an IBM HDD -- I think it was your normal 5.25" drive but it needed four people because it was mounted on a vibration dampening base ...
Continuing the shades of Monty Python theme[1]:
I remember one of my first few nights being in charge of the new IBM kit (I was a "computer operator" back then, in 1980), leaning back in the fancy new chair at the desk with its fancy "virtual" teletypes (a couple "terminals" displaying the status of the OS with a CICS system), and showing off to an "underling" by swinging a long plastic slide rule or something stupid like that (I no longer recall), and me accidentally banging it on the desk. Right "near" a recessed big red button. Or perhaps "on" the button? As I snapped my head around to look at the button and begin to understand what I may have just done I heard an ominous series of whirring and clicking sounds coming from the cpu box, right near where there was an 8" diskette drive that wasn't supposed to be doing anything while the OS was running (it was just for starting the OS). Then I looked at the console... Uhoh. They didn't fire me but it took months before they decided to let me be "in charge" again with someone else actually hovering over me...
Fast forward to when I was a coder (BCPL) in a small software startup, during the second half of the 80s, presumably 10 years before you were carrying your 5.25" drive monster, I vividly recall someone bringing a 700MB hard drive back from a local computer store. It cost an astonishingly paltry 700 quid or thereabouts. A pound a MB!
I ... do not know. That sounds very low. Look at https://jcmit.net/diskprice.htm and note the pound was 1.8-ish around this time so the price should have been well above 1000 pounds even in early 90s. We are talking of a 5.25" full height drive, here's an 1987 model http://www.bitsavers.org/pdf/maxtor/MXT8760E.pdf rare in personal computers, it was definitely for workstations / servers.
That's what the O'Reilly books were for, especially the Nutshell series.
The one thing people can't possibly fathom if they started coding after the mid-late 90s was how much we relied on the printed medium.
Rather Waite-y.
By the time our lord and savior zsh appeared on the scene Perl was already at Perl 3. And, to be fair, I do not think many used zsh before 2.1 which was some time 1991 fall and by then the Camel Book was out for half a year or something like that. So the pre-Perl and the we-use-zsh days do not really overlap.
When Perl appeared the absolute hotness was the version of Korn shell what later became known as ksh88. https://github.com/weiss/original-bsd/tree/master/local/tool...
but the only free programming languages available at the time were C/C++, various shells, and awk. everything else was expensive or not generally usable for other reasons. all the really useful languages to build complex systems didn't really appear or become freely available until the 90s. and perl was first among those.
But the thing is that today the shell landscape is much more mature for solving simple problems, and we have C/C++ alternatives that are saner and more capable than Perl (e.g. Go). So it arguably has lost its place, as shell tools are still in widespread use, while Perl is mostly underused. Raku is interesting, but it goes in a different direction, and its adoption is practically zero.
This is by design. Readability is core to the design and philosophy of python. One liners are cool and fun to write, but trying to decipher someone else's incredibly dense bash or perl one-liner is absolutely awful.
You can write hard-to-read code in any programming language.
Python lets you with mandatory whitespace so that the awfulness spans multiple lines instead.
Really talented Python programmers can do downright demonic stuff with list comprehensions.
Python appears to be simple, but is actually quite complex. I recommend reading "Effective Python" (https://effectivepython.com/) to see beneath the surface.
Just think of C: I'd argue its design is actually more akin to Python than Perl (and it definitely inspired languages like Go and Zig, NOT languages like C++). It's a small language and this is a very important characteristic of it: you can count on being able to actually master it or at least very well comprehend it. Other effects of a simple and literally straightforward language can be: easier implementation and evolvement, less mental load on the developer, easier portability among developers (both for general knowledge and actual code), etc. I'm not saying that C is the way it is for all these reasons but I wouldn't overlook this factor and I do think that languages like Python are deliberately building on these advantages.
Now, I don't dislike C++ at all but back when I studied it at university, I noticed that it was the first language for me that needed to be actually studied, unlike Pascal, C, Python and "oldschool" JS. Ever since, the only languages where I felt the same were Prolog (mostly because it requires a different mindset; other than that, it didn't seem bloated) and Raku. Not C#, not Java, not Erlang. I didn't really have to touch Perl but from all I know, Raku started off as a fresh take on the Perl approach. It seems somewhat more organized but huge nevertheless, to the extent that there literally isn't one person who really "groks the language" all around. In the case of Raku, I wouldn't even say it encourages you to write unreadable code (especially if you have a thing for APL look-alikes, lol) - it's just so rich that there is a good chance you will come across something in someone else's code you have never used before and don't quite remember how it will act in your specific use case.
There are different types of freedom than "do whatever you want". Like, the freedom to feel safe and confident about code. These days, humanity has aggregated an immense amount of knowledge and technology and "I will do it all by myself" is not that much of an option. And even people with such puritanistic tendencies will choose simple and straightforward tools, even if not for the "limitations".
Now I have barely spent any time with the oldschool Perl but trust me, I have put a lot of effort into learning Raku, the language that was meant to fix Perl. Whenever something that "seemed like a good idea at first but it's actually harmful" shows up, it's usually Perl's legacy. I'm thinking of things like the conceptual mishmash between a single-element list and a scalar value (or in general, trying hard to break down variables arbitrarily into list-alikes, hash-alikes and the rest of the world), the concept of values that try to implicitly pretend they are strings and numbers at will, or the transparency of all subroutines to loop control statements which is some next level spaghetti design. If you ever actually use something like this, you introduce a brand new level of complexity, somewhere inbetween a "goto" and a "comefrom", so I would really think about if this was worth learning at all.
Oh right... from what I remember, it was also Perl that fostered this idiotic idea that a name of a concrete thing could be overloaded to be a namespace as well, and a concrete Foo::Bar could very well be something that has no logical relation to a concrete Foo. Moreover, I'm quite sure Perl invented this nonsensical distribution-module dichotomy where you are supposed to depend on modules, despite the smallest publishable and installable unit being a distribution. There are three outcomes with that: - if the distribution contains only one module: what was the point of drawing the distinction? - if the distribution contains tightly coupled modules: you can pretend to only depend on one of the modules but in fact you are depending on the whole distribution together - if the distribution is a collection of unrelated modules: why are you trying to encouple the metadata when this will make the versioning meaningless?
I can only hope that it's somehow better than Raku but the whole principle is just an anomaly.
And you know, then these people move around in the world, pretending that all of this is just normal and you just have to learn it. Well guess what, there is a reason people might want to put that effort into something else.
I think it's possible that things that seem normal and inoffensive can become horrific simply from scale. You'll climb the stepladder without complaint, but then there's that radio tower in Canada...
Others like the idea of the family cow, then they see the 10,000 head feedlot from the highway.
Scale is sometimes sufficient by itself to induce horror.
If you can depend on a recent bash and use shellcheck, then it's actually quite a pleasant programming environment, with fewer footguns than one might think. (I want a @#$@# "set -e" equivalent that returns non-zero from a function if any statement in the function results in non-zero).
There are some things that are more awkward than they should be though (e.g. given a glob, does it match 0, 1, or many files, or the way array expansions work).
Also, there's no builtin way to manage libraries (I don't know about Perl, but Python suffers from this as well). This results in me pasting a few dozen lines of shell at the top of any of my significant shell scripts, for quality-of-life functions. Then I have to use "command -v" to check if the various external programs I'm going to use are present. Say what you will about C, but a statically-linked C program can be dropped in anywhere.
As for managing libraries, that's true, but you can certainly import and reuse some common util functions.
For example, this is at the top of most of my scripts:
set -eEuxo pipefail
_scriptdir="$(dirname "$(readlink -f "${BASH_SOURCE[0]}")")"
source "${_scriptdir}/lib.sh"
This loads `lib.sh` from a common directory where my shell scripts live, which has some logging and error handling functions, so it cuts down on repetition just like a programming language would.Moreover, it's existence being required explains why my error handling, recently, wasn't working as expected.
This works really well if your problem can be solved in one or two liners.
It go bad very quickly when, say, you have two CSV files and want to join them the sql-way. In sed, you have to use positional variables and think about shell escaping. In perl, you can at least name those variables and use \Q
My personal comfort threshold is around the 100-line mark. It's even possible to write maintainable shell scripts up to 500 lines, but it mostly depends on the problem you're trying to solve, and the discipline of the programmer to follow best practices (use sane defaults, ShellCheck, etc.).
> It go bad very quickly when, say, you have two CSV files and want to join them the sql-way.
In that case we're talking about structured data, and, yeah, Perl or Python would be easier to work with. That said, depending on the complexity of the CSV, you can still go a long way with plain Bash with IFS/read(1) or tr(1) to split CSV columns. This wouldn't be very robust, but there are tools that handle CSV specifically[1], which can be composed in a shell script just fine.
So it's always a balancing act of being productive quickly with a shell script, or reaching out for a programming language once the tools aren't a good fit, or maintenance becomes an issue.
So I don't disagree that it was needed back then, but it's important to mention the modern context it struggles to exist in.
I agree that there’s nothing wrong with composing UNIX tools. I mean, that was one of its key selling points. if you watch any early promo videos for UNIX you’ll see them talk heavily about the composability of the command line and shell scripting. It wasn’t an accident — it was designed that way.
The point of the conversation wasn’t to say that one shouldn’t write shell scripts, it was just to say that there was a massive and unfilled gulf between what was easy to do in Ksh, awk and sed, and what could be done in C.
Then just put them in a database and write a simple SQL query. If you use Perl it’s really very simple to do.
That sounds like a perfect use case for `join`.
That said, imagine a metrics system for a huge networking company that used these methods to cover all automated testing or defect analysis. Those inner loops were made of greps and seds and so forth, and each one is the invocation of a new program. It wasn't uncommon for these runs to take almost a day.
Besides performance, the other nightmare was was someone described below: each script was a one-off that didn't leverage the work from others. If the author only new C shell, then you know you're going to be doing gymnastics to catch the stderr of some of those programs (you can't capture it in the same manner that Bourne variants do).
Anyway, yes, we all adore the Unix philosophy, but there are limits.
I remember different CLI tools working differently on SysV and other variants. I remember AIX and HPUX.
Perl meant one thing that just goddamned worked mostly the same way. That and vi and emacs. :-)
Er, Perl replaced an unmaintainable nightmare? Perl, the language infamous for being indistinguishable from line noise?
If you have doubts, try writing a complicated production software stack in awk. Then hand it off to a coworker.
One need not be irreplaceably good, if one is already beating the current state of the art by an order of magnitude. Perl did this.
I've easily written tens of thousands of lines of Perl, and not a single person has complained about difficulty reading or maintaining that code. Why? Because I apply all the usual best practices for code hygiene that apply to any language.
Frankly, I think most people are just repeating a meme they heard once, and the rest just get put off by a) sigils, and b) the use of implicit variables (which I tend to use only very sparingly, and mostly in quick one-liners).
But a developer that poorly names their functions or variables, fails to modularize appropriately, fails to document their code, abuses language features because they want to be excessively terse or clever, that kind of person is gonna write crappy, difficult-to-maintain code no matter what language they use.
In fact, I'd argue the advent and popularity of Python--which was pretty radical in how opinionated it was at the time--is a direct response to languages like Perl and C that were a lot more free-form and easier to abuse by poor coders.
https://packaging.python.org/en/latest/overview/
<kidding> OCI (Docker) Containers were invented to overcome the problems inherent with deploying Python applications. </kidding>
I would argue the same of sed/awk; you can write an unmaintainable mess, but you don't have to.
Developers extremely seldom complain about other developers' code to their faces.
Instead they go to lenghts to avoid it if they dislike it severely; right or wrong.
Get the Perl source code for PangZero and you'll understand proper Perl code.
PERL being open source meant you were a "make" away from having the same environment on all those unixes.
even today when the dozens of UNIX distributions have collapsed into a couple of Lunux flavours, getting a shell based script to work reliably between say macOS and redhat can take some effort.
PS is 'okay' as a scripting language, but it's very frustrating how Microsoft's security defaults seem to be dead-set on ensuring you can't use scripts on systems without jumping through a bunch of hoops. I get how MS got to this conclusion that it needs to be locked down, but it makes me think twice about pumping out a script since I know I might need to walk my colleagues through how to actually run the damn thing depending on what the Security Admins decided to lock down that week.
# Sometimes I wake up screaming. Famous figures are gathered in the nightmare,
# Steve Bourne, Larry Wall, the whole of the ANSI C committee. They're just
# standing there, waiting, but the truely terrifying thing is what they carry
# in their hands. At first sight each seems to bear the same thing, but it is
# not so for the forms in their grasp are ever so slightly different one from
# the other. Each is twisted in some grotesque way from the other to make each
# an unspeakable perversion impossible to perceive without the onset of madness.
# True insanity awaits anyone who perceives all of these horrors together.
https://github.com/openembedded/openembedded/blob/fabd8e6d07...I barely get into loops and lists when trying to learn Python before losing interest/focus and not trying again for months. It feels very abstract and like it's not useful for day-to-day stuff without learning and writing a lot. Shell stuff is very short and powerful and feels more grounded in reality. To be clear, I am not using if/else and the like in my shell scripts either. Usually they're just a few lines with little to no logic, but they're immensely useful. I can't imagine learning perl and then using it for all this stuff.
Through my use of Usenet I learned that a better reader program was available called 'rn', which I downloaded and built. It had an amazing handwritten install script (autoconf was years off) which could automatically configure and build rn on any *NIX system. All developed by a fellow named Larry Wall.
rn was truly a joy to use and made reading Usenet swift and efficient. It would get updated with the fixes that came across Usenet that I'd apply using the clever 'patch' program, also written by Larry Wall.
Based on my experience with his other software, when Larry Wall released Perl on Usenet I immediately downloaded, built, and started using it. As promised, for scripting things not requiring a C program, it was massive improvement. Version 2.0 came out and brought many great new capabilities.
I wasn't writing software while versions 3 and 4 came out; I started using it again after version 5 and the appearance of CPAN. Over the years I've used Perl extensively for task automation and data wrangling.
Python now dominates Perl's niche because it's easier to learn and interfaces better with C. It's also less flexible, which compared to Perl is a virtue. One of Perl's mottos is TMTOWTDI—there's more than one way to do it. But many of them are bad. Much of Perl's poor reputation ("line noise") stems from this.
But when Perl was released it was a revelation, like a drab day when the clouds suddenly part letting warm sunlight pour down on the land.
But it was a critical part of computing history. Perl did a tremendous job of bridging the gap between shell commands and shell scripts and “big languages”. Way back in the early 90’s I took a complicated shell script that took about 2 hours to run this huge text processing job , rewrote it in Perl and it took about 20 seconds, and was more maintainable to boot.
It really helped boost software development in the 90s, even pre-CGI script.
The Perl source itself was also very interesting, with numerous optimizations to make it fast (again, for its time).
What killed it were the never ending eccentricities you ran into, endless foot guns, and of course Perl 6 never being delivered.
But it really was the weird shit that did it in.
*technically it was 5.005 or something to 5.6 because they decided to drop a zero or two
my $x = 5 + 6;
$x .= "0";
print $x + 5; # 16, not 115.
The result in the year 2000 was a famous cup throwing incident by Jon Orwant. And Perl 6 was a plan to separate those who wanted to create their ideal language (Perl 6) from those who wanted to maintain existing Perl.And it worked, sort of, for a few years.
I was on the Perl Grant Committee at the moment that killed Perl 6 in my opinion. We had a choice of 2 grants from Nicholas Clark that we could fund during 2006. One brought a lot of immediate benefits to Perl (Unicode fixes, reduced memory usage, etc), and the other was to get Ponie to the point where someone other than a core Perl developer could work on Perl 6.
We chose the immediate benefits for Perl 5.
They did not add support for encoding/decoding html entities or URLs were added. No standard modules for this.
No easy method of making html pages from a URL request besides running as CGI.
This made perl lose to php which made it very easy to make a simple "Hello, world" page.
php would never had any traction if perl developers had made web support high priority.
For years Amazon.com used Perl Mason.
Big PHP shops like Facebook, and PHP projects like WordPress, came in a later development generation for the simple reason that PHP made it easy to get started on shared webhosting.
It's certainly not in core, but look how PHP flubbed that up (html_entity_decode, htmlspecialchars_decode, htmlentities).
>No standard modules for this.
Obviously there are,
https://metacpan.org/pod/HTML::Entities
The changelog goes back only to 1998, where it states,
2.14 1998-04-01
HTML:: modules unbundled from libwww-perl-5.22
That wasn't a hurtle to developers, but certainly to end users who did want to just throw a php script up and have it work.Netscape Navigator was released late 1994, so this was the point of time where perl developers should have seen the light.
The thing that PHP did better than Perl was the ability to be set up on shared web hosting. Mind you, it did that mostly by not acknowledging all of the security holes that it had which allowed one user access to what should have been private for another user. But shared hosting providers had a population of people who wanted to use PHP, were forgiving of major security flaws (which PHP had many of), and would pay money.
Meanwhile Perl went the route of doing things right, and being much more secure. Which, even though it feels important to good developers, is bad for marketing. It was a classic Worse is Better situation. See https://www.dreamsongs.com/RiseOfWorseIsBetter.html if you're not familiar with that concept.
And, in the end, the commodity wins. So in the mid-2000s, PHP did catch up and beat Perl.
Well, I thank you.
Is there a summary of what Tom Christensen did? I hold him up pretty high when it comes to Perl lore. I don't know from what you wrote if he's the hero or villain in this case (and if you don't remember that's fine: ancient history)
I don't know everything that Larry did about the situation. But it was a situation that everyone on p5p was painfully aware of at the time.
https://en.wikibooks.org/wiki/Raku_Programming/Perl_History#...
And yes, I deal with non-ASCII Unicode every day.
>> But it really was the weird shit that did it in.
The current state of Perl 5 for Python fans: ----
Perl 5: I'm not dead!
TIOBE: 'Ere! 'E says 'e's not dead!
Internet: Yes he is.
Perl 5: I'm not!
TIOBE: 'E isn't?
Internet: Well... he will be soon-- he's very ill...
Perl 5: I'm getting better!
Internet: No you're not, you'll be stone dead in a moment.
TIOBE: I can't take 'im off like that! It's against regulations!
Perl 5: I don't want to go off the chart....
Internet: Oh, don't be such a baby.
TIOBE: I can't take 'im off....
Perl 5: I feel fine!
Internet: Well, do us a favor...
TIOBE: I can't!
Internet: Can you hang around a couple of minutes? He won't be long...
TIOBE: No, gotta get to Reddit, they lost nine today.
Internet: Well, when's your next round?
TIOBE: Next year.
Perl 5: I think I'll go for a walk....
Internet: You're not fooling anyone, you know-- (to TIOBE) Look, isn't there something you can do...?
Perl 5: I feel happy! I feel happy!
----
But seriously, Perl is not dead yet.
It is usually installed by default on most Unix and Linux systems.
It also rides along with most MinGW-based toolkits like Git for Windows. If you do software development, there is a good chance you might have Perl installed and not even know it.
If you are on a Linux or Unix system try running 'perl --version' to see what version you have installed. I would be very surprised if it is not there.
Some people like to make fun of it despite its wide-spread use and utility. Most of them probably know little to nothing about Perl and have not written or maintained any Perl code. There is a good chance that they are using Perl code without knowing it.
The DOM API is also a bit more robust than string manipulation.
EDIT: For the curious about the history, take a look at the perlhist documentation, https://perldoc.perl.org/perlhist it's got a lot of good info about the history of perl and it's releases.
As for as I know the only "official" reference to that is the "classified" and "don't ask" in perlhist.
Edit: one plus is that perl is as ubiquitous as awk and sed. It was there the whole time and I didn’t even realize it!
Perl gives you all the simple features from sed or AWK and adds useful (maintainable) foreach iterators, arrays, hashes, etc.
Recommend using Strict mode if you are new to Perl. It gives good guardrails against silly mistakes like not declaring or misspelling a variables, or accessing strings or numbers that aren’t the correct datatype.
Customary Generic Meaning Interpolates
'' q{} Literal no
"" qq{} Literal yes
--where {} can be any bracket pair.
my $string1 = q<a single quote '>;
my $string2 = qq<a double quote ">;This was a major reason I downloaded Perl 1.0 off Usenet in 1987.[1]
A lot of the greatness of Perl came from writings of Larry Wall, Mark Jason Dominus, Randal Schwartz, Tom Christiansen. There is so much wisdom to be gained from all their code and documentation.
The question is: should you?
I like Perl, but it itself has many warts that make maintaining a large codebase more of a nightmare than using C, and certainly more than most modern languages.
I agree with other commenters here that the tools Perl sought to replace are still used more than it. It has its niche of being excellent at text processing, and more capable than shell scripts at that task, but I'd think twice about reaching for it to build anything more complex than a shell script replacement. Especially in 2023.
I write a lot of bash for better or worse. I wonder if perl would be a better choice sometimes.
I've written Perl for many years but in the end switched to Python because the internet and resources you can find are much more quantitative and qualitative than those for Perl. Especially now when Google doesn't seem to return older results anymore.
A lot of the tutorials online in my opinion are overbaked and written for people writing large OO projects using a lot of scaffolding. But the man pages were originally targeted at sed and awk users in your position.
Worth noting “sed” is “thirst” in Spanish, which has the potential to throw off the data, especially worldwide.
https://trends.google.com/trends/explore?date=all&geo=US&q=p...
All true.
But awk is also a programming language, not just a command line application or utility.
It has conditionals, loops, regexes, file handling, string handling, (limited) user-definable functions, hashes (associative arrays), reporting abilities and more.
In fact, the name of the original book about awk is The Awk Programming Language.
Great book. I read it after it was so enthusiastically endorsed here on HN. A lot of people said it was worth reading just to take in the excellent technical writing style of Brian Kernighan.
I would offer a strong second to that.
I'm still comfortable recommending people learn sed/awk but I don't think I'd recommend learning Perl* now. It sits in an awkward spot between the simplicity of GNU utils and the expressivity of a scripting language, but doesn't do either of those things better than the equivalent sed/Python etc.
FWIW I still write Bash scripts on a daily basis, too.
* that's no comment on Raku, which I consider a separate beast.
It does scripting and gluing scripts far better than python.
It's not even close.
I'm not even going to address sed, awk, bash and the like.
but otherwise a script is either just calling a lot of commands one after the other, for which a shell script is fine. or they do a lot of data mangling without needing many external applications or none even, in which case any other languages besides perl is just as fine.
This. I always feel like bash and pipes are always holding me back and resisting.
Python is just to verbose to exec/pipe/read/write.
The real issue is the it's really easy to shoot yourself in the foot with perl.
stats wise i am sure there are still a lot of perl users. it has its fans, and there is a lot of existing systems, but it is now a choice among many, and not usually the best choice.
Can you name some? I'm guessing Rexx is one of those you mean.
with the older ones i mean languages like lisp and smalltalk which didn't become available for free or even usable on PCs until the 90s
i am not saying that they are all replacing perl, but that perl was used in areas where they are better suited, and with their appearance perl is no longer needed in those areas, hence the usage space for perl shrunk a lot.
The problem with Perl is it was still tied to Unix culture to compete with Python and it’s strong library set, so it was pretty much fazed out as an awkward intermediary.
Additionally, I think it's hard to enter the "Legacy" space. You have to un-learn some patterns, lots of: oh yea, we used to do this the hard way. The other bump is the documentation - the 1999 docs are buried under the 2009 docs which are decaying under the 2019 docs. The thing called ActiveRecord for example means like 99 things.
And, having been through many rewrites myself, I frequently surprised they could hose the deal. It feels like scope-creep is the killer but I only have feelings for the data, no metrics.
The only way I could get Java 1.x to run (among my fleet of existing machines) was by installing the windows JDK version to run inside Wine on my Ubuntu.
Native Windows 10 could not run Java 1.x. This surprised me, since in my previous experience Windows was pretty amazing for maintaining reverse compatibility.
(I guess I could grab a super-old Debian docker image like Debian 6 and see if the linux Java 1.x binary would run on that, but I wanted to stick with machines I already had running.)
When I used sed and awk a lot I also made heavy use of bash which makes things a lot nicer. And Python, of course.
Every library I needed was at my fingertips on CPAN. And just… worked. Connect to Oracle, run SQL, create a CSV, FTP it somewhere. EZPZ.
I started with Perl CGI and have been less productive every year since then.
CPAN alone is worth the price of admission. I started my career using Perl to build Excel spreadsheets from queries I ran against Oracle DB. That may sound horrible but it was actually plug and play. Really simple and just worked. I could even fire off emails from the script. Just zero hurdles (or guard rails).
I still remember seeing this line from the configure script scroll by.. Made me curious what Eunice was, which was harder than it might seem, given that this was before the web and any search engines.. (and yes, it turns out Eunice was a unix-ish environment on VAX/VMS)
Something that makes workloads usually delegated to find, grep, sed and awk coherent and easy to use, allows streaming data between functions/commands and can still be used to glue things together.
Can I use Perl instead of bash/as my command line ? Should I start learning Perl?
No. Perl is not a command line shell just as Python and Ruby are not.
>> Should I start learning Perl?
If you are doing workloads with the command line tools you list above (find, grep, sed and awk), yes.
If you are familiar with bash, sed, and awk you will find Perl familiar-looking and very handy.
Even though I worked on a lot of Unix boxes doing assorted data processing, Perl was never my goto for the same reason I adopted vi instead of eMacs.
As good as it was, it didn’t come stock on my clients machines, and I could not presume to install it or count on it being there. So, the litany of classic Unix tools prevailed and I simply became adept at working with those.
My few temptations to dip my toes in the modern (at the time) Perl waters just found indecipherable source code (to my ignorant eyes) and disastrous attempts to get whatever it was I was dabbling with out of CPAN. I was never very successful with it.
# this evaluates entirely at compile time!
if (crypt('uh','oh') eq 'ohPnjpYtoi1NU') {print "ok 1\n";} else {print "not ok 1\n";}
# this doesn't.
$uh = 'uh';
if (crypt($uh,'oh') eq 'ohPnjpYtoi1NU') {print "ok 2\n";} else {print "not ok 2\n";}Edit: this applies to any language not just Perl. I use it to learn Python things I would otherwise miss.
Curiously enough, Ventura ships with not one but two perls, 5.18 and 5.30. I haven’t yet investigated which tools there are using Perl. (It could also be that some prominent third party applications still call out to Perl, expect it to be there, and vendors asked nicely. How else would you explain the 5.18?)
It seems to be an apt illustration of the Lindy effect more than anything.
It was fashionable in programming discourse for people to parade their skillful opinions about things they never really understood. And it still is. ^_^
Even today in 2023, some people continue to disparage bash, awk, sed, grep, and Perl as though these tools form a common nightmare for all humanity. But these are very successful tools that solve problems quickly, efficiently in memory, and across many hardware platforms.
People are free to dislike the notation and to prefer Python or Haskell or Powershell. But some detractors simply behave in a bigoted way towards languages that they have little desire to learn. And as far as their objections to the notation are concerned, the same expressive problems find similar notational solutions in Powershell and the various DSLs that Python includes (e.g. numpy, pandas).
Bash, awk, sed, grep, and Perl are robust, durable, portable tools. Their developers were/are smart people. I've found that knowing these tools well has been important in my career. And I also code in Python.