Aho – a Git implementation in Awk
github.com
github.com
But you can use awk as a general-purpose scripting language [1], in many ways it's nicer than bash for this purpose. I wonder why you don't see more awk scripts in the wild. I suppose perl came along and tried to combine the good features of shell, awk, and sed into one language, and then people decided perl was bad and moved on from that.
[1] Random excerpt from NetBSD's source code https://github.com/NetBSD/src/blob/trunk/sys/dev/eisa/devlis...
Screw what people think. I found out I like perl. The last thing I wrote is a programmatic partition editor [1] - like how you use sfdisk to zero out the partitions, except I wanted to do more than zap, like having the MBR and GPT partition table to combine them and make hybrids.
I was fun, and I will use perl again (I may also use awk at one point now that I see how cool it is)
Which is not to say that nobody ever figured out those things and did them well, just that the success rate was low enough across the industry to earn Perl a really bad reputation.
I'd like to see a revival of awk. It's less easy to scale up, so there's very little risk that starting a project with a little bit of awk results in the next person inheriting a multi-thousand line awk codebase. Instead, you get an early-ish rewrite into a more scalable and maintainable language.
awk v gawk doesn’t make me want o to relive those days.
That's a fair point. I always explicitly write my scripts to invoke gawk so that I don't accidentally invoke a different version.
Taco Bell programming is the way to go.
This is the thinking I use when putting together prototypes. You can do a lot with awk, sed, join, xargs, parallel (GNU), etc. But it's really a lot of effort to abstract in a bash script, so the code is compact. I've built many data engineering/ML systems with this technique. Those command line tools are SO WELL debugged and have reasonable error behavior that you don't have to worry about complexities of exception handling, etc.
And it’s not even that Python is a great language. Or has a great package manager or install situation. It doesn’t have any of those things. It does, however, have the likelihood of the next monkey after me understanding it. Which is unfortunately more than can be said about Perl
Job descriptions tend to be looking for rails developers and forgetting that actually its Ruby developers they are looking for.
Most all “Linux” cannot even boot without python, and it is quite easy to find a minimal Linux distribution that does not have a dependency on Perl.
A historical note: Perl was that language before Python was, and it lost that status to Python through direct competition. For a while, if you had to do anything larger than a shell script but not big enough to need a "serious" C++ or Java codebase, Perl was the natural choice, and nobody would argue with it (unless they were arguing for shell or C.) That's why Perl 5 is installed on so many systems by default.
When I first started using Python, I felt a little scared for liking it too much. I thought I should be smart enough to prefer Perl. Then Eric Raymond's article about Python[1] came out in Linux Journal in 2000, and I felt massive relief that a smart person (or someone accomplished enough that their opinions got published in Linux Journal) felt the same way I did. But I still made a couple more serious attempts to force Perl into my brain because I thought Perl was going to be the big dog forever and every professional would need to know it.
But Perl was doomed —- if Python didn't exist, it would have lost to Ruby, and if Ruby didn't exist, it would have eventually lost to virtually any language that popped up in the same niche.
Personally I don’t think awk is a good choice for anything beyond one liners and personal scripts. Here it was fine because it was (initially) some write-once academic code that needed to not be insanely slow.
That's a large part of what's driving Awk's renaissance: devs that never learned Perl to begin with want something to fill the gap between shell and Python, and other devs like me who (reluctantly) abandoned Perl because it was deemed "uncool" by HN types, which means Perl and all code written in it now has an expiration date on it. But since Awk is a POSIX standard, HN types can't get rid of it.
Grab https://raw.githubusercontent.com/sabotage-linux/sabotage/ma... and clobber the perl script in the Linux source.
Then: sed -i -e 's@perl $(srctree)/$(src)/build_OID_registry@awk -f $(srctree)/$(src)/build_OID_registry@' ./lib/Makefile
They also removed the perl dependency for ca-certificates since one of the goals was to remove perl dependencies from the core system including its toolchain and kernel. It's not needed at all now.
This Aho project is neat because it has the potential of removing the perl dependency on having a git client, which was a problem prior.
To me, hell is having to debug perl scripts other people wrote. Based on experiences in the 1998 time frame.
Because you have only yourself to blame.
I stopped using Perl because my egg drop bots got laborious along with my expect scripts. They were novel early on, but maintaining them became something of a chore I wasn't inclined to do anymore. Other things started to do that better.
Personally, I think Python won because it's syntax was much more readable. It had nothing to do with technical merits. DX is a strong subliminal motivator.
[0]: If you're wondering why there wasn't adequate observability in place to not have to do this, you're not wrong, but sometimes you must live with the reality you have.
[1]: Yes, I know gawk can do a unique sort while building the list. It was late into an incident and I was tired, and | sort | uniq -c | sort -rn is a lot easier to remember.
[1].a: Yes, I know sort has a -u arg. It doesn't provide a count, and its unique function is also quite a bit slower than uniq's implementation.
$ python3 times_python.py strptime_listcomp
parse time: 45.96 seconds
count time: 0.54 seconds
count: 498621
$ python3 times_python.py slices
parse time: 9.96 seconds
count time: 0.40 seconds
count: 498621
$ python3 times_python.py isofmt
parse time: 0.80 seconds
count time: 0.38 seconds
count: 498621This is without footguns, because concatenation is a different operator entirely.
2 + "3" -- 5
2 .. "3" -- "23"
It's rather convenient, especially when you can combine it with LPeg to parse files with numbers in them, then do the maths directly.Doubtful. awk has a lot of implicit behaviour which allows the programmer to write very terse scripts. An equivalent Python program is usually several times longer.
I mean, russian is also unreadable, until you know how to read it.
awk's power for me isn't the LOC needed to accomplish a task, its power is that I can express the business logic I need very easily and very quickly, and the resulting code is really fast. I am by no means great with awk, but I can go months without touching awk, encounter some problem where awk shines, and in a few minutes or less I have exactly what I need.
you can learn to be extremely productive with awk in a few hours and it's very comforting to have this in your toolset moving forward. Essential? Probably not, but I like that I don't need to break my thought process when working with awk because it's just so natural to express what I want awk to do that I don't really "think" about writing awk, I just write it.
For the problem domain that Awk targets, its close to as good as it gets. Lot of the line reading, delimiting, chunking etc is already done for you. You the programmer dont have to deal with re-implementing that same old ceremony. You get straight to the point of writing the transformations that you want.
If you break out into Python for that, that's a bit like being at your best and formal self, your best table manners, the first time you are meeting your fiance's parents. Awk on the other hand is like being with you old childhood pal, partner in crime, from your old street where the need for such ceremonies do not exist.
With modern awks you can write extensions for them. Mawk is plenty fast too.
One thing that still grates though is that the 'split' function does not come with a corresponding 'join' and one has to iterate through the array explicitly to join.
Like any programming language, you have to get good at Perl to write good Perl.
Python has clean notation. It's a juggernaut that has changed people's expectations of language design.
But even so, Python has not decreased the amount of bad code in the world. Not even within Pypi.
Reading legacy scripts was wild back then - you had to be somewhat good at bash, unix tools like awk, C, make, Perl:-)
The first edition was published in 1988, and is available at:
* https://archive.org/details/pdfy-MgN0H1joIoDVoIC7
* Discussion: https://news.ycombinator.com/item?id=13451454
Awk could use improvement in numerous areas. Oh, for instance, you can pass associative arrays into functions, but not return them. Functions that filter array to array have to take an output array parameter.
Using extra parameters as the only way to get local variables is also a smell.
a[i] syntax cannot index into strings, what the hell?
Of course he could have died since, but as others noted, he hasn't.
function read_objfile(obj, objpath, bytes, end_of_header, header,
end_of_type, type, size,
bytes_after_header)
the parameters separated by the big white space are local variables. It's possible to pass them values, but you're not supposed to.I wrote a patch for GNU Awk to give it a let statement for binding true lexical variables, so that this could be:
function read_objfile(obj, objpath)
{
@let (bytes, end_of_header, header,
end_of_type, type, size,
bytes_after_header)
{
}
}
Unfortunately, this was rejected by the project; I was encouraged to make a renamed fork of GNU Awk, so that's what I did.If you're looking to get into Awk, and you learn well from a lecture style, I put together a talk for Linux Fest Northwest some years ago and recorded it for Youtube: https://youtu.be/E5aQxIdjT0M
It's never going to be as powerful as a specialized IDE would for a specific language, but the Linux CLI is language agnostic and even works on just text files, so it's universally applicable and doesn't change depending on the language of the project. For me it's better than an IDE, but YMMV of course because everybody is different.
Awk is an impressive tool, but putting it on a pedastal blinds people from the weak spots and why they should probably move serious tasks to more specialized, but modern and adapted tools.
That only comes into effect when the input is coming from a pty, and that only when it's in a very specific mode that is meant for interactive use.
I believe it was Linus himself who quipped (paraphrasing) that he is so arrogant he named two software projects after himself: Linux, and git.
Back in the day I created a Web-based wiki using awk. Why? Because I was using linksys router with minimal memory.
It was a great learning both how wikis work and what can be done with awk. And since there are no libraries to fallback on, I had to implement the basics and gain all the understandings.
You can also git clone from a repository in a different directory in the same computer. And push to.
BUGS
Multiple checkout in general is still experimental, and
the support for submodules is incomplete. It is NOT
recommended to make multiple checkouts of a superproject.
Source: https://git-scm.com/docs/git-worktreeIt complicates git with more cruft. A separate clone is more understandable and independent. If you trash something in its .git/ subdirectory, only that repo is affected.
In my use case, independence would be an anti-feature. git-worktree fulfills a specific desire that I cannot fulfill with any known alternative. Therefore I use it.
As for actual bugs, I certainly never encountered any. But I also don't use it with repos that contain submodules, as per the warning.
sed-chess: https://news.ycombinator.com/item?id=37896854
awk-raycaster: https://github.com/TheMozg/awk-raycaster
[1]: https://github.com/freznicek/awesome-awk [2]: https://github.com/Ypnose/ahrf
That said, it should be a criminal offense to write any tool this large and complex in any language that can’t be used in a powerful step debugger.
TBH I’m increasingly frustrated by the amount of code written in Bash. I kind of hate Python for various reasons. But if 100% of Bash was replaced with Python I think the world would be a better place.
That said, stdout/stderr is such a bloody, inconsistent nightmare. I’m not totally convinced that “chain small binary programs together” is better than “one language with useful libraries”.
Bash is admittedly nice for small things. But it always spirals out of control. And rarely gets ported.
Also my life is primarily Windows and if you want everything to “just work” across mac/linux/win it’s easier to just use Python or sometimes even Rust. I often wish I could easily write and run single-file rust scripts.
x = `git --version`
puts(x) #=> "git version 2.43.0"Bash has it's own uglies so I guess it's fair enough to compare two imperfect things, but the problem is these are two different jobs, and python just doesn't do the job bash does, and that job needs doing, so python can't be the replacement.
eg:
function run_command( c, shortopts, longopts, quiet, directory, path, errors)
Just asking as this project has kind of resparked my interest in awk.This spacing convention is meant to clearly separate mandatory parameters and optional parameters that are sometimes only introduced to "declare" a local variable.
A Git Implementation in Awk - https://news.ycombinator.com/item?id=28771841 - Oct 2021 (96 comments)
You should look at git-lfs[1] instead.
Also, I'm not a huge fan of tools that implement important functionality as an afterthought, especially if those tools deal with my precious data.
The core object storage model and data format (and many, many things on top of those) have to be changed/extended/fixed first, but it's realistically an immense change, so git-lfs and other various solutions are about as good as it'll get in the mean time.