Learn Awk with Emacs (2020)
jherrlin.github.io
jherrlin.github.io
I use org-roam, and put a lot of things in dailies, so I have a temporal log of work that is easily searchable (using deft, or just rg). So much of the mental burden of where to put things is gone, they are in my journal now, and I can always tangle them into a file if I need to, and then I just start linking commits into org to continue to keep track.
Talk about tangling, I also have one big "system config" file that contains all of my rc files and other system configurations and scripts, with sensitive information encrypted transparently with org crypt. I just keep this all in my shared nextcloud Sync folder. I even have configuration and scripts for my homelab and personal server in other files!
Not only that, but I have started relying on org attach to keep a repository of miscellaneous files. Have some PDF or zipfile I dont want to forget about? I just attach it to the daily document and write some notes about it, and wherever I am I can get it. Even have started compressing old projects and "backing them up" into org.
I work through SICP these days in an org babel document, with liberal tangling and noweb, and now I have a journal to myself of my progress, constantly linking to other nodes as I gain more concepts.
I am not a professional computer person, so I don't work with other people. I understand that this works for me because of that.
Would you mind sharing a bit more on your workflow? I am personally going through Crafting Interpreters myself, and I am struggling to organise my code and notes, since I am unable to get org-babel to work like in separate files, and build in one go.
I don't have access to it at the moment, but I wrote a small shell script that invoked emacs (without running init.el/.emacs) and ran org tangle on a file. I incorporated that into a Makefile so it ran on every org file. I placed my org files in the same places (in the file system) as normal source files and had a 1-to-1 mapping of org files to <target language> files (C, Java, Lisp, Go, doesn't matter, done it all). 1-to-1 isn't necessary, I've also done one mega-file that tangled into many source files, including into subdirectories. After tangling, you can trigger any particular build system commands needed (like in rust, run `cargo build` or `cargo run`).
Another thing I've done for smaller things is something like:
#+BEGIN_SOURCE language :tangle foo.language
...
#+END_SOURCE
#+BEGIN_SOURCE sh
build command foo.language
./foo
#+END_SOURCE
Run C-v-t (to tangle the source file(s)) and then navigate to that last block and use C-c C-c to execute it, which will execute the language specific build commands and then run the executable (adjust to particular circumstances). You can also have many of those shell blocks to run different things or in different ways (one to build a release version, another to build a debug version, another to run all the tests, etc.).And it's not just awk, there are Org Babel packages available for virtually all the languages!
- Nim (these are my most comprehensive set of notes): https://scripter.co/notes/nim/
- Tcl: https://scripter.co/notes/tcl/
- String formatting in Nim and Python: https://scripter.co/notes/string-fns-nim-vs-python/
- PlantUML: https://scripter.co/notes/plantuml/
In all the notes pages above, the result of the code blocks is seen directly in the Emacs buffer when I hit C-c C-c. Then I simply* export all those notes to Markdown and publish them using Hugo.
* Tangent: That's one of the main reasons why I went down the path of developing ox-hugo.
I'll get into this if I don't have to twiddle with each one individually but I know emacs enough to suspect that's how it'll be.
Then you need to install the Org Babel packages that don't ship with Emacs or Org mode e.g. ob-tcl, ob-nim.
grep 'foo' file.txt | awk '{ print $1 }'
Becomes: awk '/foo/ { print $1 }' file.txt
There may be times when `grep` is preferable, but this is ubiquitous enough that it's mentioned in the "Useless Use of Cat" [1] awards.I was curious how performance would be impacted, and for a large file `grep` may not be so useless.
$ du -h /tmp/file.json
161M /tmp/file.json
$ time awk -F: '/"id"/ { print $1 }' /tmp/file.json >/dev/null
real 0m17.810s
user 0m17.690s
sys 0m0.083s
$ time (grep '"id"' /tmp/file.json | awk -F: '{ print $1 }' >/dev/null)
real 0m3.617s
user 0m3.641s
sys 0m0.037s
While pushing the filter into AWK is a bit easier on the eyes, it appears there is an incurred performance cost.[1] http://www.catb.org/~esr/writings/taoup/html/ch01s06.html
It was essentially a large list of pretty-printed objects with about 20 attributes each. A few of which had multi-paragraph string values.
Personally I don't think either approach is "better". Sometimes it's easier to deal with some complex logic if it's all in a single awk program, rather than smeared across combinations of grep, awk, and other things. If that outweighs the perf drop in some specific situation, then I'll do it.
You're totally right that commands like `cat *.txt | grep foo` look right and feel intuitive, but the problem is that it's not actually equivalent to `grep foo *.txt`. With a single input file, it's probably not an issue, but with multiple files the downstream command can't actually see what file it's reading from, so in this example grep can't report filenames with matches and you can't use flags like --include/--exclude based on filenames.
I keep wishing for a Unix-like shell and utilities where `cat *.txt | grep foo` actually works the same as `grep foo *.txt`. PowerShell has pulled this off to some extent, which is cool, but it doesn't quite feel right to me compared to Bash.
But then... perl kinda died back. It's no longer default in many distros. Most new unix kids aren't learning it. Perl is the old weirdness.
And... in a world without perl, awk looks pretty cool I guess. But folks: perl is still there.
perl -ne 'print((split())[1])'
While it's true that "{print $2}" is a little shorter, it's also quite limited. The perl syntax extends naturally to different delimeters (the first argument to split is a regex) or actions (maybe you don't want to print it and want to do some tiny string processing).Essentially, you picked the One Single Task at which awk is best. And... it's only barely better, and only for the specific variant of the problem for which awk was designed.
Again, this argument was had and settled. No one in their right mind would have been caught writing an awk script in 1997. That doesn't mean awk is bad (it's software from the mid-70's!), it means it was superceded. And it's important for people interested in those sorts of problems to understand why it got replaced.
I love emacs. These days I increasingly use it for more and more tasks from reading pdfs to writing notes to email. Despite this fact, something about the "doing X with emacs" is starting to bother me for some reason. It's hard to pin down why I feel this way, does anyone else have a similar experience?
You could say that I want emacs-like behavior everywhere, but not everything inside of emacs.
I think you really hit the nail on the head with this statement. It's why I hope projects like NYXT Browser continue to improve.
I use org babel all the time and haven't had any reliability issues.
The whole interactive experience of evaluating awk script and tinkering with it in a single place, greatly helped while learning it.