Linux Text Manipulation
yusuf.fyi
yusuf.fyi
If you are early in your career, I suggest you work on these types of skills. It is surprising how often I have found myself on a random box that I needed to parse application logs “by hand”. This happens to me even in fancy, K8-rich environments.
Just FYI, Kubernetes is abbreviated to k8s, not k8.
> K8s as an abbreviation results from counting the eight letters between the "K" and the "s".
[0]: https://www.wordnik.com/words/a11y [1]: https://www.wordnik.com/words/i18n
Btw txn is similar, but they didn’t bother with numbers, they just replaced the middle with an x.
"a11y" is especially bad, since it runs counter to the very thing it is supposed to represent.
I hope the fad passes quickly.
But it's equal-opportunity annoying I reckon: no one knows what the hell `a11y` is about when they first see/hear it, but not in a way that's more onerous for screen reader and braille users than for anyone else.
It’s surprising how many times you have to ad hoc parse due to the tools being so poor. It’s endemic.
$ sgpt -s "The command 'sp current' outputs
> Album Tea For The Tillerman (Remastered 2020)
> AlbumArtist Yusuf / Cat Stevens
> Artist Yusuf / Cat Stevens
> Title Wild World
> I want just 'Wild World by Yusuf/Cat Stevens'"
sp current | awk -F' +' '/Title/{title=$2} /Artist/{artist=$2} END{print title " by " artist}'
[E]xecute, [D]escribe, [A]bort: A
E
Wild World by Yusuf / Cat Stevens
Looking back at it, the awk command it uses is actually pretty clean.More importantly, you can get the same results by just one or two dbus-send commands instead of using that "sp" script + something to clean it up.
The sp-metadata function uses tons of processes to clean the output[1], and sp-current launches a few more. If you're doing this in a loop for your WM status display this sort of stuff really adds up. Even on modern systems launching processes isn't free and relatively slow, and launching >20 of them every few seconds is going to use non-trivial amounts of CPU and will needlessly drain your battery.
I don't really know what that dbus command outputs, but I bet that with you might go a long way with "dbus-send ... | grep -o ..." or something.
So in general I'd say this is a classic case of "you're not even using the right solution, and no amount of GPT is going to help".
[1]: https://gist.github.com/streetturtle/fa6258f3ff7b17747ee3#fi...
The "sp current" issue seems less fair to bring up, as the entire article was using that throughout as well.
Out of curiosity I asked bog-standard web ChatGPT (4) if it could do the job in awk, took it five times to get it right. Whatever prompt shell-gpt is using, it works.
Shell scripting is the perfect use case for ChatGPT. Simple enough that AI can handle it, annoying enough that I don't really want to do it myself, and something I do rarely enough that I don't really know any of the commands deeply.
I'm sure awk could do that, but with Perl:
sp current | perl -nE '/(\S+)\s+(.*)/ and $d{$1}=$2;END{say "$d{Title} by $d{Artist}"}'Perl has fallen from grace as a general programming language and even as a systems administration language, but it's still absolutely the best and most ubiquitous tool for text manipulation.
Text::Balanced (for extracting the quoted text) and Text::Wrap would be the core of such a program.
awk '{ k = $1; sub("^[^ ]* *", "", $0); d[k] = $0; } END { print d["Title"], "by", d["Artist"]; }'
Personally I'm a bit of a sed fan sed -nE '/^Artist/{N;s/^Artist +(.*)\nTitle +(.*)/\2 by \1/p;}' playerctl metadata artist
playerctl metadata title
as provided by MPRIShttps://wiki.archlinux.org/title/MPRIS
> MPRIS (Media Player Remote Interfacing Specification) is a standard D-Bus interface which aims to provide a common programmatic API for controlling media players.
> It provides a mechanism for discovery, querying and basic playback control of compliant media players, as well as a track list interface which is used to add context to the active media item.
function sp-metadata {
# Prints the currently playing track in a parseable format.
dbus-send \
--print-reply `# We need the reply.` \
--dest=$SP_DEST \
$SP_PATH \
org.freedesktop.DBus.Properties.Get \
string:"$SP_MEMB" string:'Metadata' \
| grep -Ev "^method" `# Ignore the first line.` \
| grep -Eo '("(.*)")|(\b[0-9][a-zA-Z0-9.]*\b)' `# Filter interesting fiels.`\
| sed -E '2~2 a|' `# Mark odd fields.` \
| tr -d '\n' `# Remove all newlines.` \
| sed -E 's/\|/\n/g' `# Restore newlines.` \
| sed -E 's/(xesam:)|(mpris:)//' `# Remove ns prefixes.` \
| sed -E 's/^"//' `# Strip leading...` \
| sed -E 's/"$//' `# ...and trailing quotes.` \
| sed -E 's/"+/|/' `# Regard "" as seperator.` \
| sed -E 's/ +/ /g' `# Merge consecutive spaces.`
}
But as the other replier mentioned, the point was to show off an example of how text manipulation skills can solve many problems, not solve this specific problem in the best way possible. $ txr by.txr spdata
Wild World by Yusuf / Cat Stevens
$ cat by.txr
@(gather)
Artist @artist
Title @title
@(end)
@(output)
@title by @artist
@(end)
For one-liner outputs, I often use @(do (put-line `@title by @artist`)). $ cat sp_out | lines | parse '{key} {value}' | str trim | transpose -rd | format pattern '{Title} by {Artist}'
Wild World by Yusuf / Cat Stevens
https://www.nushell.sh/cookbook/parsing.html sp current | pyxr -0 -g "(Artist)\s+(.+)\n(Title)\s+(.+)" -p "{3} by {1}" duckdb -list -c "from (from read_text('sp-current') select regexp_extract_all(content, '(\S+)\s+(.+)', 2) as value) select concat(value[4], ' by ', value[2])" | sed -n 2p