Useful sed scripts and patterns
github.com
github.com
FWIW, I'd argue that the path to true sed mastery eventually goes through the hold space, which isn't mentioned here. For some real fun, and an exercise that might change how you mentally model what sed is capable of, check out SedSokoban.
You're right, I'm using -r even when it's not necessary. To my defense I think it's a good habit to have since without it regex expressions are painful to write. I didn't considered that using -E it's a better choice but I'll correct that now. (one might argue again that typing -r is easier than -E :D ).
Regarding the definition of words - I also thought of that when I wrote that snippet. I know it's not the complete regex for a word and that word regex patterns might differ. And I was probably a bit lazy - but I'll correct it presently.
I'd also like to say that I didn't write this as an absolute and ultimate reference. If I'm honest I wrote this as much to teach others as to solidify this knowledge myself. Now since it seems it gained traction I'm kinda obligated to make this better, no? Darn. :)
PS: if you'd like to help me make this better please submit a pull request or leave a comment here. Looking at your profile I see that you try to limit your online time so I probably shouldn've asked. :P
Example:
sed '10{p;q;}' myfile.txtPS: May I humbly point out that the command you provided will actually print up to line 10 and then a duplicate line. Like I said, unexpected (even though it seems logical). :)
Example "multiple replacements":
#!/usr/bin/sed -f
s/a/A/
s/foo/BAR/
s/hello/HELLO/
Save it as replace.sed;
chmod u+rx replace.sed; ./replace.sed myfile.txtAlso, FYI, this is the link for the POSIX specification for sed:
https://pubs.opengroup.org/onlinepubs/9699919799/utilities/s...
Notice that -i is not POSIX.
This document shows practices of sed, but the author doesn't seem aware of best practices. A sed script starting with '#!/bin/bash' is a bad practice.
What other good practice suggestions do you have? (if you're comfortable answering and have the time)
And if you do use them frequently enough, you might not even need to refresh/relearn anything.
I think by {} you mean an argument to the find command.
To me it's bash, sed and awk for the basic text-processing things, and then perl for anything that requires more programming. These are very similar syntaxes (and ways of thinking I'd say), so they fit together great - and also are close enough to php and js that I do at work to make the transition painless. But if I had to write it in python or ruby or God forbid C++ - which all I used a lot, but long time ago - I'm sure it would be a big struggle for me now.
I sometimes forget the exact syntax for pipes and file open/close, but it's not hard to find in 40 pages.
AWK is in busybox, and a lot of other places that Python simply cannot go.
The control structure syntax is also the same a C/JavaScript/PHP/C++ etc., so it certainly does not hurt to know it.
https://archive.org/download/pdfy-MgN0H1joIoDVoIC7/The_AWK_P...
I think the problem is that they fall into an uncanny valley between "simple tool" and "verbose programming language"; they are really terse languages wearing a simple tool's sheepskin.
And when one is using awk/sed that script is often invoked from a shell, bash for instance. And since the awk/sed can't hold its own as a general purpose programming language, now I am writing bash code, a personal sin.
Awk/sed made sense for the time and place they are from, but things have changed.
Okay, acknowledged. Python and Ruby are superior in every regard. Would you mind letting us chat about sed and awk now for a minute?
Also.. I use Awk every day, for all kinds of things. I forget the order of parameters in split, sub etc, which takes a few seconds to look up with man awk. I got into sed a while ago but never use it. Awk, together with sort and uniq sometimes, is all I need.
And it is interactive. You have an input. Print that out and look what needs to be done. Write the first transforming step in Awk. See what that did. Add the next step. See what that did. Extremely interactive. After a few minutes I have a one-liner (often 2 or 3 wrapped lines) that does exactly what I want, and easy to adapt to other similar things. I have very large bash history enabled so they stay Ctrl-R-accessible.
(I might be wrong about "everything" but that's what bomb-tossing is all about)
"Please don't sneer, including at the rest of the community."
Great way to have confidence in a port and provide for a fallback path. Eventually the new code will take over the old code and the old code will wither and die.
Don't do that. Do not assume something is of no use because you don't know how to use it. Double that recommendation for tools that have widespread use. There is a reason they're used everywhere.
In the case of awk/sed/perl: Nothing beats them for composition in one-liner text processing pipelines.
Don't assume I don't know how to use it. :)
One line verses ten, I'll take the more readable thing over the APL.
Can't find links now, but fairly certain were stories in past about how switching to AWK from other tools increased performance
I think that if you have a good mental model of how awk works, then you know times that it’s the right tool for the job. And on those occasions you’ve got plenty of time to go search for some syntax, make a coffee, maybe even go for a run all with change to spare compared to using pretty much anything else.
I will admit, that if I can do it with a few greps, seds and cuts then I tend to just use those first, because there no thinking involved for me.
What we really need is something that can "compile" down to sed for deployment in that kind of environment.
Then that compiler can contain all the weird stuff you need to remember.
curl cheat.sh/sed
Will show you several examples.
Also, LDP: Advanced Bash-Scripting Guide: https://tldp.org/guides.html
I basically use it for simple text substitutions; in that case it is the same as what I use for vim editing.
It may not be as powerful but it's just one thing to learn. Also it understands regexs like '\d', etc which are kinda universal but for some reason don't work with many of these utils.
I use sed for a lot of simple stuff, but I use Perl all the time for more advanced stuff. Perl is great for grabbing complex chunks out of a line and doing something with them that requires a bit of state. For example, writing a quick and dirty logfile parser to figure out how many requests and how many bytes were used by each IP over a time period (the time period limiting usually being a prior grep before being passed to Perl). I've probably taken 60 seconds to write out something to parse exactly what I want from a log hundreds of times.
What about setting up the project for the application, setting up instrumentation and testing, debugging, etc.?
Even supposing that all that extra work makes sense for your unique string manipulation needs, what about the next time something like this comes up but is out of scope for your first tool? Do you end up with lots of little applications, with lots of little features and arguments? When you come across that similar-but-not-quite need down the road, do you have to scour your library of full blown apps to find something that fits? Not to mention maintaining them against future system updates...
OR - you could just force yourself to get over the hump, learn the conceptual basics at play and then when you come across a novel situation down the line you can just rattle off a tailored one-liner and have your answer in a few moments...
I'm going to let you in on a little secret: many of us who have been doing development or systems administration in some form another for multiple decades are the same way. I have a terrible memory, but it turns out that does not have to be a big hindrance for tech work.
When I'm working (or tinkering at home), I keep a tab to my personal wiki open at all times. I will not usually bother to write something down the first time I do it. But I have to google it twice, I throw it into my wiki. This does two things: 1) It makes me (slightly) more likely to actually remember it for next time, and 2) The next time I need the information, I know where I can find it without having to wade through SEO-encumbered blog spam or outdated StackOverbutt answers.
A database of personal notes is a powerful multiplier for technical ability and productivity. To the point that whenever I interview a candidate for a job, one of the things I ask is what system they keep their notes in. Their system doesn't matter at all to me but if the answer is none, that's a definite strike against.
>I really don’t think non-interactive CLIs are a good way to do complicated tasks.
Many of us thrive on the Unix command line because of the raw power you can wield with it. "Why waste time trying to learn arcane shell commands when you can just write a script in Python?" is a common refrain I hear from developers with only a few years under their belt. My response to that is, "Why would I waste time writing a Python script when I can do it in one line of shell?"
The first time you look something up, it might not be obvious if you'll ever use that info again, so why waste the time and clog up your notes with a bunch of noise? But if you've had to look something up twice, odds are good you'll have to look it up again someday.
And like you said, writing a note in a way that's clear and understandable helps you to recall the thing a little better next time
Any suggestion for a remotely accessible personal wiki? (I'm assuming that access is restricted for potential "personal" stuffs).
Imagine having this attitude towards tools in general, you'd throw out most of everything. Anything that's powerful has a learning curve
I love awk and whenever I __do__ find a task where awk helps me, it's really an amazing experience and it is always a boon, never a burden.
But I really suck at remembering the syntax as I only hit tasks that awk solves well rarely. What I do remember is situations where awk definitely can likely help. Same with sed, tr, etc.
Keep in mind, every time we automate something or introduce a new tool/process, it's a cost benefit analysis; when I have dozens of gigs of logs to parse to figure out which hosts out of hundreds are having issues, it's a no-brainer for me; the 30 minutes to revisit a few awk and bash tricks is a far better investment than trying to grep/scroll my way through literally millions of lines of logs/code.
When I first learned awk I definitely went through a 'hammer' phase (i.e., when you have a hammer, everything looks like a nail); once phase passed, I find that I'm much more disciplined on when I break out awk or similar tools.
Ultimately it's about identifying when you have the "right" tool for a specific job; I might need to invest time to revisit some syntax, but if the end result is that my 30 minutes saves hours, then it's a good 30 minutes. I'm at the stage where the basic text munging with awk is a no-brainer (or at least I remember the mistakes I made previously and how I fixed them); this was achievable only after a few projects where I did have to spend a bit more time wrangling awk than I preferred, but it saved me a lot of time in the future. Just knowing what I __can__ do with a given tool helps me make educated decisions on where I should invest my time on a given issue.
sed '1~2p' file.txt #needs -n option
s '1,$ s/foo/bar/' file.txt #s --> sed and 1,$ --> 5,$
Many commands using `-r` do not need the option for the command used (for ex: `sed -r '/start/q'`). Also, using `-E` is preferred instead of `-r` since some of the other implementations support this option but not `-r`.---
I wrote a book on GNU sed with plenty of examples and exercises: https://github.com/learnbyexample/learn_gnused It is free to read online and there's a detailed chapter for learning BRE/ERE regex flavor as well.
sed -e s/a/A/ -e s/foo/BAR/ -e s/hello/HELLO/
I use -e often enough that I usually use -e whether I expect to use more than one expression or not. There's a reasonable chance that I will be going back into my history and will need to add another, which is easier if the first -e is already there.
Also, 'sed -e expr1 -e expr2' gives the same results as 'sed expr1 | sed expr2'. In both cases, the order of the expressions matter: later expressions may change things altered by the first.
$ echo foo | sed -e 's/f/g/' -e 's/goo/go/'
goThey don’t need to be ready to go, but ideally:
- natural language searchable
- add a small description
- CLI to search, examine, and copy
- not a <favourite text editor> plugin
- sync with a GitHub repo.
Use a flat text in the editor of your choice. Search as needed. Tags work because you can seach them. You delimit each section as a note. I do it like this
---
Next note
---
It seems limited and primative, you don't get any markup or fun stuff. But it's soooo good. There is nothing to break or update. Seach is the method I use most anyway in my other note taking apps.
And you wind up putting everything in there, including commands that you're staging to run. Those staged commands become snippets later.
Eventually I'd like to publish it as a microblog for such snippets ala commandlinefu.com, but priorities and laziness get in the way. :/
Every now and then I start a new sheet, dropping off the stuff I've memorized.
I have txt files for commands I want to remember for each OS I daily use (linuxcmds.txt, windowscmds.txt, maccmds.txt) and split them the same way you said but 5 dashes. Also keep txt files for install/setup notes for each os, links to every program I want to install in fresh installs, etc etc.
Text files work on every single OS/device, the format is stable, it’s easily searchable, easily shareable, and my entire 30 years of notes is measured in megabytes.
Some of them are shared via DropBox and despite the risk of simultaneous edits we had basically no incidents: the simplicity and efficacy justifies it for us.
First, set your HISTSIZE to 1 trillion. Then, tag the relevant commands in the history file with a comment after the command (e.g. sed -n '5,10p' some.log # print selected lines). It gets searchable within the shell, either by the comment or the start of the command. Always at my fingertips.
My linux only snippet manager: https://github.com/barbuk/snippy
Shell completions are essential for it, so I had to stumble through writing my own for Fish, but now that it's setup it's quite nice.
Btw if you like fish autocomplete, you might be interested in fig.io.
We've spent a bunch of time making is super easy to add your own completions for scripts or custom CLI tools. :)
(That's a web version but I usually access it on the cli or in my editor)
And I just became aware from reading the README that there's an even faster one written in Zig.
Or a personal wiki, I know people that run their own wikis for just this sort of thing
Sed is the silent version of ed. It's a full line-oriented editor. You give it commands and it executes them on lines, one line at a time. Commands always have the same form:
[lines range] [action][parameters]
You just have to learn how ranges are defined and what the actions are then you can check the documentation for what the parameters are when you need to and you will be good to go forever.As a bonus, once you realize there is more actions than just s for substitute, you start realizing that people sometimes do very convoluted things with sed because they don't know how to use the other actions. For example, there are actions to join lines (j), to delete lines (d), to prefilter lines based on a regex (g and G).
As someone with over a decade of heavy sed use, who only got into awk recently, i'd say awk is far easier and less magical. Awk is just another programming language, for the most part. Using sed for anything substantial requires you to think in strange directions, to find ways to thread state and control flow through the tiny holes the language gives you.
As you rightfully pointed, I wouldn’t do anything overly complicated with sed because it really is a line editor and anything which is not line based and can’t be explained simply in the context of using an editor will quickly seem clunky and overly complicated.
I wouldn’t do anything period with awk. I agree with you that it’s just another programming language but one which feels extremely dated with both a terrible syntax and awkward semantics. The rare times I have encountered something done with awk it was either so simple it should have been done with sed or so complicated I wish a proper scripting language like python or perl would have been used instead. As far as I’m concerned, awk is a tool without a proper use-case. It’s always better to use something else.
These days i reach for awk for a lot of log processing, particularly when it needs to be stateful. For example, i have a log file with entries for remote clients connecting and disconnecting. An awk script can loop over that and keep an array of clients, adding and removing them as appropriate, and be easily tweaked to report different things - emit an event for each connection-disconnection pair, or track the number of clients connected over time, etc. I couldn't do that at all with sed. It would be more awkward with bash. It would be easy enough, if a little bit more verbose, in Python.
Btw, there is one more use case of a variable delimiter which is more arcane (and you can combine it with the other custom delimiter) `sed -s '\_/bin/bash_s:grep:egrep:' myfile.txt`
For "keep the first word of every line", I prefer awk:
awk '{ print $1 }' < file.txtHere's a groovy script to do the same:
new File('file.txt').eachLine { println it.split()[0] }
Not as neat as awk, but not much worse either... and as I use Groovy a lot for testing, and its syntax is simplified Java (which I use a lot as well) it's just much better for me... if I had to ssh into other servers where Groovy or Lisp (my other favourite "scripting" language) are not available, then yeah, I would learn awk better (I actually like awk, and use it sometimes for the rare occasion I don't have groovy/lisp installed).But if it's going to be used a lot, and start-up time matters, then some of the full programming languages can become expensive for the temporary load they create just spinning up. In those cases, the standard Unix tools tend to be so much faster and lighter that it's worth learning them.
sed is just a DSL (domain specific language) specifically designed for for manipulating text files line by line.
Note that for matching strings even in Groovy/AWK/Python/Perl/... you'll probably use regexp which are another DSL. Would you recommend to use string functions instead?
Actually, yes, of course. Regex should be the last resort, really. It's unreadable, presents security issues depending on where it's used, and much of the time using string functions in a "proper" language will be actually easier unless you're a regex expert (which is a pretty silly thing to acquire expertise on).
Me too, though the awk version throws out any leading whitespace, so it's not quite the same.
I feel like sed is what deserves the donation…
I didn't added explanations (although I wanted too) since this was intended as a quick tips page for sed. I think there are much better guides than I could ever make (including the info page). In a way it works better since it's more digestible.
Thank you for the link you provided - it looks awesome! I'll look into it and append the existing guide if needed (while giving credits of course)
For example suppose I need to add this setting to /etc/ssh/sshd_config:
PermitRootLogin yes
One idea that comes to mind is to first delete the line with the line shown in this post, and then add it back:
Regex sed -E '/^#/d' file.txt - delete lines where regex matches
sed -E '/^PermiteRootLogin/d' /etc/ssh/sshd_config ; echo 'PermitRootLogin yes' >> /etc/ssh/sshd_config
Is there a more straightforward way with sed?
> > sed -n '10p' myfile.txt
Which line? Why p? I can make assumptions that this means “line 10, print” but easy-to-explain could go a small step further and complete the explanation.
Or you can use Org-Mode (with babel), if you're an Emacs user.
That's how I write all documentation involving some sort of code.
sed -E s_[a-zA-Z0-9_]+.*_\1_' file.txt
You're missing an opening single quote and a set of parentheses for this to work. It feels like you should have a unit test setup for these snippets, given that there are other typos found in other comments here."I made an OpenAI-powered Linux shell that guesses your bash command": https://www.youtube.com/watch?v=j0UnS3jHhAA
sed 1d
Select line 123 sed -n 123p
# or
sed -n 123{p;q}https://gist.github.com/awhileback/1fffac899d9321a6a9ec15bc9...
Requires some ocr / pdf parsing / metadata tools, see the comments.