Sed – An Introduction and Tutorial
grymoire.com
grymoire.com
Only if runtime efficiency turns out to be too slow would I re-examine choice of tools. This is like avoiding premature optimization.
(I've never found perl to be too slow in practice, btw.)
EDIT to manage expectations: the article doesn't explain why, it just provides benchmarks and one commenter made a suggestion about character handling. More insight still welcome :)
The case where I saw a 7x speedup was doing many-times-per-line, fixed-string search/replace on a file consisting of very long lines (an SQL dump where some lines had >1m characters). Perl was IO-bound (so presumably would've been even faster if I'd had better disks), while sed was CPU-bound at a pretty low fraction of the possible IO performance.
For the really complex stuff, there's rejit[0]. I wonder if LuaJIT would work; these tools also need IO tricks to be fast.
$ echo "foo foo foo foo" | sed 's/foo/bar/3'
foo foo bar foo
sed: sed -n '4!p'
awk: awk 'NR != 4'
perl: perl -ne '$. != 4 && print'
Not much between them really.
$ perl6 -pe 'next if ++$ == 2' example.txt
... prints all lines except line 2.
This is an example from Perl 6 One Liners[1].
The `$` is just just an unnamed variable that is getting incremented once per evaluation (-e is for `evaluate`) which in this case happens once per line (-p is for printing each line of input after eval'ing the code -- unless a `next` applies, in which case that line gets skipped).
And...
$ echo "foo foo foo foo" | perl6 -pe 's:3rd/foo/bar/'
... replaces the third foo with bar.
P6 regexes are far easier to read and way more powerful than P5 regexes. The `:3rd` bit is a general language feature called "Adverbs", in this case applied to the regex focused s/// built in.[2]
Quote:
"I asked on a forum what the goals are for relative size and speed of Perl 6 vs. Perl 5, and a Perl 6 developer responded that a reasonable goal would be to have Perl 6 be twice as big as Perl 5 and take twice as long to start up.
"To achieve this goal, the Perl 6 developers will have to shrink the program size by a factor of 6.1 (that is, get rid of about 84% of the code.) They'll need to reduce startup memory consumption by a factor of 13.7 (that is, cut out 93.7% of their memory use) and reduce startup time by a factor of over 275.
"Oh, and this is after they add in all the missing features required to bring Perl 6 up to production-level."
Has the situation gotten better since 2010?
Not really. Startup uses about the same RAM. It's about 10x faster.
The best docs I know about performance would be http://pmichaud.com/2012/pres/yapcna-perflt/slides/slide17.h... and http://jnthn.net/papers/2014-yapceu-performance.pdf#page=72
> "... all the missing features required to bring Perl 6 up to production-level."
The latest story is that the last major missing features (Unicode grapheme-by-default and native arrays) will land in the next few months and Perl 6 will be declared "officially ready for production use" by the end of 2015.
However, for someone who knows neither, here are some reasons you might want to choose sed over Perl:
1. sed syntax is pervasive in other tools. For example, to run a substitution from early on in the tutorial in vim, type :%s/abc/(&)/<enter> from insert mode.
2. sed is simpler than Perl. It used to be that Perl filled a unique role as a scripting language but now there are a bunch of languages in that space (most notably Python and Ruby). Python + sed for example would fulfill most of the same functions that Perl does (and there are reasons to choose Python over Perl as scripting languages, although that's a much more complicated domain to discuss).
3. sed is more performant (or so I hear). This has never been a real concern for me, but some people cite this is a concern.
4. sed usually is a bit more terse. For the length of expressions you'll typically be writing with sed, this isn't a big concern.
Disclaimer: There are probably good reasons to choose Perl over sed, too. Not being a Perl guy, I don't know those reasons. I'll leave that to someone who knows more about Perl.
When I learned sed it was because I was taking a *nix class in college, and a professor pointed me at sed and not Perl. I learned sed instead of Perl because it was put in front of me, not because of any weighing of pros and cons. That's how a lot of learning happens. Sometimes simply learning what's put in front of you leads turns out to be an obvious mistake in retrospect (I was stuck writing VBA for a little while). I don't know whether learning Perl or sed is better for what sed does, but I do know that after maybe 6ish years using sed quite frequently it hasn't turned out to be an obvious mistake.
There are regexes for which awk (maybe sed too?) will perform at a reasonable speed while perl is incredibly slow. Graph: http://pdos.csail.mit.edu/~rsc/regexp-img/grep1p.png Article: http://swtch.com/~rsc/regexp/regexp1.html
Perl...could...employ much faster algorithms when presented with regular expressions that don't have backreferences.
This is a great set of tutorials, he also wrote one about Awk: http://www.grymoire.com/Unix/Awk.html
Get to know these two tools and you'll be amazed at the hours you can save and what you can do, especially with text files.
It's obviously much slower - but I've never been in a position where I needed insane speed to quickly fix a bunch of files.
When you say you store common idioms in your Bash profile, do you mean storing the commands as Bash aliases or functions? I have similar issues with remembering syntax and building sed commands and I've been trying to do something similar to avoid spending time building a complex command from scratch when I already created a similar one some time previously.
sudo ngrep -W byline -d en6 -qilw 'get|post' tcp dst port
so `watchport 8080` will print all network traffic over port 8080 on my ethernet. I actually rarely use this as 'watch_port' directly, but it helps me remember quickly how to bend ngrep to my needs.
For a long time, I didn’t like using aliases because I didn’t want to become overly reliant on my custom aliases – and then miss them when working on an unfamiliar system. This generally worked out alright when I was able to use `Ctrl-R` with a large Bash history. Now, I think that was an irrational rationale and that aliases are very useful shell features. I’m currently trying to organise my aliases and functions into useful groups such as `home_aliases.sh`, `cygwin_aliases.sh`, etc. so that they can be loaded as needed. I then plan on adding them to a git repository so that they can easily be used – and updated – on different systems.
BTW, thanks for letting me know about ngrep. It looks like a useful complement to tcpdump.
sed is bullshit.
Do you know it is not actually possible to get output from sed that does not contain a newline?
:%!sed …
I would suggest just giving it a look directly at https://www.gnu.org/software/sed/manual/sed.html
Though be forewarned, something that neither document explains well is the actual syntax. As in how addresses and expressions can be used and how to read a script. The syntax is relatively simple to understand looking at some examples, but the lack of clear delimiters between the address, command, and command parameters can confuse beginners.
If you stick (too closely) to that one, sometimes you might use features only found in GNU Sed and think you're writing portable scripts.
I think this tutorial helps clarify what is and isn't portable.
Book: http://www.catonmat.net/blog/sed-book/
Free online articles: http://www.catonmat.net/blog/sed-one-liners-explained-part-o...
This was written in 1984 (I think) and still works with a few syntax adjustments. I think it is not bad discipline to return to these tools from time to time and remember core UNIX principles.
http://web.stanford.edu/class/cs124/kwc-unix-for-poets.pdf
I am not so sure anything that I currently am writing would/ could be relevant in 30 years. Very humbling.
I got frusturated escaping for simple replacement: https://github.com/jakeogh/replace-text
echo "/foo/bar" | sed -e 's|/foo|/tmp|'
(The article mentions this)
echo b |awk '/b/{print a}' a=4# search current and all child directories for files containing "bananas", case insensitive $ grep -ir "bananas" .
history | grep "git push"
history | awk '{a[$2]++}END{for(i in a){print a[i] " " i}}' | sort -rn | head2936 git 1166 cd 409 ll 343 ssh 302 l 279 vagrant 223 ls 201 la 151 cat 140 mkdir
history | awk '{print $2}' | sort | uniq -c | sort -rn
https://github.com/junegunn/fzf
It's a fuzzy searcher for Ctrl-r that shows all possible results, updating as you type.
But how do you then pipe that to tail? :D
What you can do is press Ctrl-R again and it will search the older match. There's also forward match but I can't recall the shortcut.
Also, for general bash stuff, these are great, and IMO much better than the TLDP site that always ranks high on Google:
Definitive Guide to sed
I found it to be well worth the money, though I wish it were available as a PDF.
Or, presuming you're on a modern browser and care that much about the content, you can just inspect the dom, find that <link type="text/css"...> in the head, and delete it.
Yeah or you could go read a better book
The site is ugly, but it's one of the best references for sed.
View > Page Style > No Style
View/Page Style/No Style
It also looks pretty good in the lynx browser, and probably other text browsers. https://en.wikipedia.org/wiki/Lynx_%28web_browser%29
[1] https://chrome.google.com/webstore/detail/clearly/iooicodkii...