Removing duplicate lines from files keeping the original order with Awk
iridakos.com
iridakos.com
[1] https://github.com/learnbyexample/Command-line-text-processi...
Thanks for linking this repo!
Somewhat related blog post which I like to refer people to: "Command-line Tools can be 235x Faster than your Hadoop Cluster" https://adamdrake.com/command-line-tools-can-be-235x-faster-...
Chapter 2 of The AWK Programming Language has incredible benefits for a novice.
https://archive.org/download/pdfy-MgN0H1joIoDVoIC7/The_AWK_P...
<line-condition> { <code> }
That's unusual enough among programming languages to call it odd. Being able to do stuff like if (/some-pattern/) { ...
and have the regex be evaluated like a condition where it matches with the current line implicitly is also pretty unique. if (/some-pattern/) { ...
This isn't really unique when you consider perl. while (<>) {
if (/pattern/) {
This does the same. Awk simply has the implicit loop.But it feels like it would be a net loss based on how seldom I currently need to write one-off scripts.
Based on experience, I'd probably have a perfect use-case for it every 2-3 years.
Sure, the days you'll need awk, you'll take 15 minutes instead of 2 writing your script. So what ?
But the rest of the year, you'll have a more versatile toolbox at your disposal for automatic things, testing, prototype network processes, make quick web sites or API, and explore data sets.
That being said, I can see the point of learning awk because, well, it's fun.
And we already have that: it's called perl. :)
IMHO knowing what kinds of tools exist and how they are used for different tasks is enormously useful. Most software projects require me to create a set of tools to solve problems in a certain space efficiently. In any long-living non-trivial project there will be feature requests you couldn't have anticipated in the beginning. They tend to be painful if your program is just a bunch of features hacked together. But if you take a tools-first approach, the unexpected features can often be solved with what you have.
Of course, time is limited and you can't learn everything. But learning one of every different kind of tool is a very good use of time.
EDIT: Note that I'm not claiming you'll build a web app with awk. I'm saying you might write code that can be used similarly to awk in some abstract sense, and that might be a core part of a web app.
awk '{ if (! visited[$0]) { print $0; visited[$0] = 1 } }'
for ex:
awk -o '{ORS = NR%2 ? " " : RS} 1'
gives (default output file is awkprof.out) {
ORS = (NR % 2 ? " " : RS)
}
1 {
print $0
}awk '! visited[$0] { print $0; visited[$0] = 1 }'
awk '!($0 in seen); {seen[$0]}''awk' is really a very beautiful little language. It's concise enough to solve many tasks in a single line, making it easy to use interactively while still being able to grow to moderately-sized scripts. It's not supposed to replace a full-blown scripting language like Python, but for processing files line-by-line it's superb.
The portable one-liner that doesn't suffer from integer wraparound is actually
awk '!($0 in seen) { seen[$0]; print }'
which can be golfed a bit: awk '!($0 in s); s[$0]'
$0 in s tests whether the line exists in the s[] assoc array. We negate that, so we print if it doesn't exist.Then we unconditionally execute s[$0]. This has an undefined value that behaves like Boolean false. In awk if we mention an array location, it materializes, so this has the effect that "$0 in s" is now true, though s[$0] continues to have an undefined value.
At least on the stock MacOS awk, you can get up to 2^53 before arithmetic breaks (doesn't wrap, just doesn't go up any more which means the one-liner still works.)
> echo '2^53-1' | bc
9007199254740991
> seq 1 10 | awk 'BEGIN{a[123]=9007199254740991;b=a[123]}{a[123]++}END{print a[123],b,a[123]-b}'
9007199254740992 9007199254740991 1
Even with one character per line, you'd need an 18PB file before you got to this limit, afaict.You might find this [2] helpful (oops, seems like it got deleted, see [3] - thanks @bionoid)
[1] https://www.gnu.org/software/gawk/manual/gawk.html
[2] https://www.reddit.com/r/awk/comments/4omosp/differences_bet...
https://www.removeddit.com/r/awk/comments/4omosp/differences...
Archive for posterity: http://archive.is/btGky
Sure, try this:
echo 1 2 | awk '{ print gensub(/1/, "3", "g", $1); }'
The logical thing for them to do would be to mention in bold and/or big and/or red font under gensub's documentation that it's an extension (e.g. try nawk), whereas looking through it I don't see any mention at all: https://www.gnu.org/software/gawk/manual/html_node/String-Fu...If I may rant about this for a bit, GNU software manuals are generally rather awful (though they're neither alone in this nor is it impossible to find exceptions). They frequently make absolutely zero effort to display important information more prominently and unimportant information less so (if you're even lucky enough that they tell you the important information in the first place). Like if passing --food will accidentally blow up a nuke in your hometown, you can expect that if they documented it at all, they just casually buried it in the middle of some random paragraph. Their operating assumption seems to be that if you can't be bothered to spend the next 4 hours reading a novel before writing your one-liner then it's just obviously your fault for sucking so much.
> Those functions that are specific to gawk are marked with a pound sign (‘#’). They are not available in compatibility mode (see section Command-Line Options)
But yes, sometimes POSIX or the C standard or whatever is too restrictive, but it's still a good starting point for figuring out how to write portable code.
You'll know anything it doesn't cover is implementation-specific, and then either decide it's not worth it to pursue it, or if it is figure out whether the implementations you're targeting support the feature.
I have twenty years of experience in getting that page to show up in search results. :)
Currently, a good way to get to it is these search terms:
posix issue 7
that actually takes us to the newer version; the above is issue 6.Haha! I love that you acknowledge this because usually people just ignore all the experience they have in getting to the right page and make it look like you're dumb for not being able to find it. Thanks for the pointer! :-)
http://pubs.opengroup.org/onlinepubs/7908799/
Ha, that didn't use frames yet! Totally forgot about that.
The same thing stands for GNU coreutils (`brew install coreutils`); here's macOS `cut` vs GNU `cut` as a quick example.
~ $ cut --help
cut: illegal option -- -
usage: cut -b list [-n] [file ...]
cut -c list [file ...]
cut -f list [-s] [-d delim] [file ...]
~ $ gcut --help
Usage: gcut OPTION... [FILE]...
Print selected parts of lines from each FILE to standard output.
With no FILE, or when FILE is -, read standard input.
Mandatory arguments to long options are mandatory for short options too.
-b, --bytes=LIST select only these bytes
-c, --characters=LIST select only these characters
-d, --delimiter=DELIM use DELIM instead of TAB for field delimiter
-f, --fields=LIST select only these fields...
-n (ignored)
--complement complement the set of selected bytes, characters or fields
-s, --only-delimited do not print lines not containing delimiters
--output-delimiter=STRING use STRING as the output delimiter
the default is to use the input delimiter
-z, --zero-terminated line delimiter is NUL, not newline
--help display this help and exit
--version output version information and exitTo be fair though, every time I have read the manual for a gawk function it clearly says "this is a gawk extension" for non-standard implementations (case in point, delete[1], which is now POSIX, although Mac's awk is too old to have that implemented).
[1] https://www.gnu.org/software/gawk/manual/html_node/Delete.ht...
Once you realise that awk's model is 'match pattern { do actions; }' everything makes a whole lot more sense.
BEGIN might be used to initialise Awk variables or print initial messages, while END can be used to print a summary of actions at the end.
There was a lot of push back - and this article is a good example of why
If I wanted to remove duplicate lines from a file I would almost certainly not use awk
I have never spent the time to get good enough with the whole new and different languge of awk, and am unlikely to need to (my large scale file processing needs seem small, and if I do it's almost always in context of other processing chains - so a normal languge like python would be the natural choice
I could whip up something like this in python in a less time than it would take to google the answer, read up why the syntax works that way and verify I have not mistyped anything on a few test files.
Basically using awk takes me out of my comfort zone - for a one off task it loses me time, for a production like repeat task I am going to reach for a slew of other solutions.
I mean the title of this page loses the exclamation mark - and it took me two goes to spot it.
Just like Python, there are users who use cli and are comfortable using grep/sed/awk/sort/etc
It's much faster to type out the awk line then to write the same in python.
Ok....
> It's much faster to type out the awk line then to write the same in python.
Is there some sort of speed-typing award that's being handed out that I'm missing? If there isn't, why would they feel superior?
We're all† smug pricks, but that's no cause for celebration. And 99% of the time we're not even justified in our smugness.
† All = a huge chunk of IT people, developers especially.
I know both Python and Awk. Do I go around telling people "stop using your preferred tool, even though it's efficient enough and works fine, use this other esoteric one instead"? Hell no.
This has not been my experience.
One-off one liners dashed off without syntax errors speaks of long and deep usage of a command line tool. That's cool. But continuing to use those one liners worries me for reasons not to do with skill
I would worry about the manual versus automation being used here. I can think of many cases where a sed/awk solution will work really well - but they almost always will be part of a larger developed and supported pipeline.
But if you using the one liner for anything not trivial you are still doing too much manual work
trying to be even shorter - if awk is your tool great! But ... at some point (and that point is much closer today than previously) anything we do needs a suite of tools we have hacked together and rewritten and passed around - from log file analysis to whatever.
And while awk can absolutely play a role in those tools, I doubt very much that anyone is good enough to make the one liners on the fly.
An quick example might be "show me all the logs for the request sent by user X in the last five minutes off the front web servers but ignore the heartbeat from that app marketing put out and ..."
I want that in my path, alongside everything else I and others working on the systems think useful.
Yes hack together your tools with any language you like. Put them in a seperate repo with all the linting turned off
But don't try and one liner them from scratch.
I was just pushing back against the sentiment that awk is undesirable because of attitutes that its users may have, which I don't think you were expressing :)
One day, nearly 20 years after it was published, I picked up a used copy of The AWK Programming Language by Aho, Kernighan, and Weinberger. Yes, they are credited in that order on the cover... I suspect intentionally. I only read the first N chapters, but it was enough. I used AWK many times within the following month, and I continue to use AWK on a daily basis. When the task is complicated, I will still use ruby, but often enough AWK is easier.
The point: you think "Why would I learn X when I can use Y?", but you won't really know the answer until you learn X. If I had never learned perl, python, ruby, AWK, shell script, vi macros, then I would probably be editing files by hand (!) like I sometimes catch developers actually doing (!!!). For a person who doesn't know these tools, that might actually be the path of least resistance. Investing some time here and there to learn new tools pays off in the future in ways that are unpredictable.
seen = set()
with open(filename, "r") as file:
for line in file:
if line not in seen:
print(line)
seen.add(line)
Often (at least in my experience) this kind of operation is either (a) part of some larger automated data processing pipeline for which it’s really nice to have version control, tests, ... or (b) part of some interactive data exploration by a programmer sitting at a repl somewhere, not just a one-off action we want to apply to one file from the command line.In those contexts, the Python (or Ruby or Clojure or whatever general-purpose programming language) version is easy to type out more-or-less bug-free from memory, debug when it fails, slot into the rest of the project, modify as part of a team with varied experience, etc. etc.
seen.add(line)
can be changed to seen.add(hash(line))
which can be significantly more memory efficient for files with long lines.This could involve saving previously seen lines in a radix tree, adding multiple layers of caching, saving infrequently seen lines to disk or over the network, etc. as appropriate for the use case.
I've been replacing some ad-hoc bash scripts (nothing fancy, just a few if conditions and some formatting of outputs for a deployment) with some AWK, and it's so much handier to write (after 10 years I still can't remember if syntax) and read (it's a proper programming language) than bash
edit: wrong markdown style
-n adds an implicit "while gets ... end" loop. "-p" does the same but prints the contents of $_ at the end. "-e" lets you put an expression on the command line. "-F" specified the field separator like for awk. "-a" turns on auto-split mode when you use it with -n or -p, which basically adds an implicit "$F = $_.split to the while gets .. end loops.
So "ruby -[p or n]a -F[some separator] -e ' [expression gets run once every loop]'" is good for tasks that are suitable for "awk-like" processing but where you may need access to other functionality than what awk provides..
I have a collection for ruby one-liners too [1]
[1] https://github.com/learnbyexample/Command-line-text-processi...
https://en.wikipedia.org/wiki/Ruby_(programming_language)
https://en.wikipedia.org/wiki/Perl
See "Influenced by" sections at both above pages.
Just the other day I helped some colleagues clean up a text file using an awk one-liner. It seemed like magic to them.
(Even though the one-liner turned out to be a bit more difficult to write than I thought at first due to '\r' characters in the input file)
Although you solved it, another way is that one can always pipe the input through a filter like dos2unix first. Very easy to write and versions can be found/written for/in many languages. Essentially, you just have to read each character from stdin and write it to stdout, unless it is a '\r', a.k.a. Carriage Return a.k.a. ASCII character 13, in which case you don't write it.
I've often found that beginners these days don't know what carriage return, line feed, etc. are, and their ASCII codes. Basic but important stuff for text processing.
https://metacpan.org/pod/distribution/App-nauniq/script/naun...
edit: https://git.savannah.gnu.org/cgit/gawk.git/tree/interpret.h
"The One True AWK" from Brian Kernighan (that is still the system AWK in OpenBSD) switched from a yacc implementation to a custom parser sometime within the last decade (fairly recently).
Busybox also has an awk; I'm not sure what they do.
GNU awk is elsewhere reported to be interpreted bytecode.
> In some ways the interesting thing is that the parser (probably for B, couldn't have been C based on radiocarbon dating evidence) was tiny and dead simple using recursive descent for most parts, a precedence table for expressions. But out of the intellectual culture-meets-culture encounter, an enduring tool was created.
If you're interested, you can read more about how GoAWK works and performs here: https://benhoyt.com/writings/goawk/
perl -nle 'print unless exists $h{$_};$h{$_}++' < your_file perl -ne 'print if !$seen{$_}++' perl -pne '$_=$#$_++?$_:""'
I'm rusty at this but shaved off six chars, five if you count the 'p' added to switches.the shortest I've got so far is
perl -lnE'say if!++$#$_'
---I don't understand what's happening with $#$_ but seems like something I should look into, thanks :)
you could remove n as p is used and would be same no. of characters as
perl -ne 'print if $#$_++'
you could save one more by removing space between e switch and single quote perl -ne 'print if!++$#$_'
seems to work alsoUsing a variable instead of 'foo' is a symbolic reference, so this is effectively using the symbol table as the associative array. This means that this solution also gets it wrong if your file contains a line that matches the name of a built-in variable in perl. That would be tough to debug!
If your file contains
This is the first line of the file
then during execution of ++$#$_
the result is the same as if you had written ++$#{This is the first line of the file}
So the variable @{This is the first line of the file} goes from undefined to an array of length 1, turning $#{This is the first line of the file} to 0.Incidentally, this is why the snippet fails to work for a line repeated more than once: for each occurrence of the expression, the value returned is in the sequence -1, 0, 1, 2, 3, ... so it is only false for the second occurrence.
Using preincrement instead of postincrement means the values returned are 0, 1, 2, 3, ... which means that inverting the test makes it false for every occurrence after the first.
A complex AWK script is very, very easy to move to a platform that lacks any AWK parsers. It can easily be done on Windows, without administrative rights, by the placement of a single .EXE - Perl can do many things, but that is not one of them (to the best of my knowledge).
Unix: No, that JSON data is too structured. But if you have a more error-prone format like CSV I can show you a neat trick to filter your bowlers by number of spares.
I'm not saying that there isn't a way to do that. Only that it can only be done poorly with a big ugly (and probably buggy) spaghetti script that looks nothing like what the expressive demo suggests it should look like.
Append .patch to the end of a pr or commit and it spits out the mbox formatted patch.
https://github.com/jiphex/mbox/commit/f139c575e306a1691a31d8...
I assume this wasn't always the case as the use case I'm referencing is a build script I'm debugging.
For a private repo you just need to set the correct options to curl.
curl -Lk --cookie "user_session=your_session_cookie_here" https://github.com/your_org/your_project/pull/123.patch
Found an example here.
These are only used when the allowed memory buffer is exhausted.
IME this one-liner can churn through 100MB of log lines in a second. Other solutions like powershell's "select-object -unique" totally choke on the file.
https://www.gnu.org/software/gawk/manual/html_node/Array-Int...
I suspect that it's designed so that hash collisions are impossible until you get to an unrealistic number of characters per line.
I'm not claiming that it will silently break! I'd be very interested in exploring the internals a little more and finding out how hard it is to get a collision in various implementations and how they behave subsequently.
EDIT: I've read chasil's comment and agree that it must be storing raw keys in the array. I guess awk uses separate chaining or something to get around hash collisions.
awk '{a[$0]++}; END{for(b in a) print b, a[b]}'
...will print every unique line in a file with the count. Obviously, that could not be done if the array index was a hash - the array index is the entire line, and the array value is the count.
The original program moves the maintenance of the array into the implicit conditional "pattern," and only prints when the array entry does not yet exist.
You might be better off switching to another scripting language that has some database API for storing key/value pairs on disk.
Or: use a 64 bit machine for a bigger address space, and add temporary swap files so you have more virtual memory.
There are people who are still on 32-bit hardware for serious work?
(Even if I was still on that, I'd probably just fire up a RV64 virtual machine (with swap space added within the VM, of course) simply to access the convenience of a larger address space when needed.)
These also have awk, almost always via Busybox, though using OpenWRT other versions are installable.
BusyBox v1.28.4 () built-in shell (ash)
_______ ________ __
| |.-----.-----.-----.| | | |.----.| |_
| - || _ | -__| || | | || _|| _|
|_______|| __|_____|__|__||________||__| |____|
|__| W I R E L E S S F R E E D O M
-----------------------------------------------------
OpenWrt 18.06.2, r7676-cddd7b4c77
-----------------------------------------------------
root@modem:~# uname -a
Linux modem 4.9.152 #0 SMP Wed Jan 30 12:21:02 2019 mips GNU/Linux
root@modem:~# free
total used free shared buffers cached
Mem: 59136 38484 20652 1312 2520 13096
-/+ buffers/cache: 22868 36268
Swap: 0 0 0
root@modem:~# df
Filesystem 1K-blocks Used Available Use% Mounted on
/dev/root 2560 2560 0 100% /rom
tmpfs 29568 1268 28300 4% /tmp
tmpfs 29568 44 29524 0% /tmp/root
tmpfs 512 0 512 0% /dev
/dev/mtdblock5 3520 1772 1748 50% /overlay
overlayfs:/overlay 3520 1772 1748 50% /
root@modem:~# which awk
/usr/bin/awk
root@modem:~# ls -l `which awk`
lrwxrwxrwx 1 root root 17 Jan 30 12:21 /usr/bin/awk -> ../../bin/busybox
root@modem:~# opkg list | grep awk
gawk - 4.2.0-2 - GNU awk
root@modem:~#
... though in this case, adblock runs on a larger and more capable Turris Omnia (8 GB flash, 2 GB RAM). The hourly sort still shows up on system load average plots.Will the BusyBox version of sort use files when there isn't enough RAM, will you have a big enough read/write flash partition for that?
The flexibility of keeping the adblock processing self-contained, rather than processing this on another box and rigging an update mechanism, is appealing.
The 'huge" file is 6MB. That's not immense, but it taxes (overly constrained, IMO) typical SOH router resources.
The flexibility afforded, for pennies to a few dollars, of, say, > 1 GB storage and 500 MB RAM, is tremendous.
I'm not sure what Busybox's sort does, though on an earlier iteration on a mid-oughts Linksys WRT54g router running dd-wrt, sorting was infeasible. I hadn't tried the awk trick.
It did have a Busybox awk though, which proved useful.
Hint: You can edit it after you posted.
[0] `man awk` on mac
[1] online version https://www.mankier.com/1/nawk
[2] gawk's man page works great as a reference https://www.mankier.com/1/gawk
Also how do you apply that command to the next file using shell history?
awk '!n[$0]++' fileName | sponge fileName dedupe.awk <file >file.tmp && mv file.tmp file
Multifile versions vary, I'd prefer listing them out, alternatively you could read from a command output (ls, find, etc.) with a 'while read; do ... done' loop: for f in file1 file2 file3
do
dedupe.awk <$f > ${f}.tmp && mv ${f}.tmp $f
done
If you want to apply to specific files on an ad hoc basis, you could wrap the whole thing in a shell function with filename or list as a parameter.Or 'gawk -i' as suggested.
Properly using tempfile would also be an improvement.
menu edit -> select all
press 'escape' and 'x' keys together and write: delete-duplicate-lines
done
Only for ones not familiar with awk.
It would make a lot of sense after you understand how awk works (as the article explains).
Sois it memory intensive or not?
If the file is large, and mostly unique, then assume that a substantial portion of the file will be loaded into memory.
If this is larger than the amount of ram, then portions of the active array will be paged to the swap space, then will thrash the drive as each new line is read forcing a complete rescan of the array.
This is very handy for files that fit in available ram (and zram may help greatly), but it does not scale.
Since then, perl and php have also implemented associative arrays. All three can loop over the text index of such an array and produce the original value, which a (bijective) hash cannot do.
I needed to remove duplicates from a sequenced CSV file yesterday but couldn't figure out the flags for "remove duplicates, output sorted by field 1 asciily, 2 numerically, 3 numerically".
The AWK version worked perfectly.