GoAWK: an AWK interpreter written in Go
github.com
github.com
There are some obvious questions around calling conventions and error handling, method invocation, etc. but nothing there seems totally insurmountable. Having a compliant implementation as a jumping off point is a great start.
Looking at the interp internals, the representation of function call expressions might need a little bit more structure to pull this off (rather than just a big switch for the awk builtins and user calls as just more awk instructions plopped inline). Furthermore there are questions about how exactly to represent go objects but I suspect with some boxing it could be made relatively ergonomic.
In my opinion /usr/bin/awk is a thing of beauty. Certainly it's the most usable out of the trifecta of scripting languages that are mandated by POSIX.
(There's /bin/sh, where merely using variables can quickly turn into a quoting nightmare. But the true nightmare material is /usr/bin/sed, which has actually been shown [1] to be a Turing complete language!)
[1] https://nixwindows.wordpress.com/2018/03/13/ed1-is-turing-co...
I had fun fuzzing the "real" awk, finding a couple of trivial segfaults. If you've not already experimented with fuzzing I'd recommend it - I found a few minor issues in my own simple-interpreter, and language, via feeding them malformed scripts.
For example, to extract and add line numbers to SQL table definitions:
t = gawk.Gawk(sys.stdin)
t.context.data = ''
@t.range(r'CREATE TABLE', r');')
def line(context, line):
context.data += (('line %d:' % context.range.line_number) + line)
if context.range.is_last_line:
print(context.data)
context.data = ''
t.run()
https://github.com/linsomniac/gawkI've used AWK for close to 30 years, but I've never achieved or maintained any level of proficiency at it. I pretty much just use it for "{ print $1, $3 }" in a filter or the like. Every time I try to do something more complicated I spend an hour or more futzing around with it and more often than not getting to: almost but not quite" where I want to be. This is, of course, a me failing not an awk failing.
But it's left me wanting something that would make doing awk-like processes easy in Python, which I'm very proficient at.
I ended up using the name "gawk" because it's an English word and nods to the AWK inspiration, but then I remembered GnuAWK so I'll probably rename it.
Your code snippet does showcase Awk's utility. It brutally cuts through all the ceremony around reading and iterating over lines.
[1] https://github.com/pharmbio/ptp-project/blob/master/exp/2018...
https://en.wikipedia.org/wiki/Comma-separated_values
For example, csv.reader in Python's csv module in the stdlib, has a dialect argument, due to this.
Less intuitive maybe for beginners but more generally useful?
One thing hope this implementation remedies is the absence of a linear time string concatenation in awk. Awk has split but no join. Only way I know is to iteratively join two strings which has a quadratic running time.
I haven't looked at linear-time string concat. Interesting point -- I'll put it on my TODO list. Though I think instead of string building you could simply use printf to write output and that would be linear time.
function join(ARRAY) {
for (i=0; i in ARRAY; i++) printf "%s", ARRAY[i];
}
But if you need the result back as a string for further processing, the obvious methods are not linear: function join(ARRAY,_s) {
for (i=0; i in ARRAY; i++) _s=_s ARRAY[i];
return _s;
}
# try to be clever and join 2 items at the same time
function join2(ARRAY,_s) {
for (i=0; i in ARRAY; i+=2)
_s=sprintf("%s%s%s",_s,ARRAY[i],ARRAY[i+1]);
return _s;
}
I for one definitely miss having a join() function, and it seems odd that this natural complement to the split() function was never implemented...Regarding concat, printf and sprintf would indeed work in some situations. Its not flexible enough in the following scenario:
Consider an array of substrings, possibly generated by 'split'. I modify and delete some of them. Then I want to put them back together. Since i do not know ahead of time how many substrings will survive, sprintf is difficult to use.
I'm curious though, why code the lexer and parser by hand? What's the state of lexing/parsing in the Go world?
As to the state of lexing/parsing in the Go world. There's a simple scanner (text/scanner) in the stdlib. I've run across this quite neat parser library that's based off structs and tags: https://github.com/alecthomas/participle ... but I really don't know the landscape very well.