Experience porting 4k lines of C code to go
blog.kowalczyk.info
blog.kowalczyk.info
That's an excellent quote. And it describes a lot of C code I've seen (good and bad).
Porting a program written in one of the "lowest level" languages still in use today to one of the newest "modern" languages which claims high level features and having only 20% code size reduction doesn't seem like a big win to me, especially if all the saving attribute to only one feature (better arrays).
Maybe it is just because it was port, and not a reimplementation in idiomatic Go style.
I would certainly like to see other solutions to the same problem written in idiomatic Go, Clojure, Haskell, C++ and Scala, just for the comparison.
There are two things at work here: how small is the parser? And second, how fast is the parser? I know you can write some extremely fast highly idiomatic parsers in Haskell though. But I doubt they will be faster than hand coded C.
BTW: if people want to do markdown in Go, the library to use is https://github.com/russross/blackfriday. It's a different one that the code I ported but Russ Ross did exactly the same thing at almost exactly the same time and was a little bit ahead, so after I discovered his work, I decided to contribute to his code instead of maintaining my own, almost identical, project.
That's not an easy question to answer.
If you were to do a faithful port i.e. using using the same techniques as C/Go code, it would be very close. upskirt uses a traditional lexing/top-down-parsing approach. There's nothing in Python syntax that makes writing such code more compact than in Go. It's a bit tedious to write but the benefit is that (in C/Go) it gives the best speed because it minimizes the number of times each source character is looked at.
If you were to use a different approach e.g. brute-forcing your way through the text several times with regexpes, which is the most popular way of doing that in dynamic languages, the code clocks at ~2 thousand lines (like the implementation I use for my home-grown blog system https://github.com/kjk/web-blog/blob/master/markdown2.py)
If you were to use this technique in Go, the code would probably end up smaller than the manual lexing/parsing approach used in upskirt, but it would be significantly slower.
Interestingly, upskirt approach would probably be slower in Python than the regexp approach, because regexpes are heavily optimized C code and looking at individual characters of the string in Python code isn't particularly fast due to python interpretation overhead.
AFAIK, the python re module does not have a C language component. It is a pure python regular expression implementation.
V8 is a notable exception. They use their code generation pipeline to JIT compile regexps to machine code, making it the fastest regexp engine around bar none (I think).
Libraries don't appear in the Comp.Lang.Benchmark Shootout, so CL doesn't do so well.
Of course, the downside is that formal grammars are often a bad fit for custom languages of the web, which tend to have incomplete/undecideable/context-dependent grammars. This would work for Markdown, though, since it does have a parseable grammar.
Using a real parser could probably cut half the code. Been meaning to make my own Markdown parser, everyone else's kind of sucks :-)
One, this was an opportunity to learn Go.
Two, writing a converter is way above my head and quite likely theoretically (and practically) impossible. I've definitely had to make some decisions that I don't think even the cleverest compiler could (like noticing that array.[c|h] and buffer.[c|h] could be replaced with native Go arrays/slices easily so I didn't have to port that at all, just change the callers to Go equivalents).
If you are porting direct C logic to Go you could be missing a lot of potential idioms that could differentiate Go from C or python further. You said yourself it was easier to implement parsing with slices.
I think a more effective task is "understand what the 4k lines of C do, and then throw it away and re-write it as if it was written in Go in the first place." Your miles may vary.