Compiled with "gcc -pipe -Wall -O3 -fomit-frame-pointer -std=c99 -pthread" on my Mac, it's about twice as fast at the blog author's version.
Time spent: coding: 30 minutes bugfixing: 30 minutes
I have a feeling there is some kind of catch in the description of the algorithm in terms of implementing the output, but I for the life of me could not grok whether they wanted me to parse the entry into these three pieces or not...
EDIT to add: The only optimization I made was inlining the tightly-called "subst" function, did it without any profiling (so the optimization process literally took about 30 seconds:). Before inlining this version was still about 15% faster than the blog author's one.
1) My initial understanding was that I do need to reverse the order, yet somehow after re-reading the article I understood the order does not need to change, and the "reverse" in the name is some kind of jargon. This is quite stupid, and probably not worth mentioning, if only to prove I was tired :)
2) missing that the first iteration of the "business logic" code in my case happens before anything is filled in. Crash.
3) forgetting about the "\n"s - with rather funny "partially correct" output effect.
Very much looking forward to see your code !
Ok, I looked at your code. What you really should do (before coding up the solution) is to look at the problem specification. Other than that I like the 'direct' approach, it isn't quite as fast as what I cooked up but yours is a lot shorter.
Enjoyed reading your code. Beautiful. Thanks!