The reason I chose to implement it that way in the article is that I modeled RLMeta after OMeta. If I remember correctly, their reason for doing so was that they thought compilers would be easier to write if all passes (parsing, code generation, ...) could be written using the same language.
I think using the same DSL for code generators works quite well and they become quite readable. Perhaps there is a better DSL suited for code generation, and it would be interesting to experiment with that as well.
You are right that the toy examples could generate code directly in the parser. There is no need for a separate code generator step. But the RLMeta compiler would be much less readable without a separate code generator step I think. The reason I introduced an AST in the toy example was to introduce the concept before I showed it in the RLMeta implementation.