We did in fact start by creating a data structure that we parsed!
The main problem we were trying to solve is essentially that we had a lot of code like:
f(x) = sma(ewm(x))
g(x) = slope(ewm(x))
And we wanted to avoid recalculating ewm(x), while also not holding memory for it for longer than necessary. Our production use cases would have tens of thousands of intermediates, so the potential saving from getting this right is huge.
Naturally a DAG is a great fit for this type of problem. We started by creating an explicit DAG, where everything like ewm(x), sma(ewm(x)) would be a node in the DAG. We could then evaluate the DAG efficiently using multithreading, dropping nodes when everything that depended on them had been calculated, etc.
And I would usually stop there, but there were two main challenges. One was that the syntax was pretty clunky, and we wanted researchers (who weren't software engineers) to be able to use it. The other was that we wanted _any_ reference to something like ewm(x) to map to the node, even if they were used in very different places – and it was quite difficult to map these dependencies without explicitly tracking.
It's hard to quickly summarize our macro, but basically what it let us do was write the short syntax and track dependencies using the underlying Julia expressions. We mapped those expressions to nodes in the DAG, and that let us calculate the dependencies without having to track them explicitly. I don't think this would have been possible without some amount of homoiconicity and macros.