Lisp works slightly different. The macro form is an expression (a list) with the macro operator as the first element. The next elements in the expression can be anything as data.
Judging by some of the comments here, it seems like the macro system has a similar approach as Lisp's macro system, which is also AST-based. Something I don't see here is macros that generate other macros, but the question is how much you really want that anyway (when I did that, I thought the syntax was horribly complicated because of all the quoting). I know Lisp also has reader macros (they run before the parser) that allow you to effectively change the language syntax, but I didn't use those.
You can do that in Nim in a readable way.
import macros
macro genMacro(name: untyped): untyped =
result = quote do:
macro `name`: untyped =
result = quote do:
echo "Foo"
genMacro(bar)
bar # Generate and perform the echo
Surprisingly I've actually used this kind of thing!In one of my projects I use a macro to parse a set of types for fields and generate constructor macros for them.
The generated constructor macro passes through the parameters it's given to the default built-in constructor but does some setup before.
The final generated code is a normal built-in construction without proc calling yet with special fields initialised automatically.
Funny, I did a similar thing! :)
For example this is a valid Lisp macro form
(loop for i below 10 and j downfrom 20
when (= (+ i j) 9) sum i into isum of-type integer
finally (return (+ j isum)))
It returns 10.
The LOOP macro parses it on its own - it just sees a list of data. Lisp has other than that no idea what the tokens mean and what syntax the LOOP macro accepts. CL-USER 142 > (defmacro my-macro (&rest stuff)
t)
MY-MACRO
CL-USER 143 > (my-macro we are going to Europe and have a good time)
T
Works. CONS
/ \
FIXNUM CONS
| / \
1 CONS SYMBOL
/ \ \
FIXNUM CONS NIL
| / \
2 FIXNUM SYMBOL
| |
3 NIL(not (eq 'has 'is))
But for an Abstract Syntax Tree for code we have more categories: function, operator, call, control structure, variable, class, ...
If our code walker dispatches on pattern matches on the nested list structure of conses and atoms, we don't need that sort of encapsulated data structuring. The shapes of the patterns are de facto the higher level AST nodes. The code walking pattern case that recognizes (if test [then [else]]) is in fact working with an "if node". That AST node isn't defined in a rigid structure in the underlying data, but it's defined in the pattern that is applied to it which imposes a schema on the data.
If that's not an AST node, that's like saying that (1 95000) isn't an employee record; only #S(employee id 1 salary 95000) is an employee record because it has a proper type with a name, and named fields, whereas (1 95000) "could be anything".
The code walker is just another parser. It needs to know the Lisp syntax. It needs to know which parts of a LET form is a binding list, what a binding list looks like, it needs to know where declarations are and where the code body ist. It can then recognize variables, declarations, calls, literal data, etc, It needs to know the scope of the variables etc. Nothing of that is encoded in the LET form (since it is no AST), and needs to be determined by the code walker. Actually that's one of the most important uses: finding out what things are in the source code. Lisp does not offer us that information. That's why we need an additional tool. A code walker may or may not construct an AST.
No, (1 950000) is not an employee record. Only your interpretation makes it one. Other than that our default Lisp interpretation based on s-expressions: it's a list of two numbers. In terms of the machine it's cons cells, numbers, nil. Without further context, it has no further meaning.
A Lisp function call does "parsing". (foo 1 2 3) has to figure out dynamically whether a (lambda (&rest args)) is being called or (lambda (a b c)) or (lambda (a b &optional c (d 42)) or whatever.
The #S(employee id 1 salary 95000) object also isn't an employee record without context and interpretation.
> it needs to know where declarations are and where the code body ist
The syntax can be subject to a fairly trivial canonicalizing pass, after which all these things are at fixed positions:
(let ((a 3) b) (foo a)) ---canon--> (let ((a 3) (b)) (declare) (foo a))
Now the variables are all pairs to which we can blindly apply car and cadr, the declarations are at caddr and the body forms at cdddr.No. You are still operating on the level of s-expressions, a data format. The type-tag of LET is SYMBOL.
Here we have some Lisp code in the form of an s-expression:
(let ((let 'let))
((lambda (let)
(let ((let let))
let))
let))
All above LET have the same type tag, but in terms of syntax they have a different purpose in the form above: we have special operators, variable declarations, variable usage, data objects. I can't just car/cdr down the lists and call TYPE-OF. This always returns SYMBOL for LET.On the level of a syntax tree we would want to know what it is in terms of syntactic categories: variable, operator, data object, function, macro, etc. Lisp source code has no representation for that and we need to determine that by parsing the code.
Essentially there's a VM that runs almost all the language barring importc type stuff, and you can chuck around AST node objects to create code, so metaprogramming is done in the core Nim language. You can read files so it's easy to slurp files and use them to generate code or other data processing at compile-time.
Several simple utility operations in the stdlib make things really fluid; the easy ability to `quote` blocks of code to AST, and outputting nodes and code to string. This lets you both hack something together quickly and learn over time how the syntax trees work.
Quoting looks like this:
macro repeat(count: int, code: untyped): untyped =
quote do:
for i in 0..<`count`:
`code`
repeat(10):
echo "Hello"
Inspecting something's AST can be done with dumpTree: dumpTree:
let
x = 1
y = 2
echo "Hello ", x + y
To save even more effort, there's also dumpAstGen which outputs the code to create that node manually, and even a dumpLisp!You can display ASTs inside macros:
macro showMe(input: untyped): untyped =
echo "Input as written:", input.repr
echo "Input as AST tree:", input.treerepr
result = quote: `input` + `input`
echo result.repr
echo result.treerepr
So it's really easy to debug what went wrong if you're generating lots of code.Since untyped parameters to a macro don't have to be valid Nim code (though they still follow syntax rules) you can make your own DSLs really easily and reliably because any input is pre-parsed into a nice tree for you.
Here's a contrived example of some simple DSL that lets you call procs and store their results for later output:
import macros, tables
macro process(items: untyped): untyped =
result = newStmtList()
# Create hash table of 'perform' names to store their result variables.
var performers: Table[string, NimNode]
for item in items:
let
command = item[0]
param = item[1]
paramStr = $param
case $command
of "perform":
# Check if we've already generated a var for holding the return value.
var node = performers.getOrDefault(paramStr)
if node == nil:
# Generate a variable name to store the performer result in.
# genSym guarantees a unique name.
node = genSym(nskVar, paramStr)
performers.add(paramStr, node)
# Add the variable declaration
result.add(quote do:
var `node` = `param`()
)
else:
# A repeat performance, we don't need to declare the variable and can overwrite the
# value in the fetched variable.
result.add(quote do:
`node` = `param`()
)
of "output":
let node = performers.getOrDefault(paramStr)
if node == nil: quit "Cannot find performer " & paramStr
result.add(quote do:
echo `node`)
else: discard
# Display the resultant code.
echo result.repr
proc foo: string = "foo!"
proc bar: string = "bar!"
process:
perform foo
perform bar
perform foo
output foo
output bar
The generated output from process looks like: var foo262819 = foo()
var bar262821 = bar()
foo262819 = foo()
echo foo262819
echo bar262821