Lisp: It's Not About Macros, It's About Read
jlongster.com
jlongster.com
> Wait a second, if you take a look at Lisp code, it’s really all made up of lists:
Haskell has `read`, most of your data types can just derive Read and Show and they'll "magically" get a representation allowing you to `read` and `show` them.
But that only works for datatypes, you can't do that with code.
In Lisp you trivially can, because code is represented via basic datatypes (through a reader macro if needed).
It's not macros. It's not read. It's much more basic than that: it's homoiconicity. From that everything falls out, and without that you need some sort of separate, special-purpose preprocessor (whether it's a shitty textual transformer — as in C — or a more advanced and structured one — as with Camlp4 or Template Haskell — does not matter).
And I don't get why the author got the gist (and title) so wrong when he himself notes read is irrelevant and useless in and of itself:
> Think of read like JSON.parse in Javascript. Except since Javascript code isn’t the same as data, you can’t parse any code
`read` does not matter if the language isn't homoiconic, you can `read` all you want it won't give you anything.
Focusing on `read` was a way to anchor my article, even if it truly isn't about read either. I tried to tie that together at the end.
A string has no additional structure, so if you want to do any transformations beyond simple string/regex substitutions you have to parse it into a more suitable format.
Yes, Lisp's power comes from the embodiment of code and data together in one manner, and the ability to treat them this way when writing code is good, but `read` is a coincidence of that power, not a demonstration of how it is used. Macros are the method by which we harness the power of homoiconicity in an efficient, powerful manner.
It's rarely the case that I want a single representation for all my data -- and if we treat code as data, do I want a representation that is indistinguisable from all my other data?
For example, the distinction between data that specifies layout (html, xaml, etc...) and that which performs logically computation (javascript, c#, etc...) seems like a useful distinction to have.
While I can appreciate the AST form of s-exprs I also do like the richness of many standard languages -- and the semantic richness of their ASTs.
Lastly, treating code as data (and vice-versa) has been the bane of many programmers of days past. Go back 40 years and you can find many developers who did treat code as data (it was all actually viewed as sequence of bits by many) and this caused no end of problems. In most modern systems there are often safeguards to specify data and code segments and ensure that you don't treat one as the other. While not completely analogous to Lisp macros, it does show that you tread dangerous ground when you attempt to treat all forms of data as indistinguishable.
Given the special purpose nature of code, I don't mind (and actually appreciate) a well thought through syntax, and a special set of functionality to interact with it -- as I do most special purpose forms of data.
(curious about how you use the language, not interested in scoring points)
Having a uniform representation for data isn't a big enough win to trump having data represented in a way that is more natural for me to think about.
With that said, if your brain thinks in XML (or sexprs) maybe Lisp will always work best for you.
Do you handle all varieties of lambda lists? recognize and descend into all the special forms? what do you do with macros? expand them (a mess)? try to walk into special-cased standard ones like 'loop' and ignore user-defined ones?
The closest you can come to doing it sanely is to use a code-walking library like the one in arnesi: http://common-lisp.net/project/bese/docs/arnesi/html/A_0020C...
`read` does not perform macro-expansion: that would break data reading and the reading of quoted forms. macroexpand expands macros at a later stage. Once expanded, macros either refer to primitive special forms or function calls, and it's trivial to determine which. Primitive special forms can have their components macroexpand'd as appropriate. Function calls can be left as-is to be compiled (well, the arguments can be macroexpand'd). Eventually there will just be primitive special forms and raw function calls left, ready to be handed to the compiler.
In idiomatic Lisp code, you either miss lots of the calls, or you have to complicate your code-walker significantly. This is especially the case if you write CL in the (common, but not universal) style that makes significant use of the loop macro, because you either ignore it as a macro, and consider anything inside it opaque until macro-resolution time (because you don't know what it does to its arguments), or you special-case it as a new bit of CL syntax, in which case your parsing is now fancier. Usually you want something like the latter, because source-to-source transformations expect to also replace things inside loops. Same with, say, special-casing setf forms, if you want source-to-source transformations to "do what I mean" in a large number of cases.
It's true that it's very easy to literally get the list representing the code, but there's precious little sensible you can do with that list unless you're willing to descend into some of the more commonly used built-in macros that most CLers treat as de-facto syntax, which requires knowing something about the syntax they in effect define.
For your specific example of replacing calls to foo with another bit of code you may be able to get away with macrolet. (Example: http://letoverlambda.com/index.cl/guest/chap5.html#sec_4 )
A code walker then is mildly complex:
http://www.cs.cmu.edu/afs/cs/project/ai-repository/ai/lang/lisp/code/codewalk/walk/new_walk.cl (read-from-string "#.(print :foo)")The linked pastebin isn't a macro system. It's merely a macroexpansion system, it needs to be evaluated. And it's not as simple as merely wrapping it in 'eval' because of subtleties in getting at the right lexical scope.
More generally, no fair claiming macros are easy because you managed to build them atop a lisp. You're using all the things other comments here refer to; claiming it's all 'read' is disingenuous.
I'm still[1] looking for a straightforward non-lisp implementation of real macros. The clearest I've been able to come up with is an fexpr-based interpreter: http://github.com/akkartik/wart
[1] From nearly 2 years ago: http://news.ycombinator.com/item?id=1468345
With the interesting critique that "objects" are better than s-expressions for representing sourcecode. (BTW, Moon did a lot of work on Lisp.)
I've too thought that s-expressions don't necessarily contain as much information as you'd want. Using Rich Hickey's word from "Simple Made Easy", maybe they're used to "complect" visual presentation and internal representation.
Then again, there's metadata...
Neither side can understand the other: one side says "why do you resist ultimate power?" and the other side says "how can you possibly think that your code is readable?"
My belief (and what I am starting to consider my life's work) is that the gap can be bridged. Lisp's power comes from treating code as data. But all code becomes data eventually; turning code into data is exactly what parsers do, and every language has a parser. The author says "it's about read," but "read" (in his example) is just a parser.
The author asks "How would you do that in Python?" The answer is that it would be something like this:
import ast
class MyTransformer(ast.NodeTransformer):
pass # Implement transformation logic here.
node = MyTransformer().visit(ast.parse("x = 1"))
print ast.dump(node)
This works alright, but what I'm after is a more universal solution. With syntax trees there's a lot of support functionality you frequently want: a way to specify the schema of the tree, convenient serialization/deserialization, and ideally a solution that is not specific to any one programming language.My answer to this question might surprise some people, but after spending a lot of time thinking about this problem, I'm quite convinced of it. The answer is Protocol Buffers.
It's true that Protocol Buffers were originally designed for network messaging, but they turn out to be an incredibly solid foundation on which to build general-purpose solutions for specifying and manipulating trees of strongly-typed data without being tied to any one programing language. Just look at a system like http://scottmcpeak.com/elkhound/sources/ast/index.html that was specifically designed to store AST's and look how similar it is to .proto files.
(As an aside, programmers have spent the last 15 years or so attempting to use XML in this role of "generic language-independent tree structured serialization format," but it wasn't the right fit because most data is not markup. Protocol Buffers can deliver on everything people wanted XML to be).
Why should manipulating syntax trees require us to write in syntax trees? The answer is that it shouldn't, but this is non-obvious because of how inconvenient parsers currently are to use. One of my life's goals is to help change that. If you find this intriguing, please feel free to follow:
https://github.com/haberman/upb/wiki
https://github.com/haberman/gazelleI crave both the power of code-as-data and nice syntax, which is why I love Lisp.
I find Lisp (particularly Clojure) much more aesthetically pleasant, in that it communicates better with me. With Paredit, it's even better to the touch.
(If there one day came to exist something even better on these metrics, then I'm sure I would start to prefer it aesthetically.)
Doubly agreed! I learned Clojure just over a year ago and will never look back. My attitude to people complaining about parents is that that should just get over it. That one hang up is actually holding them back.
(defun substitute-in-replacement ($-value replacement) (cond ((null $-value) replacement) ((null replacement) ()) ((eq (car replacement) '$) (cons $-value (cdr replacement))) (T (cons (car replacement)) (substitute-in-replacement $-value (cdr replacement)))))
From: http://www.csc.villanova.edu/~dmatusze/resources/lisp/lisp-e... with one paren moved.
I remember, when I was taking a class on AI, looking for some sort of style guideline that would help me get through the learning curve, but the FAQ (I want to say it was comp.lang.lisp) just had "coming soon." So this would be the allegro editor in 2001 or 2002. It may be obvious to an experienced hand, and perhaps if there was some sort of best practices when I was learning it I wouldn't have had the same problem, but I just remember the frustration of my mind playing tricks on me and (even with syntax highlighting) trying to match parens that I thought were there.
I haven't read a lisp style guide, Emacs just takes care of indentation - it is immediately clear when a paren is wrong because the shape of the function is wrong. If you are writing lisp with an editor that doesn't do this, get a better editor, don't blame the language.
Arguably Python -- but to get that, Python sacrificed the possibility of both usable anonymous functions and the possibility to cut-paste a code fragment and just ask the editor to reindent.
Hardly worth the price.
It's quite possible my experience as a programmer today would be different than when I started -- I mean, I made it through a few chapters of SICP without such troubles, but in the back of my head was the memory of trying to figure out my logic error in a bit of code when it was really a misplaced paren.
Actually, it's pretty obvious even without doing all that. CONS always takes two arguments.
(defun substitute-in-replacement ($-value replacement)
(cond ((null $-value) replacement)
((null replacement) ())
((eq (car replacement) '$)
(cons $-value (cdr replacement)))
(T (cons (car replacement))
(substitute-in-replacement $-value (cdr replacement)))))
CL-USER 5 > (compile 'substitute-in-replacement)
;;;*** Warning in SUBSTITUTE-IN-REPLACEMENT: CONS is called with the wrong
;;; number of arguments: Got 1 wanted 2
SUBSTITUTE-IN-REPLACEMENT
Lisp compilers able to present these error messages are in use since more than 40 years. Common Lisp has them since day one. (defn thingies [id]
(->> id
fetch
read-json
:rows
(map :thingy)))
(Of course, my code is often more complex and messier than that, even when using ->>, but some fairly significant percentage of my code does look that simple.)I'm sure there's stuff to criticize about Clojure, but we can look at real-world code in another mainstream language (Javascript+node.js? PHP? Java?) and point out readability problems too. (Python maybe being an exception in terms of readability-in-the-small, for things that fit in the mainstream style. Though as someone pointed out, there's maybe some problems with manipulability.)
I could attempt to prove to you that "conventional syntax" is inherently superior to Lisp syntax, but that would be a waste of both of our time.
Yes, trying to prove falsehoods is a waste of time.
Conventional syntax is neither conventional nor suited to humans. (If it's "conventional", why isn't there more agreement as to what it is? If it's suited to humans, why aren't there more than 100 who actually know it for any given language?)
add 1 and 2 and 3
The syntax is a bit terse but if you teach people a good way to read it it becomes much more readable than 1+2+3
The only reason we prefer that way is that we are thought that syntax when we do math in school, I have found it much easier to teach people lisp who have no or very little formal education in math.
(/ (+ (- b) (sqrt (- (* b b) (* 4 a c)))) (* 2 a))
Yeah, so it divides (the addition of (-b and the (sqrt of (the difference between (the product of b and b) and (the product of 4, a and c))) by (the multiplication of 2 and a))Right, that's much easier than
(-b + sqrt(b*b - 4*a*c)) / (2*a)
(-b plus the sqrt of ((b times b) - (4 times a times c))) divided by (2 times a)I see you omitted some parenthesis in the "conventional" expression, relying on the fact that multiplication takes priority over substraction. Making this fact explicit is exactly what makes Lisp better, especially for more complex domains: delegating priorities to the notation, freeing brain capacity for the actual problem.
(defun foo (a b c)
#I(
(-b + sqrt(b*b - 4*a*c)) / (2*a)
))
CL-USER 8 > (foo 1 2 3)
#C(-1.0 1.4142135)It may be bikeshedding, but I would not let 'blue is better than red, because the sky is blue' pass either.
(/ (+ (- b)
(sqrt (- (* b b)
(* 4 a c))))
(* 2 a))
It tells me:* there's a quotient of 2 things
* the first thing is a sum of -b and a sqrt
* the second thing is a product
and so on. Pretty nice. Of course, mathematical notation is more terse.
You sound as if you think that putting salt on grapefruit is inherently strange, while in actuality there's a very good reason to do so: it reduces the perception of bitterness.
I'm not saying that it's impossible to like Lisp's syntax, but empirically most people prefer the ALGOL-like syntax
Empirically, most people prefer what they are already familiar with, so I'm not sure what this is supposed to prove, other than most people are already more familiar with Algol-like syntax.
For me, Lisp syntax has the definitive advantage that the first identifier in every expression tells me what to expect. I.e., I don't have to scan to the right to figure out what kind of expression this is. For me, this makes code much more readable. And this makes Lisp syntax more "nice".
Closer to "the human"? Do you know more than 3 people who know C++ operator precedence?
Humans don't handle operator precedence very well.
You don't have to know an entire operator precedence table to read and write idiomatic infix-notation code. Precedence is defined such that common expressions evaluate as people intuitively expect (a notable counterexample is "x & y == z" in C). Parentheses are always available to clarify more complicated expressions.
Come to think of it, humans usually add and subtract by stacking numbers vertically. I don't think you can point at infix notation as "the" human-friendly notation.
This feels like a discussion based in fiction...
(/ (+ (- b) (sqrt (- (* b b) (* 4 a c)))) (* 2 a)) ?
I take it you think there is something intrinsically wrong with that idea?
I find it very hard, without bracket counting, to see exactly what the '+' and '/' bind to. With the more traditional:
(-b + sqrt(bb-4ac)) / (2a)
I find in only a glance I can tell what everything is binding to.
It must be nice to live in a world with only 4 infix operators and expressions that have only 3 infix operators.
For example, lots of folks think that sqrt should be a prefix operator, not yet another function. I suppose you're going to assume that the top bar will serve as parentheses.
BTW "-b + sqrt(bb-4ac) / 2a" is the interesting expression. Is it "(-b + sqrt(bb-4ac)) / (2a)" or "-b + (sqrt(bb-4ac)) / 2a)" And, are you certain what "bb-4ac" means? (There's at least one major language where it doesn't mean "(bb)-(4ac)".)
And that's how the exceptions swallow the rule. And, it's also how we get infix programming languages where that's definitely not true, and so on. Where should we make the switch?
Also, only four? What about set operations?
(/ (- (sqrt (discriminant a b c)) b)
(* 2 a))Other notations are used, but with a frequency similar to pre-fix (lisp) and post-fix (forth). "Associativity" (not affected by order of evaluation) only makes sense for in-fix.
But it really could just be familiarity, I guess. I can't see how to determine it either way. But regardless of the cause, there's overwhelming evidence that people, in fact, prefer in-fix.
How many people have seen anything other than in-fix? Of those, how many got a fair shot at an alternative?
If it's "idiomatic", why is there such disagreement?
> Why then has math (which is read and written only by humans)
Convention has a lot of value. That said, mathematicians don't have to worry about getting things wrong. It's just paper, and they're happy to let humans fix up the errors.
> Parentheses are always available to clarify more complicated expressions.
Unnecessary parentheses are how humans deal with the fact that they can't handle infix.
here's IPL, an influence of lisp, also a list processing language (c/p from wikipedia)
IPL-V List Structure Example
Name SYMB LINK
L1 9-1 100
100 S4 101
101 S5 0
9-1 0 200
200 A1 201
201 V1 202
202 A2 203
203 V2 0
How human LISP feels now ;) ?But there are other factors which would make me happier with a language than closing the gap between expressive power and great syntax. For example, I would love if there were a language with nice syntax and good metaprogramming (eg python) that also had an unambiguous visual representation (something like eg Max) that you can switch between at will. Dunno how realistic that would be without adding complexity or ambiguity or ruining code formatting)
They were great for messaging . . . but we found ourselves using them /everywhere/. And since our stuff worked in many different environments (C++, Java, Visual Basic were the ones we directly supported), you could have your choice of language.
It's flattering to see this rediscovered, several times over :-)
Another way of putting it is that I'm trying to beat Greenspun's Tenth Rule by making that "half of Common Lisp" separable from Common Lisp so that C programs (and high-level programs too) don't have to keep re-inventing it. As a bonus, this will help make languages more interoperable too.
So what I'm saying complexity is significantly reduced when you have a small/consistent core. As for readability I think Clojure makes this better by providing different literals for vectors and maps and those literals have consistent meaning in similar situations so it provides nice visual cues. But immutability by default, clear scoping and functional programming make things like using significant whitespace and pretty conditional expression syntax bikesheding level details.
https://github.com/andrewf/fern
It's a very much a prototype, and my ideas have evolved a lot (towards lisp), but I'm at least curious what you think of the ideas in the README. #id > p a.red:visited {
background: url(foo.png) white;
margin: 0 3px 5em 80% ! important;
}
There's a lot going on here. CSS isn't just key/value maps.Also, I don't think you want a data language to be Turing-complete. PostScript was Turing-complete but PDF is not; this makes PDF easier to deal with because it's easier to analyze and there's no risk of it getting into an infinite loop.
I don't intend it to be just a data language. What I've been moving toward in my daydreams is a DAG of, for lack of a better word, function calls (some interesting data doesn't really fit in a tree), some of which are generators. If Turing-completeness is a problem in your context, you can reject some or all generators and/or just not evaluate them, i.e. take them as pure data. But I don't want to limit myself. I would have no problem if it turned into a general purpose language with a nice data-oriented subset.
There are two larger problems in adding Lisp-style macros to non-Lisp languages, one social and one technical.
The social problem is that language designers must be persuaded to publish a specification of the internal representation of the AST of their language. This makes the AST a public interface, one which they are committed to and can't easily change. People don't like to do this without a good reason.
The technical problem is more difficult, though. To make a non-Lisp language as extensible as Lisp would require making the parser itself extensible. This is not too hard to implement, but perhaps not so easy to use. If you've ever tried to add productions to a grammar written by someone else, you know it can be nontrivial. You have to understand the grammar before you can modify it.
And if you overcome the difficulties of having one user in isolation add productions to the grammar, what happens when you try to load multiple subsystems written by different people using different syntax extensions which, together, make the grammar ambiguous?
I don't know that these problems are insurmountable, but a few people have taken a crack at them, and AFAIK no one has produced a system that any significant number of people want to use.
It's worth taking a look at how Lisp gets around these problems. Lisp has not so much a syntax as a simple, general metasyntax. Along with the well-known syntax rules for s-expressions, it adds the rule that a form is a list, and the meaning of the form is determined by the car of the list -- and if it's a macro, even the syntax of the form is determined thereby.
Add a package system like CL's, and you get pretty good composability of subsystems containing macros. You can get conflicts, but only when you explicitly create a new package and attempt to import macros from two or more existing packages into it.
Applying these ideas to a conventional language gives us, I think, the following:
() While the grammar is extensible, all user-added productions must be "left-marked": they must begin with an "extension keyword" that appears nowhere else in the grammar.
() Furthermore, those extension keywords are scoped: they are active only within certain namespaces; elsewhere they are just ordinary names. This requires parsing itself to be namespace-relative, which is a bit weird, but probably workable.
I think that by working along these lines it might be possible to add extensible syntax to a conventional language in a way that avoids both the grammatical difficulty and the composition problem. And if you do that, maybe you can then get the relevant committees or whoever to standardize the AST representation for the language.
I've never taken a crack at all this myself, though, because I'm happy writing Lisp :-)
http://magpie.stuffwithstuff.com/index.htmlMy goal is to make AST's as available and easy to traverse/transform as they are in Lisp. This is the foundation that makes things like Lisp's macros as powerful as they are. And easy access to AST's enables so many other things like static analysis, real syntax highlighting, and detecting syntax errors as you type.
In a way, Lisp-like macros are just a special-case of tree transformation that puts the tree transformer inline with the source tree itself. But this is not the only possible approach. You could easily imagine an externally-implemented tree transformer that implemented GCC's -finstrument-functions. This tree transformer could be written in any language; there's no inherent need to write it in C just because it's transforming C.
It's true that a complier/interpreter could be reluctant to expose their internal AST format. But there's no reason that the AST being traversed/transformed has to use the same AST schema that is used internally; if you can translate the transformed AST back to text it could then be re-parsed into a completely different format. And with a correctly implemented AST->text component, this would not be a perilous and fragile process like pure-text substitution is.
Also: I recall Matz said Ruby was lisp with friendlier syntax (but I can't find the quote right now, so maybe he didn't).
Is this considered a macro? Is it homoiconic? It's code as data and using input variables to generate code based on that input. It struck me as weird the first time I read through it but figured since I'm pretty stupid that there's a good reason for it.
I'm willing to give up that 20-25% to enjoy and be happy writing the other 80%.