The Grep Test
jamie-wong.com
jamie-wong.com
The "smartass" way of programming is sometimes overused but it does have its benefits. When you're metaprogramatically setting the attributes on the User, you're also avoiding needless and error-prone repetition and making sure that this central piece of code will either crash all the time or work all the time for all attributes. This has tremendous value.
So while I understand the point about this article, I might want to add a pinch of salt to the dogma underlying it.
I think there are definite ways of adding metaprogramming functionality without breaking this test. For instance, in the first JavaScript counterexample, if the iteration was over [{attr: "position", fn: "getPosition"}, {attr: "direction", fn: "getDirection"}] instead, the Grep Test passes, and you get much of the same benefits, with a very minor duplication that I'd argue is worth the cost.
my %to_generate = (
'sub first' => { ... },
'sub second' => { ... }
);
generate_from_spec(%to_generate);
generate_from_spec then strips off the 'sub ', and errors out if it is not found.The second case I have is a function that itself then generates a function, but is trying to look like a sort of normal function itself, in which case I end up with:
generate_sub routine_name =>
... whatever arguments ...
"abusing" Perl's => operator, which functions like a comma except that it forces stringification of the left argument, to once again make the literal, greppable "sub routine_name" appear in the codebase. Here "routine_name" is then just a standard string argument, the alterative being generate_sub "routine_name",
which is then harder to grep for. (Still possible, obviously, with a different grep query, but only if you already know up front you need to add the other possibilities.)Note this actually goes a step beyond what you are proposing in that it makes the declaration site clear; my counterproposal for your JS case would be
["function getPosition", "function getDirection"].each(...)
and using string manipulation to do whatever you need to do to get the right info out of the function name.> Seems a bit extreme to me.
That reasoning has absolutely zero argumentative value behind it, in any context that it's used. It ought to be treated like a logical fallacy.
> add a comment mentioning the methods called in there.
The problem of comments getting out of sync with code is omnipresent. "Never fail the grep test" seems like a much more easily-enforced and -maintained practice (both for yourself, and for teams) than what you're suggesting.
Simply using it to save keystrokes is pretty lame since you'll spend much more time maintaining code than typing out the original and explicitness is valuable when returning to a piece of code. Additionally, if you find yourself that you need a lot of metaprogramming for a lot of things it's often an indication that you could just refactor your code using static idioms and be better off.
Not to mention that the more you use metaprogramming, the more likely it is that you or someone else will kill some runtime optimizations of the JIT.
IMO, syntax matters way less than people tend to think it does, and the additional implementation complexity and astonishment cute syntax introduces tends to make the trade-off not worth it.
I've been having a lot of fun recently writing code in wat-pl (my perl port of Manuel Simonyi's wat-js interpreter)
For example this CAD system from PTC is based on 6+ million lines of Common Lisp code and under development for two decades:
Well, Arc is actually very maintainable, and quite fast. After all, HN is powered by Arc, and it serves >100k daily uniques on a single core.
~ $ time (echo '(do (= n 0) (for i 1 100000000 (++ n i)) prn.n (quit))' | arc)
Welcome to Racket v5.3.5.1.
Use (quit) to quit, (tl) to return here after an interrupt.
arc> 5000000050000000
real 0m28.561s
user 0m28.308s
sys 0m0.249s
----
~ $ time (echo -e 'n=0 \ni=0 \nwhile (i <= 100000000): \n n += i \n i += 1 \n\nprint(n)\n' | python3.3)
5000000050000000
real 0m30.244s
user 0m30.230s
sys 0m0.013s
It would appear to be competitive with Python on my machine on this particular task. (Also you can make it faster by dropping into Racket.) $ time python -c 'print sum(xrange(100000000 + 1))'
5000000050000000
real 0m1.398s
user 0m1.383s
sys 0m0.012s
Comparison to baseline: $ time (echo -e 'n=0 \ni=0 \nwhile (i <= 100000000): \n n += i \n i += 1 \n\nprint(n)\n' | python)
5000000050000000
real 0m33.140s
user 0m32.939s
sys 0m0.023s arc> (time:xloop (i 0 n 0) (if (> i 100000000) n (next (+ i 1) (+ n i))))
time: 9121 cpu: 9130 gc: 0 mem: 480 ; the times are in msec
5000000050000000
Or perhaps a "higher-order function": arc> (time:sum idfn 1 100000000)
time: 19889 cpu: 19908 gc: 0 mem: 1224
5000000050000000
Or use a deforestation macro that I wrote, which is closest to your Python example: arc> (time:thunkify:reduce + (range 1 100000000))
time: 17971 cpu: 17985 gc: 0 mem: 3592
5000000050000000
Also, here's what you can get by dropping into Racket: arc> (time:$:let loop ((i 0) (n 0)) (if (> i 100000000) n (loop (+ i 1) (+ n i))))
time: 402 cpu: 403 gc: 0 mem: 920
5000000050000000
I suppose Python has an analogue of that--dropping into C, or at least loading C libraries. Which Racket can do too. Mmm.Using metaprogramming should not make you "feel bad", however it should not be used excessively. Metaprogramming gives you a lot of power, it allows you to extend the language in any way you please. With this great power, of course, comes great responsibility but I would rather have the option to use this power than to not have it at all.
I would also argue that metaprogramming can enable your code to be more efficient, if the macros/templates you write are evaluated at compile-time. This is the approach that the Nimrod programming language (http://nimrod-code.org) takes and it works well. This also means that in most cases compilation will fail if the macro is incorrect instead of causing some silent errors at runtime.
I think a better rule is if the tooling can't look at the source, intermediate compiler stages/ AST and bytecode/binary and tell you unambiguously where something came from, then you can characterize it as a "last resort"
And there's the counterargument that for statically typed languages with good type systems (scala, haskell, ocaml, F#), there's REPL prototyping/testing and code generated from template haskell or camlp4 has to type check, so you have multiple safety mechanisms
you could just refactor your code using static idioms
you or someone else will kill some runtime optimizations of the JIT.
syntax matters way less than people tend to think ... astonishment [sic] cute syntax
Many intelligent, design-savvy people would strongly disagree with you. Besides the obvious gains in the ability to create DSLs, there is a power and fluidity in metaprogramming that is not possible any other way.
Speed (as, say, provided by runtime optimizations from a JIT compiler) is not always critical. If it was, one probably should be using a language other than Ruby or Javascript; they're meant to be flexible, powerful languages, and their emphasis is not on pure speed.
Syntax is not "cute." Humans are going to read and use this code, and the more it reads like English, the more likely it is to be understood.
I am unsure what you mean by the "obvious gains" you get from DSLs. I see many DSLs as a code smell -- that the runtime environment is almost expressive enough to express the syntax construct the DSL creator wants, but not quite.
The implementation of DSLs tends to judiciously use closures, operator overloading, dynamic getters and proxies such that it's non-obvious what's going on under the hood. Sometimes that trade-off is worth it. Usually it isn't, in my experience.
I believe (I could be wrong) that DSLs were coined around 2004. I don't think that is a long enough history for us to even think about simply accepting that they are a good idea in and of themselves -- garbage collection was invented in 1959 and is still being debated!
I see the speed argument all over the internet being presented as a binary argument: either you care about speed or you don't. That's simply false -- I care about order of magnitudes for speed. I personally don't care if I can encode h264 faster in JS than I can another way, but I do care that my event handlers execute within 16ms so I don't drop frames in the browser.
Some syntax is certainly cute. Additionally, making syntax readable like English should be a non-goal IMO; for many people, Lisp is far more readable than SQL, which is much closer to English.
As an earlier example of an Embedded DSL, the book PAIP[1] included prolog embedded in common lisp in 1992.
It could be argued that the loop macro in Common Lisp is a DSL for describing iteration. If not loop, then certainly regular expressions are a common EDSL for describing a regular language, and performing operations with those languages against strings?
I found this snippet from Computers in Crisis [2] which seems to describe domain-specific languages in familiar terms from 1975
Most domain-specific programming languages can be categorized in one
of two ways; either as a "sugared" general ... of programming, in
fact the style of problem solving, embedded in and supported by that
language remains unchanged.
[1] http://www.norvig.com/paip/README.html
[2] http://books.google.co.nz/books?id=QndQAAAAMAAJ&q=%22Embedde...configuring Spring Beans configuring Log4J defining URL mappings configuring Hibernate mappings & queries etc.
Now some of those are cool, but at least some of them replace other techniques that I already know, and find much more readable and understandable. Configuring Log4J, for example. I'd much rather simply jam a log4j.xml in there and forget about the Grails DSL.
The thing is, none of these DSL's is, in and of itself, necessarily bad in any way... but there's a bit of "cognitive overload" in having to deal with 3 or 4 or 5 new DSLs, on top of the base Groovy stuff. The flip side is, these DSLs mean that your configuration is done mostly in .groovy files and you can use normal groovy syntax for looping and accessing variables, etc. So it's easier to code up more dynamic configurations than if you were using XML files or .ini files.
Anyway, the point of this rant is not to say "DSLs are bad" but just to lend weight to the suggestion that they aren't universally Good either.
OK, sure, but after you've grokked the first "find_by" usecase...do you really need documentation for all the other kinds of "find_by"'s that you'll use?
In any case, I'd agree that meta-programming is too often abused, but the proposed grep test is far too strict. And, inability to create complete documentation for every single kind of method token is not really the main reason to avoid meta-programming...I'd say performance and the propensity for abuse are better motives.
This pull request re: removing most of Rails' dynamic finders in 4.x covers the topic nicely:
I agree. In a CMS I'm coding, the core MVC-like mark up tags work by using function tables in Go. Such code would fail the grep test - but it makes a lot of sense in the project and it actually very readable as well. In fact in that particularly use case, not only is the code more readable, but it's also more efficient.
But this is the age old problem with having rules in languages (both human and programming) - more often than not, there are exceptions that are perfectly legitimate use cases.
I also agree that Rails gets a pass. It's a little different when you have a stable (well, sort of :-) API with dozens of books and thousands of blog posts. It was annoying to learn what was going on at first, but that knowledge is more long-lasting so worth a bit more pain.
For example, try to step through the code that figures out which validation callbacks to trigger for an ActiveRecord model. You'll be led through at least 3 dynamically eval'd methods.
I had to do this recently and probably would have given up had there not been Foo.method(:bar).source_location to point me in the right direction.
I'm still trying to figure out how Python projects scale to any appreciable size; I suspect maybe they don't. I've worked on several million-line codebases in statically-typed languages, are there any truly large Python projects?
The general underlying problem is that a lot of people stick with those "principles" (which are rules of thumb, often weak ones at that) they have read somewhere without developing a real understanding of software design. Slogans like DRY or SRP seem to blind people to often simple forseeable consequences of their design decisions.
Software design boils down roughly to two abilities: one is imagining a great many ways of structuring a program, the other is the understanding of practical consequences of a given program structure. So a good software designer must ask him- or herself: why code repetition is a bad thing? And the answer is that there is absolutely _nothing_ wrong. What is wrong is that a logically or structurally common aspect of the problem didn't receive the recognition in the form of an abstraction (a method, class, variable, interface ...). Then you waste time by not being able to reuse this abstraction, when something about this common aspect changes you are forced to go through fifty places and modify the code, but the worst thing is that there is a limit of the amount of details a programmer can keep in his or her head, and the less you are able to structure your program the more restrictive will be limit of the problem complexity you will be able to tackle. That's the reason software design is at all important.
There are cases however when there is no real logical common denominator to two pieces of code and they are similar practically by accident. It is just as wrong to abstract an accidental similarity as it is not abstract an existing one; soon the requirements change and the abstraction will have to be abandoned, code copied and developed in divergent ways.
Finally, people start attempting metaprogramming way before they had learnt enough about structuring programs using basic means, like breaking things down into classes and methods appropriately, data-driving your programs etc. Yes, this is actually a skill, and unless you rewrite some of your own programs 5-10 times and compare their structure you won't learn it. I also recommend reading some classical books: SICP, Refactoring, Refactoring To Patterns, Effective Java, Programming Pearls.
And the ultimate lessons come from maintaining your own code for a few years.
I read it has him saying "hey I found this easily identifiable test which seems to correlate with code that's possibly too clever for it's own good".
I also believe this test is worth thinking about because I tend to use grep/Find in Files a lot on my own code. It's often just simpler and faster and equally effective to the sophisticated IDE code comprehension stuff.
> a lot of people stick with those "principles" (which are rules of thumb, often weak ones at that) they have read somewhere without developing a real understanding of software design
We all have to start somewhere. I don't think it's possible to get "a real understanding of software design" without making some mistakes along the way. At best, we can help newcomers avoid the most painful ones and wasting too much time on dead ends.
But (IMHO) all good programmers learn by failure (their own as well as others).
In the example posted, position and direction are both vectors, recognizing them as such would provide a much better way of checking for "zeroness" and for negation. It's a nice educational aspect of this toy problem the author seems to not have noticed himself. There is often a similar elegant way out of real-world complex problems that doesn't involve fancy meta-programming but just good abstractions, putting right data in the right place etc.
My understanding is the point of doing the grep test is to avoid capturing uninteresting abstractions not likely to see much usage as the program gets more complex. If you are a programmer who sometimes loses sight of the program's real goal because you're worrying about code duplication, the grep test might be useful for you.
Ok, it's actually right under one false assumption: that we are limited to using our existing text-based tools like text editors, grep and so on. As long as we use these existing tools, this advice is quite valid.
However, we are absolutely not limited to existing tools, we can and should make new ones that augment how we work and allow us to reach more dryness without compromises. You could have tools that let you work with ASTs, that expand code on the fly, show you higher level information, let you debug and step through sections of code effortlessly.
In the end, dryness is good, because it means you have less repetitive work to do when things change (and software is all about changing, unless you don't want to make progress). So it's very well worth investing into being able to achieve dryness more naturally without the downsides.
To quote Bret Victor, stop "blindly manipulating symbols." (You don't have to do do so overnight, just as long as it's a goal you work on achieving as time goes on.)
Yes, grep is primitive for this task. We have much better tools for static languages, and they are great. It will be great to get them for dynamic languages. But they will fail in pretty much exactly these cases where this primitive test fails. Because halting problem.
The fact that most dynamic languages were hacked together doesn't imply that the better systems also fail.
A statically typed language can do arbitrary stuff at run-time too that will not be amenable to static analysis. This has nothing to do with the type system.
Furthermore, a dynamic language can be perfectly amenable to static analysis. For instance, in lisp+slime, slime-who-calls can identify callers of a particular function, even when grep would not, e.g. if the function was invoked via a turing-complete macro. The key thing here is that the AST to be analyzed does exist at compile-time; it just happens to be generated rather than hand-coded.
On the other hand, the typical Ruby ORM defines methods that corresponds to the current columns of tables in the database it is told to connect to. That fancy AST parsing tool won't be able to do anything in that case without being able to connect to your database and issue queries...
You might consider this "hacked together" but those kinds of api's is a major part of the draw of Ruby for a lot of people.
Common Lisp has plenty of well-defined macros with simple patterns any tool could learn, but it's definitely dynamically typed. (defstruct foo ..) generates foo-p, make-foo, copy-foo; what parser couldn't understand that? All it takes is a few good macros to really make the DRY work. Either give up like the author and repeat yourself or move past the problem with a macro.
Patterns mean "I have run out of language." — Rich Hickey
If the tool has to learn (i.e. be written to know about, yes?) these patterns, then it will not know about those you wrote.
However, as someone else has pointed out, the compile-time execution of macros actually does make them accessible to a sufficiently sophisticated "grep", the kind of tool shurcooL asked for.
I work in C++ all day and loathe when functions are fully implemented in the class declaration. Keeping them seperate with full ClassName::FunctionName(...) scoping makes finding functions exceedingly simple and friendly. Keep implementations separate means the full class declaration is easy to read, parse, and understand. I don't want to scroll through hundreds of lines of code just to see what functions are available.
I found myself using a pattern of declaring a one-off interface for complex classes and then implementing it in the real class. I'm still undecided when/if this is a useful thing to do.
A published & concise interface wins over an undocumented redundant interface. Explicit-only just takes you down the path towards Java and COBOL.
Published interfaces may also not be possible in certain code bases or frameworks -- think of a Chef recipe for instance where most (ruby) code is being written to directly interact with the system rather than provide a service to other parts of the system. Dynamically generating variable names and key/value pairs in Chef recipes is a common pattern I've seen which makes grepping and isolating issues difficult if not impossible.
The rest of the time it is still a huge benefit to be able to easily navigate the source. And frankly, a lot of the time the source is a more concise source of truth than decently written documentation.
But I don't think the OP argued against either documentation or fixing bugs, so I fear that both of our comments are irrelevant.
It's an artefact of the design. You simply need to know how the abstraction is done, and then you look for invocations of that abstract more general form. It's not that difficult, and if it's being done, it is almost without exception a better way to do the thing.
If metaprogramming is not a better way to do the thing being done, and adds confusion and decreases reusability, then you shouldn't do it, but that's a tautology and doesn't mean we have to be able to grep for everything we ever write.
But I don't see why metaprogramming should inherently fail the grep test anyway. Sure, your property/thingy/whatever (an access method for a field automatically generated from a database, say) might be generic and generated at runtime, but its name should still be written down somewhere authoritative, even if it's in a schema or (if nothing else) a documentation file. That stuff should still be in your tree and greppable, and if it's not your code has discoverability problems.
Other languages in which I've enjoyed REPLs include Python, JavaScript and Haskell. Two languages I have not found a good REPL for are Java and C++. Is anything of this sort available there?
You can grep for every instantiation of the struct in the code-base.
If anything, macros are more problematic:
#define VTABLE_INIT(prefix) \
(struct vtable){ .add = &prefix##_add, .sub = &prefix##_sub, and so forth }That said, dynamic languages still have a great deal of value, and a significant portion of the programming population uses them, so I still think it's worthwhile to set down some useful rules of thumb.
I used to think this too.
- library code publishes an interface and can do what it wants
- your client/application code should be greppable
There isn't a sharp distinction between the two, but it's a pretty big headache if developers are using heavy metaprogramming everywhere in your project.The reason this difference is important is that metaprogramming is structured and explicitly defines the change in the program's semantics.
Thankfully I have never seen code like what was posted in my experience with Python so far. Explicit is better than implicit which the code example in this post would fail to pass IMO. Dynamically dispatching to names that don't exist in a class' published interface at call-time is a big no-no in my book.
Although practicality beats purity.
I haven't seen it yet but that doesn't mean there is a practical reason to use this method of dynamic dispatch. If there were one and it gets a problem solved NOW rather than waiting to find a better solution -- it might be worthwhile.
However it's a price you have to pay.
I think the grep-test is at least a good way to test the waters with a bit of code. I don't think it's a universal end-all-discussions rule.
This would also have the side benefit of being able to produce better designed search results.
EDIT: I suppose searching the AST wouldn't be enough, you would have to evaluate the code to some extent to be able to search these properties.
The situation is uniquely bad in JavaScript, since JavaScript doesn't formally have modules, functions can take any number of arguments, there's the funky scoping, and so forth. Not that good tools don't exist, but it makes it difficult for a JS IDE to enable working across files like a Java IDE can.
That being said, I've found the language a pleasure to work with (aside from browser incompatibility issues) and with the proper application of design principles such as modularity (check out require.js and the r.js optimizer) and, yes, DRY, I can be just as productive, if not more so, in JavaScript as I am in any statically-typed language.
EDIT: Also C# (so many languages, I can't even remember which ones I've used sometimes). I'm also not including TypeScript or Dart (optional static typing), which I've technically "tried", but don't know much about.
For example, in Haskell you can write code as if you're in a dynamic language without annotating all the types and yet you get all of the benefits of types. It's the best of both worlds.
Concluding that static languages suck after using java and c++ et al is a common mistake, and is understandable.
What I did say was that I can be just as productive in those languages as I am in JavaScript, if not more so. It remains to be seen whether I can gain more productivity with TypeScript or Dart.
Imagine the 'incredible boon to productivity' you could get from tools (text editors, IDEs) that can deterministically show you all references or definitions of any method, variable, or class throughout your codebase whenever you want. (Hint: this isn't a dream--this is a huge plus of working with statically typed languages.)
For all of the pop-trendy love that dynamic languages seem to get for 'being fast for development,' (read: hacking) it's tragic how much they slow you down when you need to start hunting down where the hell something was magically (or meta-) declared or changed, especially when you're working with someone else's code (read: real life.)
The grep test seems like a great approach if you're stuck with a dynamic language. Of course, we don't always have the luxury to choose the technologies or platforms that we work with, but you've got the choice, a statically typed language solves this problem out of the box.
(DESCRIBE #'functionname) from the REPL will give you information about the function; you can also pick up source if you configure it right.
The key idea is that software as written only loosely defines software images as they are live. 'Static' languages attempt to ensure that there is a tight correspondence, but dynamic linking defeats that in part.
Image based software ideas take this idea and run with it: that's why you download Smalltalk images, not smalltalk source.
Anyway, food for thought. :-)
One thing I think the post might emphasize more thoroughly is that dynamic function invocation (at least when the function is defined in the project scope), is probably far more problematic than dynamic function declaration, particularly if sometimes the function is explicitly invoked, and sometimes dynamically. In this situation, I'll generally try to document that the function is dynamically invoked with the function declaration, but I'm always unsure what exact information I should put there. Simply saying something like "Dynamically invoked -- grep will not find all usages" is a minimum, but I often want to add more than just that.
But yeah, I'd say it's more about http://en.wikipedia.org/wiki/Dynamic_dispatch than static/dynamic typing. You can have dynamic dispatch in a statically typed language. I don't know how OP would feel about virtual methods in C++ since he resists using an IDE.
You can do about anything you want with cglib / Javassist. Dynamically subtyping or byte code re-writing a class is how many Java ORMs work.
#define readinto(obj, fd) (isatty(fd) ? readbuffered(obj, fd, findttystate()) : readfromfile(obj, fd))A better title would be "Don't use a fastening device if it fails the hammer test."
Are you don't use a tool that generates SetFunkyColumnName from funky_column_name data spec ? Or "I can't find the click handler in the .html, quit using jQuery" ?
I have been at the bottom of steep learning curves multiple times where I couldn't figure out where stuff was coming from. In most cases I got over it
Use the appropriate tools that get the job done and make you and your team productive. Tools and ideas evolve at different rates. Not all tools or ideas are implemented properly the first time, and not all tools or ideas are necessarily good/useful. But we'll never get any further if we don't try...
Deleted comment
So in this case, writing longer, "grep-friendly" JS code can actually reduce the size of the JS payload you serve to your users.
It's madness to classify all use of metaprogramming as abuse. Perhaps you need to pour a glass of wine and learn to savor the source code, and to appreciate the power of modern programming languages.
(See Yegge's grok project, unfortunately he doesn't blog anymore)
All the Textmate users now seem to be burned and always fear the software they use will die, so they stick to open source tools only. Those are okay for dynamic languages and smaller scope projects but i still find IDEs indispensable for large codebases and languages like C# or Java.
The idea is that programming languages allow too much flexibility in form and syntax for any indexing system to accomodate the full range of possibilities. And "coding style" rules are apparently too restrictive on creativity: they are too difficult to enforce. And hence a solution would be better to focus not on the code, but on the comments. Force programmers to adhere to a uniform commenting system.
The OP mentions ctags. It's not a perfect system, but it's still in the BSD base systems so everyone who has BSD has a copy of the needed programs. That's a start.
What about cxref? Another old system that's probably not perfect, but seems like it was aiming in the right direction.
I've never understood why programmers obsess about things like verbose function names (that make lines go way over 80 chars and bend across the page... with identation it becomes almost unreadable to my eyes) instead of just providing an index of all functions and including the verbose information in the index, not the code. vi (and no doubt emacs too) allows you to jump around easily so you could look things up in an index quite quickly.
Why do Wikipedia pages have an index of footnotes and references at the bottom? Why not stuff all this information into the words in the body of the article? Why do books have indexes? Why do academic papers use footnotes? I don't know. But I'm accustomed to these conventions.
I also don't know why code uses verbose function names and generally lacks an index or footnotes. But I guess programmers have just become accustomed to these conventions.
This entire post is silly.
1. Not being able to grep for code has nothing to do with being DRY.
2. Being able to grep for something doesn't make it correct. Nor does it make it less maintainable if you don't used named functions.
Moving on...
The idea is that when you look at code, the computations you see are the ones actually executed. Meta-programming may be used but only to add behavior to existing code, not to nullify it.
A good example is an 'execute around' method. The method you see in the code is executed, but some things may happen before it and some things may happen after it. What you can't do is replace the body with something else.
An interesting thing about the 'No Lie Principle' is that aligns with good practice around inheritance also. It's better to override abstract methods than it is to override concrete ones for a number of reasons.
function do($class, $action, $param, $value) {
$method = $action.$param
$obj = new $class();
$obj->$method($value);
return $obj
}
$my_name = 'Arnor';
$$my_name = do('Entity', 'set', 'Name', $my_name);
// The entity is now stored in the variable $Arnor
Thanks PHP... Don't even get me started on __call.It's fun to come up with clever solutions and puzzles, but it harms your code base. If you're proud of how new and clever your last 10 lines were, you probably need to refactor it.
As for "Not grepping the right file", I meant it should be findable in a project-wide grep, not in the file I assume it to be in.
For example, Guice fails the Grep Test hard, but it incredibly helpful.
All data-driven fail the Grep test. Your browser fails the Grep test (you can't grep for javascript content)
You just need to replace Grep with an xref tool that understands your programming language, including your metaprogramming. This may require you to commit your configuration files and standard data objects into your source control, or build an indexer that can read your CMS as well as your code.
It's more involved than just grepping for the name, but the same is true of tracing where any other argument value came from.
Basically the important aspect is that even if you have higher order functions, the things passing the functions to these higher order functions will still be greppable.
Lua 5.1.4 Copyright (C) 1994-2008 Lua.org, PUC-Rio
> function f(g) g[1]() end
> function thisFunctionCrashes() undefinedVariable() end
> f({thisFunctionCrashes})
stdin:1: attempt to call global 'undefinedVariable' (a nil value)
stack traceback:
stdin:1: in function '?'
stdin:1: in function 'f'
stdin:1: in main chunk
[C]: ?
In the example, the source line is available, but in the genuine article the thunks are for native functions.It's a difficult trade-off, one one I've had to make calls on several times. Do you use the full capabilities of incredibly gifted and talented programmers, and then allow code into your codebase that maintenance programmers can't understand?
Tricky.
The solution is to learn how to hit the nail with the hammer, not your thumb.
There's nothing wrong with using metaprogramming to generate methods, as long as you are writing tests for those methods.
Grepping the application code is usually much less useful than grepping the tests to see how the system is intended to behave.
user=> (meta #'leiningen.gnome/uuid)
{:arglists ([project]), :ns #<Namespace leiningen.gnome>, :name uuid, :column 1, :line 10, :file "leiningen/gnome.clj"}The point of DRY isn't to mindlessly remove code duplication. It's to remind us to look for code duplication and keep us mindful of coupling between the various parts of our code.
Where two identical chunks of code that represent the same kind of work are used in multiple places we introduce "algorithmic coupling." That is, whenever the work being done in one location changes, we have to make sure to change the work being done in the other location. Anyone reading this code -- whether the author, the author's future self, or teammates -- has to remember this extra fact, increasing the surface area for bugs.
There's also "name coupling," viz., for every vector _foo_ associated with the Ray we want methods like foo_is_zero?, negative_foo, etc. Here it's important that the naming convention be consistent, so there's coupling there, too. If the names aren't consistent anyone else reading the code would then have to remember that fact, and anyone changing the names would have to remember to update all the other names, too.
The irony of his example is that style of metaprogramming is a great way to get rid of name coupling, but he did nothing to get rid of the much worse algorithmic coupling. Indeed, algorithmic coupling screams for DRY whereas name coupling requires it on a more case-by-case basis.
That is, when all the names are highly localized, e.g., three short methods that share some naming pattern are all defined in succession, it's much less important to remove the duplication. Anyone reading or editing that code will quickly see what the pattern is and why it exists.
Here's a comment I left on the blog:
Hmm. I don't think the lesson here is about to-DRY-or-not-to-DRY your code. Instead, it's about using the appropriate abstractions.
Using the Ray example, both position and direction are vectors, not points. Make them vectors! Then you'd be able to say things like
ray.position.zero?
# We'd usually say "the inverse of" not "the negative of"
ray.position.inverse
Furthermore, if you wanted to define the same methods, well... class Ray
def position_is_zero?
position.zero?
end
def direction_is_zero?
direction.zero?
end
end
There's less repetition, now, because you're only repeating names, not logic. This means the code will only break when the names change, i.e., zero? becomes is_zero? or something.In a world where all of your logic is also duplicated in the Ray class, the code would break when either the names or the logic changed.
Making code that passes the grep test allows many more programmers who are vaguely familiar with the codebase to make changes.
Making code that fails the grep test allows a team of a few highly skilled developers who know the codebase inside and out to do the work of hundreds.
It's like mathematical notation, you generally need experience in that sub-branch of mathematics to understand the notation.
Besides that, building teams of generalists is far easier if it's easy to dive into new code bases. Even with small teams highly specialized on individual projects, there are lots of advantages to having people be comfortable digging into other teams' code, which is much easier if it's more discoverable.
Much like languages have more successful ways of representing numbers, some languages are more successful at metaprogramming. In particular, Lisp takes the view that the act of metaprogramming is really the act of writing application-specific compiler extensions. The default behavior simply doesn't involve ripping apart strings and concating them back together.
The practice of placing any general functionality in lambdas or blocks will make thing more messy.