Especially for programming, the tools have to be narrowly tailored to the examples since you're constantly wrestling with the specter of Turing-completeness. Any given program is an instance of an infinite number of more general classes of programs, and it's the tool designer's job to choose which dimensions of the design space are meaningful and important enough to be worth simultaneously visualizing the consequences of possible alternatives. I think it's going to take some very judicious integration of recent work on modularity-enhancing programming paradigms into the design of reflective language implementations before it becomes tractable to build responsive special-purpose reflective tools on top of a generic infrastructure. Or at least that's the strategy I'm trying.