Searching a Million Lines of Lisp
wilfred.me.uk
wilfred.me.uk
"Programs are not text; they are hierarchical compositions of computational structures and should be edited, executed, and debugged in an environment that consistently acknowledges and reinforces this viewpoint."
It is unfortunate that most major code editors and IDEs today do not store units of code (like Lisp SEXPs) as databases to be updated as you edit, which would make problems like this or other operations like metaprogramming, formal analysis, or documentation much easier to solve.
It's cool, but it has some major downsides too. For example, MPS stores the source as XML, not text (since it isn't text, it's a tree). This makes lots of basic tools we've taken for granted a lot harder, such as git merging etc. They've had to make a custom mergetool just to make basic collaborative coding feasible.
I bet there's other ways around that, all I'm saying is that text has major, major upsides because of the enormous ecosystem support.
If only there were a way to write your code using a uniform tree syntax in the first place...
Doesn't solving this just require a text-to-AST, AST-to-text input and output step?
Anyplace outside the editor the programmer just sees regular text.
Text is probably the next biggest mistake in programmer productivity after null.
With a rich structure editor that can do merges, the undo history of edit and refactor operations could be persisted and merged into the VCS. Currently this isn't possible. Text is a projection for the page and a lowest common format.
Directly operating on structures would mean that you'd have to write an editor, which had to enforce correctness as well. And then you'd have to write a generator to save those structures in some kind of format that could be written to a file and passed around, and a parser to read such format. And check for correctness again, since who knows what generated that file.
As for constant compilation, that already exists, many IDEs have it. That's because parsing text is not actually hard, the other stages are.
With a rich structure editor that can do merges, the undo history of edit and refactor operations could be persisted and merged into the VCS. Currently this isn't possible.
Of course it is, you could write a plugin for any IDE that would record edits and refactor operations and save those in or alongside the text (much like they've have to be save alongside the AST). Of course, that doesn't help if the user does a manual refactor, but that's no different than they choosing a node in the rich editor, deleting it, then manually recreating it in its refactored form.
Consider that every programming language and every config language first invents a new syntax to encode a tree like structure (typically using a combination of curly braces, other brackets, keywords, indentation etc.) but the code itself is saved as 'text'. This is a lossy encoding - all a generic reader such as `git` or `grep` can now infer is that the file contains a 'sequence of lines' and can then only offer line based operations (git diffs are line based, grep searches are line based, etc.), when in fact a more meaningful operation would be the tree structure based.
If a tree based format was the canonical baseline, diffs could display the location of the node added (e.g. 'Added <Class X> -> <Function Y>'), without having language specific parsing knowledge. Similarly, most editors could provide 'tree view' and 'jump-next', 'jump-up' etc based on context, again without knowing language specific details. Further, many internal representations of programs (e.g. intermediate representations in compilers) also use trees, and could potentially be exported into one of these forms, to make the plethora of tools work with them.
(BTW, I'm not saying a tree is the best generic structure to replace text, but just using it as an example to argue for advantages of a generalized extensible structure over plain text.)
Why do you argue so vehemently against someone perusing an avenue of research?
It would finally settle the formatting arguments and all the tooling would be able to leverage each others projects.
Even PHP has an internal AST representation these days.
Having taken Teitelbaum's compiler class as an undergrad, I was happy to ditch the IDE he inflicted on us and go back to text editors. Structured code editing is a tricky UI problem and I didn't find a really good IDE until many years later.
There's no one way to edit code. Sometimes refactoring tools work well, but typing text can be quite efficient too. Getting locked into a tree editor at the expression level is no fun.
If they are they either aren't offering user/programmer access to that database, they aren't doing semantic/type binds, or aren't advertising those features well.
>I was happy to ditch the IDE he inflicted on us and go back to text editors. Structured code editing is a tricky UI problem
It's reconcilable with normal text editing, just update the structure once its valid. I agree that things like block/visual programming can be absurd.
But this indexing is only on one user's workstation and tends not to scale up well. Updating dependencies or switching to a different branch means rebuilding large parts of the index.
Also, part of the problem is that there is little standardization. Many ecosystems are language, platform, build tool, and/or editor-specific. When you do something new you end up reinventing the wheel.
[0] The Cornell Program Synthesizer: A Syntax Directed Programming Environment https://core.ac.uk/download/pdf/21750999.pdf?repositoryId=14...
https://help.github.com/articles/searching-code/#considerati...