Difftastic, a structural diff tool that understands syntax
difftastic.wilfred.me.uk
difftastic.wilfred.me.uk
Those 2 massive innovations are leading to an explosion of tooling improvements like this. Now every editor, diff tool, or whatever can support dozens or hundreds of languages without having to duplicate all the work of every other similar tool. That's freaking amazing.
Just a bubble right now. It will come back to its natural uses after it. Everyone is doing AI now and I am pretty sure it is to attract investment even if some might know their product will go nowhere.
I spent the 1990s building actual AI software, but we had to call it something else because if you even whispered "AI" in the 90s your funding would dry up instantly.
Often when writing a linter, you need to bring along the runtime of the language you're targeting. E.g., in python if you're writing a parser using the builtin `ast` module, you need to match the language version & features. So you can't parse Python 3 code with Pylint running on Python 2.7, for instance. This ends up being more obnoxious than you'd think at first, especially if you're targeting multiple languages.
Before tree-sitter, using a language's built-in AST tooling was often the best approach because it is guaranteed to keep up with the latest syntax. IMO the genius of tree-sitter is that it's made it way easier than with traditional grammars to keep the language parsers updated. Highly recommend Max Brunsfield's strange loop talk if you want to learn more about the design choices behind tree-sitter: https://www.youtube.com/watch?v=Jes3bD6P0To
And this has resulted in a bunch of new tools built off on tree-sitter, off the top of my head in addition to difftastic: neovim, Zed, Semgrep, and Github code search!
Vs code is better at debugging and maybe slightly better at remote connections, that yes. But for the rest of things I am way more productive with Emacs than anything else.
Mac only, for now.
I'm getting very annoyed by things that don't mention they only work on Mac until you go to install them.
- Hyperlinked C++ BNF Grammar (https://alx71hub.github.io/hcb/)
- EBNF Syntax: C++ (ISO/IEC 14882:1998(E)) https://www.externsoft.ch/download/cpp-iso.html
The second doc has a year in the title, so it's ancient af. The first one has multiple `C++0x` red marks (whatever that mean, afair that's how C++11 was named before standardization). It mentions `constexpr`, but doesn't know `consteval`, for example. And doesn't even mention any of C++11 attributes, such as [[noreturn]], so despite the "Last updated: 10-Aug-2021", it's likely pre-C++11 and is also ancient af and have no use in a real world.
Who might have thought. /s
It feels kind of as foundational as YACC.
LSP truly solves the M editors to N languages needing M * N many integrations by using a standard interface for a query oriented compiler. Tree sitter doesn't solve this problem, it just makes it way easier to write N many integrations for your editor/tool.
1. Clone someone's tree-sitter grammar off GitHub.
2. Build it into a Mac .dylib.
3. Create a Nova extension that says "use this .dylib to highlight that language."
4. Use it.
I don't have to make any changes to Nova itself, and the amount of configuration I have to write is so tiny that Nova could have a DIY wizard if they wanted it to.
The source for Difftastic discussed here (at https://github.com/Wilfred/difftastic/blob/master/src/parse/...) is also very simple: for each of a list of supported languages, import the tree-sitter parser and wrap a teensy amount of configuration around it.
How is that possible if the different tokens emitted by tree sitter don't have standardized names? Isn't there some kind of configuration that maps the rules in the grammar to whatever convention Nova uses for their token names?
Now tree sitter does make this super easy, but my point was you still have to have some kind of per-language configuration/logic to work, whereas the entire point of LSP is to have none.
The maintainer of the tree-sitter grammar is usually one who maintains that mapping. At least, every time I've wanted to use it, it's been the case that all of that was already done and part of the grammar's repo.
[1] https://www.masteringemacs.org/article/how-to-get-started-tr...
https://news.ycombinator.com/item?id=39762495
(1) The top comment is from the author of difftastic (the subject here), saying that treesitter Nim plugin can't be merged, because it's 60 MB of generated C source code. There's a scalability problem supporting multiple languages.
The author of Treesitter proposes using the WASM runtime, which is new.
(2) The original blog post concludes with some Treesitter issues, prefering Syntect (a Rust library that accepts Textmate grammars)
Because of these issues I’ll evaluate what highlighter to use on a case-by-case basis, with Syntect as the default choice.
https://www.jonashietala.se/blog/2024/03/19/lets_create_a_tr...
Other feedback:
(3) The idea of a uniform api for querying syntax trees is a good one and tree-sitter deserves credit for popularizing it. It's unfortunately not a great implementation of the idea
(4) [It] segfaults constantly ... More than any NPM module I've ever used before. Any syntax that doesn't precisely match the grammar is liable to take down your entire thread.
---
I think some of the feedback was rude and harsh, and maybe even using Treesitter outside its intended use cases. But as someone who's been interested in Treesitter, but hasn't really used it, it seems real.
One problem I see is that Treesitter is meant to be incremental, so it can be used in an editor/IDE. And that's a significantly harder problem than batch syntax highlighting, parsing, semantic understanding.
---
That is, difftastic is a batch tool, i.e. you run it with git diff.
So to me the obvious thing for difftastic is to throw out the GLR algorithm, and throw out the heinous external lexers written in C that are constrained by it, and just use normal batch parsers written in whatever language, with whatever algorithm. Recursive descent.
These parsers can output a CST in the TreeSitter format, which looks pretty simple.
They don't even need to be linked into the difftastic binary -- you could emit an CST / S-expression format and match it with the text.
Unix style! Parsers can live in different binaries and still be composed.
The blog post use case can also just use batch parsers that output a CST. You don't Treesitter's incremental features to render HTML for your blog.
At the same time, I believe that there needs to be a corrective about what tree-sitter should and should not be used for. There are companies building security products on top of tree-sitter which I think is an objectively bad idea given its problems and limitations. Difftastic is to me a grey area because it could lead hypothetically to a security issue if it generates an incorrect diff due to an incorrect tree-sitter grammar. Unlikely but not impossible.
Your point about batch vs incremental is spot on, though even for IDEs, I think incremental is usually overkill (I have written a recursive descent parser for a language in c that can do 3million lines per second on a decent laptop which is about 60k lines per 20 ms, which is the window I look to for reactivity). How many non-generated source files exceed say 100k lines? Incremental parsing feels like taking on a lot of complexity for rather limited benefit except in fairly niche use cases (granting that one person's niche is another's common case).
That being said, it is impressive that their incremental algorithm works as well as it does but the cost is that grammar writers are forced to mold a language grammar that might not fit into the GLR algorithm. When it doesn't work as expected, which is not uncommon in my experience, the error messages are inscrutable and debugging either the generator or the generated code is nigh impossible.
Most of the happy users have no idea how the sausage is made, they just see the prettier syntax highlighting that works with multiple tools. I get that my criticism is as welcome as a wet blanket, but I just think there is something much better possible which your comment hints at.
That is, the big win is getting people to buy into the concept of syntax (and analysis) as a library and not as a feature of one specific editor. Once we're all spoiled by that, perhaps a better implementation or an nice API will come along and astound us all.
I'd understood that incremental was used so that as someone writes code the IDE can syntax highlight the incomplete and syntactically incorrect code with better accuracy. Is that not the case?
I don’t believe that’s true, but it’s likely correct for the common use case of files a few pages long, written in well supported languages.
I'm not saying that incremental is bad per se, but that the choice of guaranteeing incrementalism complicates things for cases where it isn't necessary. I am not super familiar with lsp, but I can imagine lsp having a syntax highlighting endpoint that has both batch and incremental modes. A naive implementation could just run the batch mode when given an incremental request and later add incremental support as necessary. In other words, I think it would be best if there were another layer of indirection between the editor and the parser (whether that is tree-sitter or another implementation).
Right now though, you have to opt in whole hog to the tree-sitter approach. As mentioned above, incrementalism has no benefit and only cost for a batch tool like difftastic or semgrep to mention two named in this thread.
I do wonder how much of a range there is on non-brand-new computers though. I'm typing this on an M2 Max with 64GB of RAM. I also have a Raspberry Pi in the other room, and I know from hard experience that what runs screamingly fast on my Mac may be painfully slow on the Pi.
I could also imagine power benefits to an incremental model. If I type a single character in the middle of a 30KLOC document, a batch process would need to rescan the entire thing where a smart incremental process could say "yep, you're still in the middle of a string constant".
I have no doubt that interactive editors like Atom/Zed can really make use of incremental parsing, and also lenient parsing.
Syntax highlighting and parsing isn't the only thing they do -- they still need the CPU for other things.
But yeah the problem is incremental is very different than batch, and lenient is very different than strict, so basically every language needs at least 2 separate parsers. That's kind of an unsolved problem, and I'm not sure it can be solved even in principle ...
This is because visual line diffs for an essay is bonkers. Usually the sentence changed starts in the middle of a visual line.
--word-diff-regex=<regex>
... A match that contains a newline is silently truncated(!) at the newline.
If I understand this correctly, if you use newlines inside a sentence (if you are writing a fixed width document, for example), this won't work.[0]: anyone have a good source for this? I’m not sure where I first encountered it
There is no philosophy more important in this age.
You can often achieve your goals much more quickly by using tools they way they best support being used.
- One Sentence Per Line (OSPL) - Semantic Line Breaks (SemBr) - Semantic Linefeeds - Ventilated Prose - Semantic newlines
Reading through the pages below was helpful in getting a better idea of what language people use to discuss this. They're mostly historical retrospectives or arguments for the merit of semantic newlines.
https://rhodesmill.org/brandon/2012/one-sentence-per-line https://ramshankar.org/blog/posts/2019/semantic-line-breaks https://vanemden.wordpress.com/2009/01/01/ventilated-prose https://discuss.python.org/t/semantic-line-breaks/13874 https://discuss.python.org/t/one-sentence-per-line-for-peps-... https://sembr.org https://asciidoctor.org/docs/asciidoc-recommended-practices/...
(Actually I think one-sentence-per-line denotes something slightly different from semantic-line-breaks, not that I know what that difference is).
Mock outrage aside, whimsy and play in written language is vastly cheaper than in industrial programming environments. Provided, of course, the author can yet communicate while horsing around.
One sentence per line doesn't mean your sentence has to be limited in length.
I have set ",',(,[,{ in visual mode to cut the selection insert the pairs then paste it back as a very hacky solution, but it gets the job done. If you want something more advanced to add or change anything around the selection tpope has solved that with vim-surround[1].
I just wrote a language parser a few months ago in tree sitter and it’s probably the most delightful software I’ve used apart from ffmpeg.
Yeah, decided to check this out to see if it could help review in our massive C-based project. Unfortunately, in a recent patch, of the 90 "hunks", 88 of them had fallen back to "normal diff" because "$N C parse errors, exceeded DFT_PARSE_ERROR_LIMIT").
One appeal of the general idea of a structural diff tool, for me, is ignoring the ordering of things for which ordering makes no difference.
x = 4
y = 7
are independent statements and the code will be no different if I replace those two statements with y = 7
x = 4
However, this information is not actually present in the abstract syntax tree. If I instead consider these two statements: x += 3
x *= 7
it is apparent that reordering them will cause changes to the meaning of the code. But as far as the AST goes, this is the same thing as the example where reordering was fine.What kinds of things are we doing with our new AST tooling?
Not always, e.g. in a multi threaded situation where x and y are shared atomics. Then unless we authorize C++ to take more liberties in reordering, another thread will never see y as 7 while x is not yet 4 in the first example, but not the second. This kind of subtlety can't be determined from syntax alone.
The page has several examples:
1. Understand what actually changed.
This appears to show that `guess(path, guess_src).map(tsp::from_language)` has been changed to `language_override.or_else(|| guess(path, guess_src)).map(tsp::from_language)`. The call to `map` is part of a single line of code in the old file, but has been split onto a line of its own in the new file to accommodate the greater complexity of the expression.
The bragging associated with the example is "Unlike a line-oriented text diff, difftastic understands that the inner expression hasn't changed here", but I don't really care about that. I need to pay close attention to which bits of the line have been manipulated into which positions anyway. I'm more impressed by ignoring the splitting of one line into several, which does seem to be a real benefit of basing the diff on an AST.
2. Ignore formatting changes.
This example shows that when I switch the source from which `mockable` is imported from "../common/mockable.js" to "./internal.js", the diff will actively obscure that information by highlighting `mockable` and pretending that `"./internal.js"` is uninteresting code that was there the whole time (because it was already the source of some other imports). This badly confuses a boring visual change ("let's use the syntax for importing several things, instead of one thing") with a very significant semantic change ("let's import this module from a completely different file"). I'm not comfortable with this; there must be a better way to present this information than by suggesting that I shouldn't be worried about it.
(A textual diff, in this case, has the same problem. But when the pitch is that your new tool is better than a textual diff because it understands the code, failing to highlight an important change to the code is worse than it used to be!)
3. Visualize wrapping changes.
This shows that when I change the type of some field from `String` to `Option<String>`, the diff will not highlight the text "String", because that part hasn't changed. This is a change from a textual diff, but it doesn't appear to add much value.
There's a second example to do with code that belongs both before and after other code, in this case an opening/closing tag pair in XML, but in that case the structural diff appears to be identical to a textual diff.
4. Real line numbers.
"Do you know how to read @@ -5,6 +5,7 @@ syntax? Difftastic shows the actual line numbers from your files, both before and after."
I agree that that's a real benefit, but again it doesn't seem to have anything to do with the difference between textual and structural diffs.
------
I think the conceptual appeal of a "structural diff" is that it fails to highlight changes to the code that don't change the behavior of the software. Difftastic clearly believes something different; in the second example, they are failing to highlight a change to the code that does change the behavior of the software. And in the other examples, they are failing to highlight things that haven't changed from some perspectives, but could be argued to have changed from other perspectives -- and that in either case don't derive much benefit from not being highlighted. If changing `String` to `Option<SpecialType>` produced a diff that highlighted `SpecialType` in a separate color from the surrounding `Option<>` wrapping, indicating that the one line of code contained two relevant changes, that might be interesting, but otherwise I don't see the point of not highlighting the inner `String` along with the new wrapping.
So... what is the appeal of structural diffs?
I was just replying that if you want to not get a diff for your example to which I replied you have to use a more advanced representation of the code, and AST won't be able to do it.
All tree-sitter gives you is a _different_ grammar, so that a structural diff can operate on different trees given the same text as diff.
A parse tree still doesn't know anything about the meaning of a program, which is what you need to know in order to determine that those assignments to x and y are unordered.
Built on the shoulders of giants.
cargo install cargo-update
cargo install-update --list
cargo install-update --all
Other fun Rust projects available via cargo:https://mise.jdx.dev/ mise-en-place, a drop-in replacement for asdf https://asdf-vm.com/ that is really fast and flexible.
https://github.com/ajeetdsouza/zoxide is a fantastic cd replacement, which stores where you cd to, and you can then do a partial match like "z hel" might take you to "~/projects/helloworld".
https://github.com/bootandy/dust is a compliment to "du", shows which directories are using the most disk space.
Looks great, thank you for the recommendation.
But ncdu is a fully interactive file browser that lets you navigate through the tree, and crucially it lets you delete things without requiring a full rescan. It's amazing for freeing up disk space by deleting things you don't need anymore, which is probably 95% of the reasons I run `du`.
I am currently using it with direnv, but there's enough functionality in mise to replace that too. I keep meaning to spend some time making mise work without direnv, but it's not been urgent since everything pretty much just works now.
I like the active development since the dev really seems to care about doing a great job, and I've been lucky enough so far that my (simple) workflows haven't been impacted.
https://github.com/jdx/mise/discussions/1525 for an example of how I use direnv with mise.
- https://github.com/eza-community/eza (ls with some added visual sugar)
- https://github.com/ClementTsang/bottom (htop but with graphs)
- https://github.com/sharkdp/bat (cat with syntax highlight)
Now, it works great! I tried it with a mise update, and it pulled down the right binary with no proxy problems.
Thank you for the reminder and recommendation, much appreciated!!
I am curious if there’s been any work on _semantic_ diff tools as well (for when eg the syntax changes but the meaning is the same). It seems like an intractable problem in the general but maybe it’s doable and/or useful for smaller DSLs or subsets of some languages?
if you do this your difftool becomes a compiler
You probably have a mental image of catching something really simple, and, yeah, "1 + 1" -> "2" is reasonably easy, but in reality there aren't a lot of those super easy changes. Most of the time there is something confounding the situation.
Truly neutral refactorings are pretty uncommon in their own right. You can see that when someone is discussing semantic versioning and pointing out that if you define a "major version" as "there exists at least one possible use of the code whose behavior will be changed as a result of this library change", almost any API change is automatically a major version change, which isn't really what anyone wants. E.g., in Python, the mere fact that introspecting on an object's methods will show one more method than it used to isn't really what we want a major version change for. In general, proving refactorings are actually 100% safe is equally difficult; even simple arithmetic changes can result in things overflowing at different times or in different ways, it's virtually impossible to rewrite an expression involving floats without the change being witnessable somehow, extracting a function could make it so that code that previously didn't overflow the stack now does, memory allocation changes can be the difference between OOMing and not and may interact with GC in unpredictable ways if you get really precise, etc.
TL;DR verifying that 2 functions have the same output is really freaking hard.
You can edit a function you've committed into the Unison code repo, and if you didn't change the semantics of the function, it's actually stored under the exact same hash... All places using the function refer to it by its hash, so nothing needs to be recompiled either, and no tests need to be rerun.
Things like renaming variables, reordering code whose order doesn't matter (common in functional programming) and things like that do NOT change the hash.
I believe this is only possible because Unison is a Pure Functional Language. If it's not, it becomes a NP problem to decide if two programs are exactly equivalent, probably.
I wonder if Unison could provide the actual semantic diff you're thinking of, it's probably not much more complex than actually knowing the meaning of the code did change. Maybe create a Feature Request :) https://github.com/unisonweb/unison
Some linters and formatters are effectively compilers already, so that doesn’t seem completely implausible in itself. Finding canonical representations of common coding patterns so you can quickly and reliably determine that they are equivalent is a different question, though.
But I'm glad it's easy to change that default.
So when using such a diff tool you can spend hours refactoring something, and then git will refuse to commit your changes because your refactoring was successful in not changing the behavior of the code? I understand what you mean, but if we arrive at that point maybe we should stop calling it "diff", to avoid confusion...
It does use diff to generate patches, however. I know in today's GitHub-dominated landscape, that's considered a bit of a dusty feature, but it would be a pity to break it.
We are working on https://semanticdiff.com/ which detects basic semantic changes like converting a literal from decimal to hex or reordering keys within JSON objects. It is not a command line utility but a VS Code extension and GitHub App. You can check out https://semanticdiff.com/blog/semanticdiff-vs-difftastic/ if you want to learn more about how it works and how it differs from difftastic.
Preach!
Just dropped it in and did a git diff.. works like a charm!
Do you not?
The one starting with a minus sign is for the original file, the one with the plus prefix is for the new file.
See https://www.gnu.org/software/diffutils/manual/html_node/Deta... for a canonical source and more detail
[1] https://github.com/tree-sitter/tree-sitter/issues/130#issuec...
This just does diff but not merge, but at least it's open source - and the diffs look a lot nicer, I've already made it my default.
Any plans to extend it to merging?
The GitHub readme:
> Can difftastic do merges?
> No. AST merging is a hard problem that difftastic does not address.
> AST diffing is a also lossy process from the perspective of a text diff. Difftastic will ignore whitespace that isn't syntactically significant, but merging requires tracking whitespace.
https://news.ycombinator.com/item?id=27768861 (297 points | 3 years ago | 61 comments)
https://news.ycombinator.com/item?id=32746258 (698 points | 2 years ago | 90 comments)
https://news.ycombinator.com/item?id=30841244 (983 points | 2 years ago | 219 comments)
I didn't find in the documentation how it is possible to change the style of the diff, or to ask for another color in the bold case.
Any idea?
I chimed in on this issue[1] to express support for that. You may wish to also chime in on that one, or open a new issue if you think that the feature you're looking for is sufficiently different from the one discussed in that thread.
Processed 1 file, 614 regular extents (614 refs), 0 inline.
Type Perc Disk Usage Uncompressed Referenced
TOTAL 14% 10M 77M 77M
none 100% 1.1M 1.1M 1.1M
zstd 12% 9.8M 76M 76MAn executable is not loaded into RAM completely and then started. It's memory mapped and only the parts that actually get used are loaded (when they are needed).
This is why large binaries get slower when you compress them, because then the demand-paging doesn't work anymore
I did the test with a 280mb Windows EXE some years ago, compressed down to like 70 megs or so but took multiple seconds longer to start up than the original
For some scenarios might make sense (running binaries across the network maybe) but in most cases an uncompressed binary will start up faster
* the company was acquired by Splunk years after we shelved that product
Because xml files are very often not using the extension xml and treating them as plaintext means losing a lot of the structural goodness.
https://github.com/Wilfred/difftastic/blob/master/CHANGELOG....
difftastic.enable = true; cargo install --locked difftasticMost diff tools throw out something that is quite difficult to parse, but difftastic gave me the most concise diff so far.
In general, we're overflowing in TMI which makes it hard to suss out what matters. For example at work I often read docs that describe what we do for customer X vs customer Y and it takes a ton of work to suss out the 1% of text that is different between those two, which is really what you want to understand and validate.
So anything that makes just the impactful change stand out is beyond welcome.
(I don't expect you to say "great idea, sir, here's your coupon!". I'm sure you've done your analysis and you're at the price point that works best for you. I'm also not saying it's overpriced, just that it's more than I'm willing to pay for it given the functionality I want from it. This is just in the spirit of friendly user feedback.)
I tried it, but unfortunately it's not as seamless as it could be, so I reverted back to Jetbrain's native diffing, which is quite good anyway.
Since Emacs is widely used for parsing, and can parse using tree-sitter, Emacs doesn’t seem to benefit from difftastic. But perhaps I’m overlooking some capability of difftastic.
Also, on Arch there doesn't seem to be a man page.
Some quick googling turns up https://github.com/tree-sitter/tree-sitter-embedded-template which may or may not meet your needs.
Is there a patching tool that can apply difftastic diffs?
Looks like it wasn't all that hard. Honestly the harder part is getting my shell to remember the changes to PATH.