Difftastic: A diff that understands syntax
github.com
github.com
e.g. see https://groups.google.com/g/golang-nuts/c/1BlZDNBLiAM
Having said that: if a Go compiler for a given architecture decided to change its layout algorithm, I'm pretty sure it would earn a changelog entry.
Language designers need to contend with the fact that the ultimate final say in whether a thing is or not is whether that behavior is observed.
> If one or more of the communications can proceed, a single one that can proceed is chosen via a uniform pseudo-random selection.
https://go.dev/ref/spec#Select_statements
In an early implementation it would pick in lexical order, IIRC (and the specification did not mention how a communication should be picked). Not only could this lead to bugs, apparently some people were relying on it and they didn't want that.
The tl;dr is that there's an almost infinite number of ways to atomize/conceptualize code into meaningful "units" (to "register" it, in my supervisor's words), and the most appropriate way to do that is largely perspectival — it depends on what you care about after the fact, and there is no single maximal way to do it up front.
Just thinking about it makes my head spin. I spend a lot of time working out font/color hierarchies, supplementary to coding and data viz. Arguably what you're bringing up is a case for a carefully colored diff that visually cues whether something is a true semantic change or indicative of a lower level issue. I'm comfortable with reading a plain ol' diff that just shows me what changed, superficially, and interpreting it. While I think OP's idea is awesome, it also might create more confusion than it resolves; and resolving confusion is the point of a diff.
https://victorcmiraldo.github.io/data/MiraldoPhD.pdf#page=24
The real difficult part is in how you represent AST-level changes, which will limit what your merging algorithm can do. In particular, working around moving "the same subtree" into different places is difficult. Imagine the following conflict:
([1,3], [4,2,5]) <-- q -- ([1,2,3], [4,5]) -- p --> ([1,3], [2,4,5])
Both p and q move the same thing to different places so they need a human to make a decision about what's the correct merge. Depending on your choice of "what is a change", even detecting this type of conflict will be difficult. And that's because we didn't add insertions nor deletions. Because now, say p was:
([1,2,3], [4,5]) -- p --> ([1,3], [2,5])
One could argue that we can now merge, because '4' was also deleted hence the position in which we insert '2' in the second list is irrelevant.
If we extrapolate from lists of integers to arbitrary ASTs the difficulties become even worse :)
The practical problem, though, is that the Haskell compiler is limited/buggy, so you couldn't implement this for C, and you settled on a small language like Lua. If you _do_ extend this to other languages (perhaps port your implementation from Haskell to something else?), please post it on HN and elsewhere!
One thing I do find interesting (and a wish were different) is that only programming languages are supported, rather than data formats as well.
For example, two JSON documents may be valid but formatted slightly differently, or a common task for me is comparing two YAML files.
Comparing config files that have a well defined syntax and or can be abstracted into a tree (JSON, YAML, TOML, etc.) would be absolutely lovely, even and including (if possible) Markdown and its ilk.
HTML and XML are missing, too.
Sadly YAML, TOML and the others I mentioned are not there (yet?)
But re JSON:
> object keys in a different order
They can't be "in a different order" as JSON keys are not ordered. They can be whatever order, and would still be considered the same.
> array items are out of order
Then it's different, as JSON arrays are ordered. ["a", "b"] is not the same as ["b", "a"] while {a: 1, b: 1} and {b: 1, a: 1} is the same.
> you could UTF-8 encode characters or write out the same character using backslash notation, numeric or boolean data that might be wrapped in a string in one file but not in another
Then again, they are different. If the data inside is different, it's different.
I understand that logically, they are the same, but not syntax-wise, which is why I included the "differently formatted" "disclaimer", it wouldn't obviously understand that "one" and "1" is the same, but then again, should you? Depends on use case I'd say, hard to generalize.
This is what GP is saying, I'm pretty sure. Object member order is non-semantic in json, so in order to do a semantic diff (one that understands structure), you need to canonicalize the order of the two sides. Simply diffing the output of jq doesn't do that, because (afaik) jq doesn't alter the order.
Basically, if you want this to come up the same:
{"a":"b","c":"d"}
{"c":"d","a":"b"}
you need more than just `diff $(jq) $(jq)`.Can argue about whether a tool like difftastic should do that, I guess, but I would personally lean towards that it should be smart enough to see this because it's precisely the sort of thing that both humans and line-based diff can be awful at seeing.
It's more simplistic than difftastic though: it considers `1` and `[1]` to have nothing in common.
If a format has a tree-sitter parser, it can be added to difftastic. The TOML tree-sitter parser looks good, but there isn't a mature markdown parser for tree-sitter. There are other markdown parsers available, so in principle difftastic could support markdown that way.
The display logic might need a little tuning for prose-heavy formats like markdown though. I'm not happy with how difftastic handles block comments yet either.
I'm not sure about formats that contain more prose, such as markdown or HTML.
He was in love with and endlessly curious about English slang, it’s basically all we talked about.
I remember explaining to him why my uni friends and I referred to things as being “craptastic”, starting with American marketing’s love affair with the portmanteau.
He got it pretty quickly and enjoyed using it in conversation.
The saying that was harder for him to understand was “fuck all”. He always wanted fuck to be the verb, rather than using “fuck all” as the adjective, so he would say things like “I fuck all my money last night at the pub”.
Profanity is just delightful in general, and non-native English speakers come up with some of the best profane idioms in English.
I wonder if it’s the same in other languages?
Honestly, I cannot imagine going back to the standard emacs help.
Once I discovered Helpful, all of those things seemed so obviously useful that I can’t understand why nobody else thought to put them there, including myself.
Not to discredit Wilfred (it looks like he's taken over the project as the maintainer), but, based on the historical contributions [1], it looks like it was originally developed by Max Brunsfeld, who also created Tree-sitter. [2]
[1]: https://github.com/Wilfred/difftastic/graphs/contributors
[2]: https://github.com/tree-sitter/tree-sitter
[3]: https://github.com/Wilfred/difftastic/commit/958033924a2dea7...
My apologies to Wilfred.
Edit: And after installing cargo, watching it fail to build, then determining I must need a newer version of cargo, so I built that from source... it fails. Apparently I need to install `rustc-mozilla` and not `rustc`. "obviously".
This is all a testament to how much I want to try this tool...
MOAR EDIT: even with rustc-mozilla cargo fails to build. running `cargo install difftastic` gives me an error about my version of cargo being too old ;.;
Dear author: Let us run your tool.
Edit: Reinstalling Cargo worked!
It'll confirm that you want to install it, because it's already installed I think, and I just selected 1. for Yes.
hard pass :)
Why? You're willing to run some random open source project, but you're not willing to run the official Rust installation script?
Even if this specific instance of curl'ing into sh is safe, or if I download and then run it, it's still extremely poor practice and gives me serious doubts about the developers and their security practices in general.
I also do not like when every project decides to poorly reimplement the package manager. If every software used it's own package manager my system would be a complete mess with dozens of different package managers fighting each other and it would be a total nightmare to update the system or manage non-trivial dependency chains when installing something new.
Rust is one of my favorite languages but this is definitely my least favorite aspect of it all. It really feels like the developers "optimized" for systems with no package manager.
A get started guide with all the required commands easily copy-pastable? (A popular option these days) Something else?
I don’t mean to be critical, I’m simply curious.
Also, `cargo install difftastic` AIUI pulls it from a central location, if I'm gonna poke at software for the first time, I enjoy building it myself first, so I can get my hands dirty in the source. :)
EDIT: Also, the build fails. :(
"error: unexpected token: `include_str` --> /home/loxias/.cargo/registry/src/github.com-1ecc6299db9ec823/radix-heap-0.4.2/src/lib.rs:2:10 | 2 | #![doc = include_str!("../README.md")] | ^^^^^^^^^^^
error: aborting due to previous error
error: could not compile `radix-heap`.
sad trombone
curl https://sh.rustup.rs -sSf | sh
Restart shell to get $HOME/.cargo/bin in PATH, then did: cargo install difftastic
And ~4 minutes later, difft executable is ready.Agree though that some pre-built binaries would be fantastic!
It's considered really sloppy and unmaintainable to admin a system like that. Things quickly get out of hand.
That strategy _does_ work if you isolate it to a chroot or a container, but littering /usr/local with all sorts of locally compiled upstream is just asking for future pain. Security updates, library incompatibilities, &c.
Prebuilt binaries might be nice, but I don't expect them for random projects. (and I wouldn't have used them if offered) I do think it's a reasonable expectation to be able to build software w/o essentially setting up a new userland just for that tool though. :)
I'm sorry, and retract my ignorant assumption! Going to try it out now.
I've also had requests from Alpine Linux packagers to allow dynamic linking to parsers. This is something I want to support in future, once I'm happy with the basic diffing logic.
I've documented the minimum rust version required today, although I'm looking at lowering the minimum version.
spinned up a Ubuntu 18.04 instance -> git clone, git checkout 0.24.0
installed rust using curl | sh method
build fails:
removed the instance and gonna check it again 6 months later
= note: /usr/bin/ld: cannot find Scrt1.o: No such file or directory
/usr/bin/ld: cannot find crti.o: No such file or directory
Have you tried googling for "ubuntu crti.o: No such file or directory" ?Depending on the project, there is a certain threshold of trying-to-make-something-work which I'm willing to undertake in order to test an app.
But you are right. I'm sorry if my OG comment may come arrogant to the devs who do stuff for free. (♥ to the devs)
[edit]: ok, I tried again, `sudo apt update && sudo apt install build-essential` before installing rust and `cargo install`ing.
Error again:
vendor/tree-sitter-haskell-src/scanner.cchttps://doc.rust-lang.org/reference/linkage.html#static-and-...
$ cat /usr/local/bin/difftastic
#!/bin/sh
source $HOME/.nix-profile/etc/profile.d/nix.sh
nix run nixpkgs.difftastic -c difftastic "$@"
and then it'll install on first run: $ difftastic
these paths will be fetched (1.17 MiB download, 9.38 MiB unpacked):
/nix/store/wn74xn0w60xcwsly6nqaibn205hh2qms-difftastic-0.8
copying path '/nix/store/wn74xn0w60xcwsly6nqaibn205hh2qms-difftastic-0.8' from 'https://cache.nixos.org'...
Difftastic 0.8.0
Wilfred Hughes
A syntax aware diff.
USAGE:
[etc.]https://github.com/cregit/cregit https://lwn.net/Articles/698425/
That standard format is commonly known as source code - although it lacks a normal form.
Tools like prettier, gofmt and black can be thought of as a way to produce a normal form of source code.
This is (IMO) a reasonable incremental approach towards exactly what you describe - if a project checks in only source code that's formatted using a standardised format, then you're free to work on it using whatever equivalent representation you like - as long as it's converted back at commit-time.
The challenge for a tool like difftastic is that I can't guarantee that syntax is well-formed. You might be using new syntax that my parser doesn't support, you might have merge conflicts, or you might have a plain syntax error in your code.
Tree-sitter handles parse errors gracefully, so difftastic handles syntax errors pretty well in my experience.
Since moving to short lived feature branches it is less useful to me.
Maybe they don't want to any more? And this is just their subtle way of pushing everyone interested in using it away?
git-difft:
#!/bin/sh
GIT_EXTERNAL_DIFF=difft git diff "$@"
git-showt: #!/bin/sh
GIT_EXTERNAL_DIFF=difft git show --ext-diff "$@"
Then you can run "git difft …" or "git showt …" if you want to use it.Since I use VSCode as my editor, I created this oneliner in my .bash_profile:
# VS Code Diff
diffcb () { "/usr/local/bin/code" -n --diff $@ > /dev/null 2>&1 ; }
With it, I can "diffcb filename1.json filename2.json" to get a visual editor with contextual awareness based on installed lint modules.
However, I'm not sure if Org markup lends itself to structuring that would allow proper diffing—even with just the headings.
e.g. A merge that knows properties files support the same property added in different places but only once is needed. And another strategy if order is significant.
Cool to have an HTML merge that recognises the tree structure and supports merging tags and having the indentation follow some rules.
I believe git supports merge strategies, its been on my todo list forever.
I would recommend putting an installation guide in your readme, and it being a full installation guide.
I followed the link to your manual and then it told me to install your tool using a tool called "cargo" with no reference on how to install cargo. At this point I gave up. Lazy, maybe, but for a convenience tool like this I want a convenient installation.
Only diff is I got to the point where it said I needed "cargo", On a whim, I typed "aptitude install cargo", and it did something. Now waiting for the >1GB source repo to clone to see if it works.... ;)
(I have an example workflow here if anyone from there is interested https://github.com/conradludgate/wordle/blob/main/.github/wo...)
But I wish there was more of a convention in the F/OSS community that if your software isn't written in something universal (C, C++, shell and maybe python), then it also comes with a container of all that's necessary to run it.
It's frustrating to pollute my nicely packaged managed system with hundreds of locally installed python modules just to run one tool. Or, in this case, backport and rebuild a language specific build tool simply to compile. :)
>universal
* laughs in Windows, then cries *
But it's just once a year :) And the last time I was deep in windows was win7, whenever that was. I tried to use a win10 machine and gave up.
Besides, I thought the big new feature in modern windows was that WSL improved to the point you can run unix tools! ;)
It's really not that bad, the AD-IPA cross-forest trust is really solid as is the native sssd-ad integration if IPA is too much. Honestly I can't really imagine it any other way now, so much work has been put into AD support that it's actually the best login experience on Linux at the moment. OpenLDAP is definitely showing its age -- dgmr I use it for all my personal infra because it's free and my use-cases are dead simple but we got to delete so much bespoke code after migrating off it at work.
I'm not sure, and you undoubtedly know more and are more up to date than I, but I don't believe any of these things existed in 2005, when I was on the aforementioned team. Or, maybe they did exist but management decided an internal implementation was better.
Getting Windows to accept the user profile in an AFS path I recall being particularly vexing.
FreeBSD would be up your alley. Its native ACLs are NFSv4 format, a superset of NTFS ACLs. You need to enable it explicitly on UFS2, but it's default on ZFS.
Finally, even if you have a syntax tree, that is just part of the solution, probably the smaller one. Detecting three lines of code wrapped in a new if statement is easy but also doesn't benefit much from a syntax-aware algorithm. But once you changes names and signatures, extract methods, introduce constants, and so on it will become progressively harder to match subtrees and one is probably quickly approaching the territory of NP-hard and undecidable problems.
I've thought before this is how diffing should be done, and speculated that tree-sitter would make it more feasible.
At this point, whenever I think some language-aware tool ought to exist, my first thought is "Does the language server protocol or tree-sitter make this more feasible?"
Languages usually change slowly, though, so once a good baseline grammar is in place, maintenance is unlikely to be a huge load.
Furthermore, with tools like tree-sitter and the language server protocol, multiple communities benefit from their continued existence, so there's a bigger pool of contributors to the parser.
I very much agree. I feel there has been a trend recently where people (re)discovered how cool and useful ASTs are and now expect everything be using them. I suspect old-school computer scientists might be secretly laughing at this while programming with some Lisp-like languages they invented for themselves.
Jokes aside, I do wonder how modern IDEs manage to parse broken source code into usable ASTs --- is this trivial (CS theory-wise) or are there a lot of engineering secret sauce involved to make it work?
[1] And also harder because if you want to parse the file after each key stroke, you have to be fast. This probably also makes incremental updates to the syntax tree the preferred solution and that might align well with using prior result for error recovery.
Don't we call such heuristics "test suites"?
var foo = bar baz
there are many ways to change it and make it parse including the following reasonable ones var foo = barbaz
var foo = "bar baz"
var foo = { bar, baz }
var foo = bar // baz
var foo = bar
//var foo = bar baz
var foo = bar * baz
var foo = bar + baz
var foo = bar.baz
var foo = bar(baz)
but also unreasonable ones like var abc = 123
and therefore a parser that can handle malformed inputs has to make educated guesses what the input was actually supposed to look like. And don't be fooled by this simple example, imagine a long source file with deeply nested code in a language with curly braces and randomly deleting some of the braces. Now try to figure out where classes, methods, if or try statements begin and end in order to produce a [partial] syntax tree better than just giving up at the position of the first error.And you actually jumped over the hard part that requires the heuristics, how to modify the input in order to make it parse. Take a 10 kB source file and delete 10 random characters - how will you figure out which characters to put back where? With 100 possible characters, 10,000 positions to insert a character, and having to insert 10 characters, you are looking at something like 10^60 possible modifications. You are certainly not going to try them one after another, each time checking if the modified source file parses, compiles, and passes the test suite.
Not sure what this whole straw man is about. I definitely didn't suggest anything like that. Of course you can only compare two compiling versions of a source file using a test-suite-based heuristics. I thought this whole thing was about "heuristics that identify reasonable changes which fix the file" mentioned above? "Reasonable changes that DON'T fix the file" are clearly recognizable by NOT passing the test suite, just as if it was a human trying to make those changes and finding out that the change that he just did didn't in fact yield the desired results after running the test suite.
> With 100 possible characters, 10,000 positions to insert a character, and having to insert 10 characters, you are looking at something like 10^60 possible modifications.
If you're working with an AST, you're almost certainly not working with characters. That would be immensely wasteful. In fact working with an AST is pretty much the only way in which the set of changes is sufficiently reduced for almost any change to NOT be rejected outright. With character-level modifications, you're facing the problem that almost every edit will be outright rejected as early as at the stage of parsing.
class foo
{
function bar() {
function baz() { }
}
it should be able to parse the file as if bar() was not missing the closing curly brace. If the parser just gave up or inserted the closing curly brace at the end class foo
{
function bar() {
function baz() { }
}
}
making baz() a nested function inside of bar() the result would be worse than using a character-based diff algorithm. But I never intended to say anything about making code functionally correct, that is none of the business of a parser or diff algorithm.Seems there's a good open market for such a lazy reason.
But why? Shouldn’t the code you push into a repository be at least syntactically correct? And even if it is not, one can simply fallback to textual diff.
> Or, if you only pick a small set of supported languages, your diff tool will not work on most files or have to fall back to a structure-agnostic algorithm.
I don’t see how it is a blocker.
(1) Parsing an arbitrary language is hard. Without tree-sitter, difftastic would probably be a lisp-only tool. You also want a parser that preserves comments.
(2) Inputs may not be syntactically well formed.
(3) Efficiently comparing trees is extremely difficult (difftastic is O(N^2) in time and memory).
(4) Displaying tree diffs is equally difficult. Alignment is particularly challenging when your 'unchanged before' and 'unchanged after' are syntactically the same, but textually different.
On the other hand, even though my diffs are usually not that huge, sometimes they might be, and I don't want to switch tools every time that happens (I have just git alias and I don't even remember my exact config, nor should I care). So being slow is not great.
I'm using it like "git-ydiff-s" script in my PATH to use "git ydiff-s":
#!/bin/sh
git diff "$@" | ydiff -s --wrap --width=0
Installation is "sudo dnf install ydiff" or
curl -fsSL https://raw.github.com/ymattw/ydiff/master/ydiff.py > ~/bin/ydiff
chmod +x ~/bin/ydiff # (and change to python3)Build instructions? Nope.
Minimum system requirements? Nope. But if you check out cargo.toml, you'll see it says it needs Rust 1.56.
My system has 1.48.0 . And it the latest Debian release! I don't see how a diff tool can expect you to have a bleeding-edge development environment. I mean, ok, you chose a new language - I can understand that; I won't demand that it build with just a C compiler and Make. But come on, this is not supposed to be just a toy for new systems.
Anyway, I still cloned it, tried to build with "cargo build", and got stuck with:
error: unexpected token: `include_str`
it couldn't even tell me "get Rust 1.56" :-(
Meld takes a diff, and applies syntax highlighting over the diffed files. It additionally highlights the changed characters in a line. Git diff, vimdiff and probably others, do this as well.
From the demo, I understand that Difftastic first applies syntax and then rebuilds the patch over that. Being aware of line wrapping, changes in nesting, moving codeblocks into functions and so on.
Is this just a pretty website, or is the software actually available anywhere?
https://www.plasticscm.com/pricing
Looks like no locally-run-binary/non-SaaS version. I was hoping it'd have SublimeText like model. I have no interest in trying to get my team to switch nor having to deal with the security team when it turns out I was using a free cloud account.
> Difftastic output is intended for human consumption
Why not separate the human-consumption part and the underlying parsing part? Or at least provide both in the same utility?
Difftastic then converts the tree-sitter parse tree to a simpler s-expression style format (see https://difftastic.wilfred.me.uk/parsing.html#simplified-syn...), and computes differences on that.
I'm just trying to clarify that I'm not generating conventional 'unified diff' patches, so I can provide a nicer interface (e.g. line numbers).
It looks very handy though. I still do a lot of C and C++
> Non-goals
> Patching. Difftastic output is intended for human consumption, and it does
> not generate patches that you can apply later. Use diff if you need a patch. brew install rust
cargo install difftastic
Worked for me without any problems.With a textual diff today, your only choices are 'highlight all whitespace changes' (e.g the git default) or 'ignore all whitespace' (e.g. diff --word-diff).
If difftastic says there are no changes, then both files have the same parse tree and the same comments.
The indented basic block won't show as a difference, only the start and end of the block.
I disagree. I struggle to replicate it right now using a simple test, but I've seen the following rather infuriating and counter intuitive behaviour from Git/GNU diff. If you have a simple if statement such as:
if (bla) {
// do something
}
And you were to add another statement at the end, after the closing curly brace, e.g.: if (bla) {
// do something
}
if (bla2) {
// do something else
}
Git/GNU diff will sometimes show the following diff: diff --git 1/left 2/right
index c2ea6f1..dc0e1c2 100644
--- 1/left
+++ 2/right
@@ -1,3 +1,6 @@
if (bla) {
// do something
+}
+if (bla2) {
+ // do something else
}
This is basic example, but there's other similar things. For a simple change like the above, this isn't a huge issue, but for a bigger patch sets, it can take a minute to understand what is really going on.[0] https://git-scm.com/docs/git-diff#Documentation/git-diff.txt...