Difftastic, the fantastic diff
wilfred.me.uk
wilfred.me.uk
> This IMO really highlights the power of having an ecosystem built around tools like tree-sitter because they allow for powerful dev UX tools to be more democratized
LSP is also an example of this and it is also an example of how dangerous it can be when these tools are backed by big companies: https://news.ycombinator.com/item?id=31760684
If microsoft, tomorrow, pulled the plug on the pylance LSP and removed all access, and the only thing left is the opensource version, the world is still at a better position than if they hadn't invested any money or opensourced anything.
This is strictly different from google owning chrome, because with that ownership, they have the ability to dictate protocols for the web.
The LSP is a protocol, which now is quite widespread, and will continue to survive and be enhanced, regardless of microsoft's involvement or not. The fact that they can choose to make some language servers close-sourced sucks, but being able to access a closed source component is _still_ strictly better than having zero access in the first place.
Microsoft has the resources and ability to make better language servers which could be closed source now, so people end up switching to those tools, owned by Microsoft, with better language servers, killing other editors if they are not able to catch up.
E.g.: Why would a C# programmer use Neovim if the language server is worse than the one that VSCode has? (Which is now proprietary and closed-source). The difference right now is insignificant, but the future tools and features that they are adding to the new proprietary C# language server will not be available for other editors.
[1] https://en.wikipedia.org/wiki/Embrace,_extend,_and_extinguis...
they can only extinguish by forcefully bankrupting alternatives. Unfortunately for microsoft, there's no such thing as bankrupting an opensource project.
If neovim's feature set can't keep up with the might of microsoft's development budget for vscode, and many users switch over, then it's a net-win for those users. If microsoft then removes the C# LSP afterwards (or charges money for it), neovim is still around. Those users who don't want to pay can move back to neovim (and may be someone would then start contributing to it again, and over time, the features of neovim will improve again).
Microsoft killed netscape by depriving it of a business model. But you cannot deprive opensource of a business model, because that model is not profit oriented.
Unless microsoft patents the LSP protocols and stops other people from using it (and even then, other editors do not die, unlike netscape), there's no real way for microsoft to EEE.
Google can deprive firefox of a business model (directly by stopping their payment to mozilla, and indirectly, by making the web protocols so restrictive that firefox cannot implement them properly, such as DRM). Therefore, google is in a position to EEE other browsers.
Lots of open source does, in fact, come out of a profit oriented business model. Open source is harder, but not impossible, to kill by killing the business model of whoever runs the current project because open source provides rights to other people to pick it up even if the original sponsor’s interest becomes to end it, so as long as there is someone or some group with the resources and interest, it can continue.
Admittedly, this is a somewhat far fetched scenario that I have posited but I am somewhat concerned that the proliferation of incorrect tree-sitter grammars is going to lead to the worst kind of problem down the road: tools that work great _most_ of the time.
I don't know if any languages do this.
in reality damn near all languages are parsed easily enough that there is no difference.
A better interface would be a simple binary that produces JSON output from AST, and another binary that produces AST from the same JSON output. Then you'd do a diff on the JSON and converting it to AST before displaying. Then you’d only need to modify the compiler to print out the AST as JSON (and read it in as JSON), instead of reimplementing the parser.
I'm curious exactly why A* failed here. It worked great for me, as long as you design a good heuristic. I imagine it might have been complicated to design a good heuristic with an expanded move set. I see autochrome had to abandon A* and has an explanation of why, but that explanation shouldn't apply to difftastic I think.
I use delta as my daily driver but sometimes when I want the contextual info, switching to `env GIT_EXTERNAL_DIFF=difft git log -p --ext-diff` gives a better picture.
That's probably the only reason, otherwise, you are correct. I would want the context always.
+1 insightful
The VSCode extension "Diff & Merge" give you the right/left arrows to merge lines if anyone is looking for a tool that does that. I haven't needed another one since I found that.
Install the CLI, run the command (alias diff='diff2html -s side') - I run this at least every time before committing to quickly see all I've done.
Seems like an unnecessary extra step.
https://www.plasticscm.com/semanticmerge/documentation/intro...
(I don’t know enough about the internals of Git to answer this myself.)
Yes, you can turn it on without having any side-effects for others.
1. The nesting/unnesting/merging/splitting problem that many tree diff algorithms have a hard time with are handled here by allowing the diff to see insertion/deletion of delimiters. I guess one way to see it is that the tool is doing a text level diff augmented with the tree structures to calculate the cost of diffs and choose the cheapest one it can find.
2. The problem of syntax errors. I think this just depends a lot on how well tree sitter copes with syntax errors (or weird syntax that’s hard to parse or eg a committed merge conflict) but my understanding is that it is designed to cope ok with syntax errors.
I basically felt that tree diffing was not very viable because of these issues but seeing this project I think I’ve changed my mind. I guess it remains to be seen how good performance is though (maybe this isn’t good enough yet if it can sometimes be very slow).
Though, to be clear, I think you put the word in quotes because you think the problem isn’t trivial.
Ultimately what you really want is many different kinds of specialized error (and warning) nodes for different mistakes users might make. For example, in C++ the statement `auto x = 1,000,000;` is technically correct but probably not in the way the user wanted. So you put a warning node there to record that they might want to use `1'000'000` instead. If you see something like `auto x = 42\n if (y) {…`, then you might guess that there is a missing semicolon before the newline. If a speculative parse of the next line succeeds, then you put a MissingSemicolon node into the AST (instead of one saying that the “if” was unexpected), followed by whatever the speculative parse found.
Now repeat that process for five years as you polish your compiler into a sparking user–friendly gem.
In practise I haven't noticed any issues yet. I suspect that difftastic doesn't see many syntax errors because users have fixed most of them by the time they run the diff. When you look at diffs of committed code you hopefully have no parse errors at all.
The tree-sitter parsers could still reject valid code I suppose. I worry slightly more about the parsers getting precedence/associativity wrong, but it would be hard to construct an example that produce identical parse trees due to incorrect precedence.
>Autochrome and difftastic represent diffing as a shortest path problem on a directed acyclic graph. A vertex represents a pair of positions: the position in the left-hand side s-expression (before), and the position in the right-hand side s-expression (after).
>The goal is to find the shortest route from the start vertex (where both positions are before the first item in the programs) to the end vertex (where both positions are after the last item in the program).
I don't understand what this means at all? What is a "position" here? Position of what? If it's a graph, what do the edges represent? The diagrams afterwards aren't very helpful either, I can't make head nor tail of them. They talk about a "start" vertex and an "end" vertex, but before that it said that a vertex is a pair of start-end positions ... I'm totally lost.
Autochrome has a brilliant worked example https://fazzone.github.io/autochrome.html and it still took me several readings before it clicked.
Every graph vertex represents a pair of pointers (or positions) to AST nodes. So in the example, the start program is `A` and the end program is `X A`. The positions point to AST nodes in these programs.
I try to use 'vertex' consistently for graph vertices, to avoid confusion with AST/s-expression nodes. If you have any suggestions for better terminology I'd be very interested too :)
Writing this comment it occurs to me that the structural diff doesn't translate to plaintext very well, and thus is not accessible to folks with red-green colorblindness.
Advanced diff tools usually allow configuring the colors, so that if the red/green is a problem, you can change it to red/blue or blue/yellow or whatever.
(foo (bar))
(foo (novel) (bar))
The balanced parentheses around bar “haven’t changed”, but the symbol between them “has”, and the balanced parentheses around bar were “added”. It’s a perfectly valid diff given the approach, even if it’s not ideal, as the author found.The only way you could arrive at “novel) (” given the algorithm, I think, would be something like
(foo (bar))
(foo novel) ((bar))
(This might be off too, hard to write code on mobile)I agree that the classic red/green colour scheme of diffs isn't great for colourblind users. I've asked a few colourblind developers and they were happy with terminal ANSI colours (which is what difftastic uses), because they can configure each colour individually.
The approach from Chawathe et. al splits nodes from the before/after trees into chains by their label in the syntax grammar, and then runs myers’ longest common subsequence on each pair of chains. Some parameters t, f are used to have an approximate ‘equals’ method for subtrees.
This iteratively builds a set of matchings between equivalent nodes from the old and new trees. Here’s the paper https://dl.acm.org/doi/10.1145/235968.233366
I’d be curious to see if this approach handles re-ordering of nodes better. The ‘fastMatch’ algorithm described above will typically miss matching cases where a node that is not order sensitive (i.e a function in a namespace can be moved somewhere else in that namespace).
I know that sdiff by Arun Isaac is based on applying that paper to Scheme: https://archive.fosdem.org/2021/schedule/event/sexpressiondi... , although he also reports performance challenges.
Most tree diffing papers that I've seen focus on either (1) providing a minimal diff and accepting the performance cost or (2) providing a relatively minimal diff and focusing on the performance.
I've generally found that you need a minimal diff to get a good result, so papers in (2) are less applicable. I've also found several cases where there are several possible minimal diffs, but there's a clear 'correct' answer from the user's perspective.
Difftastic doesn't handle moves: the edit set is add, remove, or replace similar comment. If you reorder functions, it will take the largest unchanged subset. Moves are hard to model in a diffing algorithm, but they're also very hard to display coherently in the UI.
I know a few code forge websites (e.g. Phabricator) show moves in a fairly comprehensible way, although they're all based on line-based diffs.
Are the algorithms open? Does the tool work reliably?
I find it very sad that potentially revolutionary progress is lost like this...
https://github.com/cregit/cregit https://lwn.net/Articles/698425/ https://www.youtube.com/watch?v=iXZV5uAYMJI
Is there any chance such an alternative differ could be used in Git (and adjacent tools like GitLab), or are we stuck with line-based forever?
Due to this constraint, it is unlikely to be the most convenient way for a human to input that something.
Take arithmetic expressions, for example. The usual infix syntax as edited in a conventional text editor is very good, probably the best option that humans have for editing arithmetic expressions. Having to edit an AST structure explicitly would almost certainly be a loss in terms of convenience and productivity.
Or take a (raster or vector) image editor. They provide interfaces to editing image data in a much more convenient way than editing AST or other internal representations directly.
https://www.diffnow.com/compare-clips or http://incaseofstairs.com/jsdiff/
Does anybody know any better alternatives which work with pasting?
In that case, assuming that one file is OK and only the other text needs to be pasted, you can diff against stdin. On macOS there is `pbpaste`, which prints the contents of the pasteboard to stdout. This allows you to do `pbpaste | diff the-file -`, with `-` being the commonly used "filename" for stdin.
I use Araxis Merge at work, and Meld at home, both of which can work this way. (Araxis is much better, but I can't justify the cost personally. Meld is fine, just a bit slow.)
You can use it like `cursh difftastic {{1}} {{2}}` and it will prompt you for each {{#}} to hit enter once you have the thing you want in your clipboard and it passes the content to the command as files