The Tyranny of the Diff
michaelfeathers.typepad.com
michaelfeathers.typepad.com
At our current level of expressiveness, the refactoring tools that make these semantic changes mostly aren't a huge step up from brute force text editing, and the mechanical changes we can automate are usually trivial compared to the way a developer might conceive a required change in the behaviour of his code. Trying to reverse the process to identify the semantic significance of changes after the fact is far beyond us today.
In fact, it's hard to see how we could ever move beyond that level without developing some new and much more semantically rich way to represent our programs. At that point, the idea of diffing raw text code side-by-side might seem like it comes from the dark ages anyway...
What has succeeded for many is scriptable text editors, cscope-like tools, autocomplete-like features, and refactoring support in editors. But as you point out, these are better tools to work with program text.
I think text is safe for the forseeable future. Partly because we have thousands of years experience with it and partly because it's a winning combination of a extremely powerful representation format and a KISS solution.
I don't think the hypothetical alternative representation I mentioned is going to happen off the back of a single grad school research project, or anything even close to that scale. It's more like something that's going to take the R&D lab of an industry heavweight a decade to develop and refine and then launch into widespread industrial usage with the backing of at least one major platform developer when it's ready for prime time.
Sadly, as long as text-based (or, perhaps more accurately, line-based) programming languages are good enough to produce acceptable software, there isn't a truly compelling motivation to develop a completely new model. I wonder whether increasing pressure on the software industry to provide quality and security as a backlash against the current cheap-and-nasty trend will drive us toward more radical programming models first, and perhaps get us close enough to make the quantum leap to an entirely new kind of representation from there.
Chunk format Smalltalk change logs are a text file, but they're combined with a mechanism for treating code changes like db transaction logs, in a language with very little syntax, with very fine grained codebase change (method level, with generally small methods) built into the language and environment. It's also as rock solid as traditional flat text files, and even more robust in some regards.
Code syntax which requires all information relevant in a class to be defined together interferes with a change log scheme. But such syntactic structures are helpful if dealing with code in traditional flat files. In a way, it's like an Evolutionary Stable Strategy. It's mostly a flat file world, so the languages are mostly implemented with that in mind. This general situation makes it almost impossible for anything else to develop its won ecosystem.
I think forward progress will be made by leveraging "code folding" schemes. Eventually, this will amount to the same functionality people are aiming for then they discuss non-text representations.
I say this because such a tool may be easily adopted into existing workflows. This is key.
The benefit here is that you don't need to know about the parse tree, don't depend on special characters to delimit elements (newlines), and don't get LCS type diffs that aren't based on any sort of semantic intuition beyond that the minimal change is probably what was meant.
Just last week, however, I had a small code change that produced a crazy diff that would give my code reviewer an unnecessary headache. I tried Patience diff on a whim and it produced a clean diff. Admittedly, it just entirely -removed the old code block and +added the new one, but at least it was readable.
I know its sort of a silly argument to make, but LCS's algorithmic intiution leaves me cold.
I've been itching to figure out how to combine perceptual hashing, computer vision like opensurf, and stable marriages to do something really interesting for image diffs, but until I can shed some less sexy responsiblities I don't think I'll get to play with that.
One example - the use of the trinary (?:) operator as a replacement for if/else assignment statement can quick and easy when programming. The problem is that diffs with it can look like a total mess because a whole lot of things are happening in that one line.
Similarly, certain languages where program flow or control structures make it so that people are inclined to make many things happen in one line (inline regex, lisp or scheme syntax lanagues) can diff in a confusing manner.
Diff is great when everyone is using a coding/whitespace standard, and things tend to atomically happen on one line.
I'd encourage that when refactoring code that's not to the standard, you do two passes - one to clean up the code to the standard, then another to make the actual changes.
Or maybe, a "quick snapshot" facility could be developed to make it easier to save off intermediate steps and commit the series of them automatically.
A "diff" is also incredibly useful prior to a commit to make sure that you changed only what you thought you did. Programmers should go to great lengths to adopt a style that is "diff compatible" so they aren't likely to miss anything important.
We currently have pep8 (python style checker) as a git commit hook, and the first few times we used it, we'd have 10 lines of legitimately modified code and then 100 lines of "style corrections."
Usually I commit my changes first, and then do another pass with style changes. Maybe it would be easier if I did it the other way around - whitespace pull request first, followed by the actual change.
Has "coding to a diff" changed your coding style?
When I know that I will have to prepare a diff for code review, I find myself writing code in a way that produces a cleaner diff. I use more vertical whitespace to collect related code into logical "paragraphs" and keep from stepping on other code when diffing.
If one believes (as say that Literate Code movement does) that code alone isn't sufficient to convey an author's intent, why should diffs alone, which mark changes in a codebase, be any different?
Comprehensible commit messages and comments explaining the purpose of a method itself (and perhaps the mechanism around which a block of code functions if necessary) are no vice and go a long way to mitigate or counteract any such "tyranny".
Seriously: in most cases it's possible to split up a 30-line patch into 5 separate patches that, applied in order, monotonically improve your code base and are easier to review, in total, than one larger patch. Changesets are cheap, and we should be optimizing them for easy review, so any eyeball we get can see what's going on.
I suppose I could create patches as I am actively solving the problem, but at that point my code may very well be a mess that I'd have to clean up each time.
Of course all of that is moot if your problem naturally segments into several patches, if nothing else than simply by the virtue of being larger than a 30-line patch or involving several mostly independent components.
git-cola is a good GUI to visually stage chunks into separate commits. I haven't really used git-cola's other functionality, but I really like its visual staging features.
If you use, say, git, you might commit every single, tiny change separately, on a private (local) branch; then use `git rebase --interactive` to reorder/merge the commits as necessary. This is the easiest way I have found, but it still involves more work.
If you wanted to take it further, you could have your editor automatically commit on every save, with a post-commit hook that takes your changes into a staging area, compiles/tests, and provides a report for each commit.
Even aside from the artifact of diff, I don't think this situation is analogous to short functions at all. The big difference is that commit logs are much more serial in nature. If I break one commit down to a sequence of 5 simpler commits they will almost always be read in that order, with context preserved.
I'm ambivalent about long functions, but am much more draconian about complex commits: http://github.com/akkartik/wart
If there were people doing this, I'd see the odd little tool to make it easier, to annotate logs and so on. But there are no signs of this.
If nobody is looking to commit histories for narrative, perhaps it's because the commit history is the wrong place for it. Since history is immutable it's really hard to have a coherent narrative of a project. Trying to do that ends up with commit messages like "final version", "really final version", "final version this time for sure", etc. My attitude has been to leave the narrative up to the reader to reconstruct. All I can do is talk about this particular point in time.
Hmm, even if the globally coherent narrative is impossible, perhaps it's worth trying to keep a piecewise-coherent history. I tend to have 'section boundaries' where I start a new feature/subsystem/narrative[1]. Perhaps I should demarcate them with "=== " or something so they're easier to see in the log.
[1] Again, I never attempt to demarcate where a feature or narrative ends, because that's impossible to judge without hindsight.
Emacs has tools that make this fairly easy with any version-control system supported by Emacs.[1]
Presumably some of the web-based browsers provided by version-control tools offer the same functionality, but every one I have seen lacks the crucial feature of re-running "annotate" again starting from the revision just before the revision that last changed line X. Otherwise you're looking at "annotate" output for a whitespace change, and to get to something useful you must manually get the log for the file, find the previous revision to what "annotate" was reporting, and run "annotate" again from there.
I've also toyed with storing documentation in commit messages themselves. For example, I wrote a blog post[2] where all the code samples in the article reflect files tracked in a version-control system, evolutions in the code samples are different commits, and the text of the article itself is taken from specially-annotated text in the commit messages. Turns out (surprise surprise!) that this is totally unmaintainable, but it was a fun exercise.
Do you have any concrete examples/writeups of your approach?
I've tried to put some tools together several times (trying to integrate with vim) without success. Mostly my approach boils down to being more aware of the commit history as I navigate. When I find myself in a new codebase I start with browsing the initial commits. Then as I go over the codebase I aggressively use git log <path>, and might drill down to look at specific, tantalizing commits.
One of my project ideas is a 'wikipedia for open source' where anybody can browse the code for open source projects both in space and in time (like emacs seems to allow), and add annotations to specific revisions. While reading, annotations from previous revisions are rendered as well, but every annotation would give some indication of age (like '350 days ago' on HN) which would help the reader gauge if it might be out of date. This would allow readers to collaboratively say, "read this snapshot first if you're new to the project" and so on.
I also try to make the logs in my projects easier to read. Here's how I got started with that: http://akkartik.name/codelog.html
IMO the big reason lots of great hackers mistrust code comments is that when you add a comment it hangs around forever by default, unless someone takes the time to decide to delete it. Attaching comments to the commit log or to annotations on a specific revision helps with this.
---
"I wrote a blog post where all the code samples in the article reflect files tracked in a version-control system, evolutions in the code samples are different commits, and the text of the article itself is taken from specially-annotated text in the commit messages."
Compare http://akkartik.name/countPaths.html :)
You can make your approach more maintainable if you give up on keeping the prose coherent. The biggest problem with documentation is that it doesn't get written most of the time. I focus on mechanisms that make it more likely I will provide that one key sentence for future readers. And -- like in wikipedia -- I assume the reader is reading critically enough to be able to handle glitches.
git diff can use --word-diff(=color) and --word-diff-regex=...
There is also the venerable "wdiff" program.
I've only found one program that can do "word patches" though: "wiggle" http://freecode.com/projects/wiggle . It works well, though I've found the interface to be slightly confusing. But turning it into a git diff and merge driver isn't that hard.
I somewhat agree with the author, that maybe there is some better change viewing paradigm we aren't seeing for these complex cases, that could benefit everyone.
Well there are multiple ways to do that. This doesn't really have anything to do with the diff but how you chose to view it.
No idea how you would define the criteria or implement it though.
A common example I see is something like a one-line list of files to build in a makefile. It is certainly possible to put every file on its own line and backslash-escape each line ending, and doing so produces a very readable "diff": if someone adds a file you see "+ xyz.c" (or whatever) and that's it instead of a mangled mess of file lists repeated with one word that's different.
Sometimes I refactor before making my changes, and sometimes I do it afterwards. In any case I try not to mix a change in the functionality and some refactoring in the same commit. That way, in retrospect, it's easier to me to understand each commit: the ones related to changes in functionality have simple, easy to understand diffs, and the ones related to refactorings have messy diffs but at least I know that they don't change any functionality.
emacs has a mode which allows one "logical" line to wrap and to be edited as many "physical" lines, but when I tested it few years ago, it was rather broken. Fortunately, for editing Latex and such, I don't really need the diffs, I'm just interested in archival.
What I do with my Latex files is have one sentence per (logical) line. I've found diffs at the sentence level much more helpful.