Bram Cohen: I have a question for the version control experts
facebook.com
facebook.com
http://stackoverflow.com/questions/14077470/is-simplified-se...
Is archive.org even allowed (by Facebook's robots.txt) to archive this?
It's sad if technical people post important information on Facebook
Edit: facebook's robots.txt has this: User-agent: *, Disallow: /. :-(
On the other hand, there's also possiblities such as; Facebook losing or deleting the content, Bram Cohen deleting the content or his account.
"Page cannot be crawled or displayed due to robots.txt."
So no, the Internet Archive will most likely not archive public Facebook posts - because of Facebooks robots.txt. The IA respects and honours the domains robots.txt.
I've taken a snapshot of the post with my own script which will put it into a IA Wayback Machine friendly format (WARC).
If nothing else, torrents with version control would frankly be mind-blowing. Think of a single TV show torrent, which could be updated again and again. Then think of TV broadcasting and content distribution of the program being shown through such an above torrent and an overlay, and ads, being played in on the client-side?
Thinking away the rights issues and DRM, this might be the future of live TV and content distribution.
Think of it like sailing when the mast breaks, you don't who last touched the mast, you care how to get the mast back up.
In truth the totality of world events leading up to the moment the mast broke is what lead it to break. This is not important, what is important is how to right it.
If the question is viewed in terms of righting the ship then the point is to identify the last revision that worked, conversely the 'broken' revision is the next one.
Realistically, to solve the problem you'd want a genetic algorithm that could turn on and off each change in each revision to find the combination that yields the most fit iteration of the code.
As far as the stack exchange comment that it's useless because what you usually see is white space fixups, (a) that's why I fix up the whitespace before I commit the final patch into the source tree, and/or insist that those problems be fixed before I pull a commit, and (b) this is why a strongly discourage whitespace-only cleanup patches. (Or checkpatch.pl style changes, in general.) With a little discipline, "git blame" can be incredibly useful.
Blame information is useful, not for pointing fingers, but for reasoning about changes. Often, I'll be browsing code, see something change, blame that line, then go and read that entire revision so I can find out why the line was changed. Normally for me, my goal is to find out the new argument ordering or whatever so I can continue writing my code using the new pattern.
But there are plenty of occasions where it is, because the changes broke the system. A broken system is not going to sit around for 10 years.
There are plenty of times where I want to find the changeset that last touched a line of code so I can read its commit message, find an associated bug tracker ticket, etc. What do these uses have to do with the last working revision of a project?
As others have said, it is often very important (at least when working in a collaborative manner on real-world complex code bases) to work out who added code and more importantly why it either exists or was written that way. More likely than not (assuming it's a fairly high-quality code base) there's an edge case somewhere that it's dealing with.
Bisecting to the revision before it broke doesn't tell you anything about this.
Ideally, there should be loads of comments too, but then, that doesn't happen nearly as often as it should in the real world...
else if {
but I don't think that matters because you are unlikely to care where such lines come from. (You could fix this anyway by doing unique lines first, then using the blame of those to make better choices for the non-unique lines.)See also the more complicated variant at the end of this post from the git mailing list:
Version control systems are not storing what actually changed between files, they are storing the smallest set of differences between them.
IE Given two versions of the same file, and the history graph A->B they store how to reproduce the bits of B from the bits of A.
This is completely unrelated to how B actually got that way. So in turn, they use textual diff algorithms to approximate how B was formed from A.
Even moving back into the text world, there is still no right answer. It only tells you one of the possible ways that A was transformed into B. It would be perfectly valid for the text diff algorithm to say "every line in A was removed, every line in B was added". This in turn would give you a blame that pointed to that rev for everything.
Most textual diff algorithms "try" to do something sensible, but blame is essentially trying to turn applesauce back into apples.
Even git's more "advanced" blame can be completely messed up by the internal text diff doing dumb things.
All that said, one of the reasons you may like it is because it's basically what everyone actually does.
Fractal Designs Painter (now Corel Painter), back in the day, used to have a feature where the entire creation and edit history of a "painting" could be captured and replayed in full detail, right down to the level of the brush angle and pressure used by the artist with every stroke. This was useful for creating art at low resolution, then re-creating the piece automatically (and much more slowly) at high resolution. Corel's version may still have this feature; I haven't used it in years. But if it was practical ten years ago in manipulating multi-hundred-megabyte files, certainly it would be possible now, for what are usually plain text files.
Instead of just doing a diff or git bisect, imagine being able to load up the commit of a file with one of those ambiguous changes, grab a slider (or your favorite keyboard equivalents), and scrub back and forth through a condensed replay of someone else's actual changes exactly as they were keyed.
I'm not sure this would always be a good thing (a form of surveillance?), but it would certainly be useful when the original author of the code is unavailable.
Many version control systems work that way, but there are other possible strategies. For example, Darcs' patch theory is very much based on recording what actually changed in each revision (and consequently achieves better results in some awkward cases than a purely text-based VCS).
Besides not having the goal of figuring out what changed, they often are heuristic and give up (IE they stop trying to align the original files, and just say "removed here, added there").
I think that's a little unfair. Darcs does store what actually changed between the files, at least at the points they were committed, in a qualitatively different way to the flattening effect of cumulative commits in a system like say Git.
Of course, you can defeat even that approach if your edits from one commit to the next are ambiguous, for example if you have two verbatim copies of something next to each other where you had only one before, and this will subsequently lead to ambiguity if you try to merge a change from someone else to the original copy since the merge has no way to determine whether the first or second duplicate (or both) should be modified.
It sounds like you want something that is directly tied into your editor, so it is aware of changes between commits, or perhaps something that has semantic understanding so that instead of recording half a dozen text edits, it records "variable foo was renamed to bar". Tools that worked on that level would be fantastic to work with, but until you've swapped a text file representation of code for some sort of database-backed semantic model and your edits/refactorings can all be expressed in terms of that model, I don't see how any VCS could possibly achieve it.
As for the rest, i don't want anything. I just don't pretend that textual displays of blame/diff/etc are actually showing me a correct history, instead of one possible history of textual changes. If it says bob wrote/changed some code, i don't assume that's really correct (unless of course, the changelog says "wrote code" :P), I ask bob.
I will point out that historically, there were version control systems that were integrated like you describe, even for C++ (IBM Visualage C++). But people are happy with what things like git/svn/etc provide, and that's fine by me.
Remember that I worked on a system whose sole goal was to provide a qualitatively better experience than CVS, so I don't have very high standards :P.
Please correct me if I'm misunderstanding, but Bram is proposing to do blame without using the diffs, a nice simplification and not something I'm aware of other version control systems doing.
Revision 1: x
Revision 2: xx
Which x was added in revision 2?In real life, you get things like replacements of a statement by an if-else block with statements similar to, but not identical to, the original statement in each branch. Are there still parts of the original statement there? If so, in which branch(es)?
Also, suppose I add code in version 6, you remove it in version 14 and someone else resurrects it in version 23. What commit contributed that code? How can you know whether that someone else resurrected the code, rather than write a new copy?
Finally, what is 'contribute'? Does a commit that 'only' moves lines around in online help contribute to a file? What does it contribute? How do you show that contribution in a diff? What if the move also necessitated some minor changes like adding punctuation?
Your RSS reader would simply merge with your news feed.
Do not trust.
Even if Facebook somehow managed to deploy Wordpress or wiki quality content systems, I still won't go there.
If you host your own Wordpress install, the ad dollars go to you.
If you're Mark Zuckerberg, bloggers blogging on Facebook is better than bloggers blogging elsewhere. That's the circumstance under which blogging on Facebook is better.
I'd certainly run my own Wordpress, but not everybody wants to.
I suppose I just connect more so with a facebook profile than with a traditional blog.