Hey Dropbox, why can't I compare file versions like this?
web.eecs.utk.edu
web.eecs.utk.edu
In all seriousness, though, I wanted this so badly that I started (and failed) a startup with 12 employees nine years ago to build it. It was conceived for use in working with "big data" but the system essentially provided Etherpad-like scrubbable versioning of all common office document formats as a side-effect. All of it was structured in an environment more similar to the social aspects of Github than Dropbox, but you could sync up to your filesystem via a FUSE wrapper. That is, people could easily follow or fork your work in progress. If we'd continued, you'd have been able to accept the equivalent of PRs on your Word docs.
It was so awesome that we couldn't find anyone to pay for it, sadly. Armchair quarterbacks would fairly accuse us of failing to do proper customer development.
I can't speak to the technical limitations of Dropbox's versioning implementation, but given that they already have both viewing AND versioning running for a decade, I honestly can't believe it would take more than a few months for a small team to implement Etherpad-like editing functionality for the office suite document formats.
If it was mine, people dramatically underestimate the effort and cost of supporting the OSS release of something like this.
It was also written on what I'm guessing was Rails 3.2 (circa 2011). There was no Docker, no Kubernetes; it was a monolith with many moving parts.
I also don't have access to the source code anymore, not easily.
It's also true that the company that we started continues to exist with a new product, and I suspect that some of our original code lives on.
It's just super messy.
I see some issues with it. No matter how smooth you make this, any software with pull requests is going to be considered "technical". Christ, people think basic excel skills are "technical". SO you have to get over that hurdle. But personally, if I've already gotten people over that hurdle, I might as well just use git and LaTeX documents.
I don't know, I think it sounds awesome, but I also think it might be tough to sell.
Our mistake, as I said elsewhere, is that there was no market. We had many sales conversations and zero takers. For context, we were fully under the spell of The Great Big Data Hype of 2011 (see the current AI hype for reference) and convinced ourselves that there would be so many opportunities that we'd have our pick of which path to take.
In fairness to our past selves, for a while this seemed true; valuations were insane for companies with vague value props in the space. And we met with dozens of influencers in the data world and they all professed to be excited to use it. Most of them ultimately logged in once, realized they had no hair-on-fire problem for us to solve, and stopped returning our emails. It was frustrating in the extreme.
I just found one of our product videos on Youtube: https://www.youtube.com/watch?v=EWMjQhhxhQ4
git wasn't yet supporting large binary commits very well (not that it's amazing support today) but I remember us going very deep on this python library that fudged support for large binary repos with the glorified equivalent of symlinks. I'm not sure we ever really got this working well and my memory is honestly fuzzy enough that it's getting harder to sort what we did from what I desperately wanted to do before we ran out of cash and I got depressed enough to take two years off and run away to Europe with a crazy person.
Humans are funny.
> It was so awesome that we couldn't find anyone to pay for it, sadly. Armchair quarterbacks would fairly accuse us of failing to do proper customer development.
I get the impression that it was extremely complicated and didn't fit into anyone's workflow. IE, if you approached someone using Dropbox, they'd have to change far too many habits just to switch to you.
It's actually far more appropriate to say that we were competing with a culture where it's engrained in people to make those FINAL_FINAL2 versions on Samba shares. Or worse, to email them.
I am biased but I'd give our UX a 7.5/10, and if we'd have continued it would have gotten smoother. The FUSE wrapper was not the primary interface by a long shot, though... in fact, I'm not sure it was used by many people outside of our team, in the end.
> Please tell us something surprising or amusing that one of you has discovered. (The answer need not be related to your project.)
> The ridiculous things people name their documents to do versioning, like "proposal v2 good revised NEW 11-15-06.doc", continue to crack me up.
https://www.ycombinator.com/apply/dropbox
And yet here we are a decade and change later and Dropbox, while having solved "a" problem, sits like a ridiculous behemoth leaving it's users hungry for so many other pain points to be addressed by another savior, including especially this one problem they said they're gonna solve.
I introduced a lot of people to Dropbox like 8-9 years ago and after using it to share files with other people I found out the hard way it's a terrible tool for that. I then used it for a couple more years to share files between my machines but they haven been introducing so much crap in their desktop app that I moved to sync.com.
The reason they're not doing it is because they want a piece of the productivity pie.
They're not getting it from me. Ever.
As OP said, it took a lot of effort to get the UI ok. Probably takes even more effort to create a great UI, but I guess Dropbox has some resources, right? Shameless plug: Landing page at https://julesdocs.com
If anyone is interested in pushing this forward, I'd love to hear from you (mail address on the landing page)!
And I will never ever forgive y'all for what you did with mailbox! (Like seriously what did they do?)
Enterprise customers pay more money to have more features with checkboxes in the feature matrix. Telling them they don't get a feature because their needs are too complicated is a tough sell. (something, something, opens up opportunities for low-end disruption, something, something)
Sadly, we never developed it.
But, yes, I totally agree with your rationale. This is the kind of rabbit hole that can quickly turn into a distraction.
EDIT: This comment is almost a word salad, I need to sleep lol.
In other words, someone wants a diff tool for Microsoft products but specifically wants Dropbox to implement it.
ssh user@rsync.net diff some/file .zfs/snapshot/yesterday/some/file
We can implement this later today if it sounds useful ... I was sort of surprised that 'diff' was not already a whitelisted command[1] ...[1] https://www.rsync.net/resources/howto/remote_commands.html
You may optionally set any arbitrary schedule you like (day/week/month/quarter/year) and you simply pay for the bits on disk that those (efficient, changes only) snapshots take up. Sometimes they take up almost nothing.
My favorite part of all of this is that the snapshots are immutable, or read-only. No matter who attacks your rsync.net login or what password you lose, the snapshots cannot be destroyed by any outside action.
This allows for some interesting insurance against ransomware / Mallory ...
Any other way to implement would probably also involve one common format (or a small few) behind the scenes.
Dropbox could just buy one of these companies and work on integrate the solution with its platform.
All arguments about the complexity of this feature are bogus when it has been solved several times by different vendors over the last decades.
One such tool found via DDG: https://draftable.com/compare
Adobe Acrobat also has a diff (including visual diff) feature that can be used to do advanced comparisons if necessary.
Granted, author’s suggestion is more user friendly and integrated.
It does? Didn’t know that. How does one activate it?
You can do "Combine" which is effectively a merge interface.
I wrote up about it here for the curious: https://www.vivekkalyan.com/using-git-for-word
> ...but this is useless. Timestamps??? Tell me what changed! Let me see the changes over time. Word has a change tracking feature, but my PhD in computer science isn't enough for me to figure it out.
> But but but Austin, you should be using a proper version control system! Just use Git and GitHub!
Found that aside curious, as track changes in Word is a first class versioning implementation with word processing and editors savvy, just as Git is a first class versioning implementation that's code lines and commits savvy.
Surely headspace around track changes is less "PhD" than git.
Have you tried track changes? :)
Imagine how github could push for MS Office integration and become a versioning powerhouse for non-code-stuff.
But I can't see it standing as a stand-alone product that people would really pay for. It has to be part of something else.
IMHO, if it could be smoothly integrated to e.g. Git then there would probably quite a few companies that would pay good money for it.
The solution is to dump Office and use text files, if you can get away with it.
[0]: https://etherpad.org
[1] https://en.wikipedia.org/wiki/Document_comparison
[2] https://Last10K.com/compare.gif
[3] https://www.bloomberg.com/opinion/articles/2018-05-22/10-k-c...
Update: The site in reference is https://Last10K.com
https://www.joelonsoftware.com/2008/02/19/why-are-the-micros...
The "zipped XMLs" you mention are basically XML dumps of the former binary format that evolved organically from the 1980s, when resources were scarce and they had to hack together a working office solution.
• Create a simple Excel document.
• Clone the document and change the text value of one cell.
• Unzip both .xlsx files into two different directories.
• Now launch Meld/WinMerge or similar and diff the directories.
Now tell me if you still think diffing this format is trivial.
(Have written streaming XLSX parser in the past.)
AFAIK there are no ready-made solutions for that so far. Would be very useful![2]
[1] It would be interesting to dive further in to this subject but personally I can’t currently find the time for that.
[2] Now that I think of it, this might be an interesting project for someone participating in Google Summer of Code. Not sure if the Git project will participate this year or not.
Sure, it's hard to diff and merge tree data structures, but it doesn't have to be perfect. Text diffing and merging is already imperfect anyway, yet it's very useful.
Not sure why we would need this at a file system level. You’d need diff tools for all sorts of file type.
If you would simply render each version to plain text and compare them (which is a solved problem), it would already be very useful.
It’s obvious UI on the level of pinch to zoom and mouse input. Hard to come up with but obviously the right choice once suggested.
[1] https://stackoverflow.com/questions/23140240/git-how-do-i-ad...
add x y = x + y
left B: add z y = z + y
and right C: add x w = x + w
succeeds without a conflict?But the problem with that idea is that two different people explicitly made a change that looks meaningless. That tells us that we're evaluating "equivalent" incorrectly, which means we don't actually have any remaining justification for picking one over another, and the conflict is hopeless without further input.
add z w = z + w
My justification is that if you put each identifier on a separate line like this: fn add(
x: i32,
y: i32,
) -> i32 {
x
+
y
}
then as far as I know, Git would happily merge B and C into D.For example... .docx sometimes actually contains binary blobs...
Unzip a few MS Office documents and have a go yourself.
That was an inspiration for a tool I built called Yestercode [1] (though it uses undo history, not version control).
[1] http://web.eecs.utk.edu/~azh/pubs/Henley2016VLHCC_Yestercode...
Also Zoho Writer has a combine feature, that lets you upload a docx and combine it with another docx - with the changes highlighted as tracked-changes. Pretty handy for comparing docx files.
I don't think I'd want a scrub bar like that though, maybe? I suppose I've never tried it.
OHTF Vg'f uneq gb hfr guvf pbzznaq jvgubhg fvatvat.
--
*often (not in MS Office probably) the undo buffer is managed by the OS. It’s conceivable that some rethinking could happen at the OS level.
It supports docx, excel and whatnot. But .txt file? That's too complex.
In short, it is probably possible now, using what DropBox already exposes.
Are you sure? I’ve never heard about that.
macOS includes a built-in version history since OSX Leopard. Sadly the flashy version UI with 3D effects is not the best to find differences and many programs doesn’t use the native frameworks that bring this feature (e.g MS apps, Adobe apps)
I had no idea. Do you have any pointers for further reading? That would be very interesting.
http://osxdaily.com/2015/06/16/revert-to-prior-version-file-...
https://blog.macsales.com/47415-auto-save-and-versions-an-of...
I totally agree with the comment from dreamcompiler, the file history feature is a great idea but the execution needs a lot of improvement.
The main issue to me is not the change to Save as.. behavior. To me the problem is that most of the apps didn’t adopted it. In particular cross platform apps ignore it. So you never get used to the behavior change.
The version history UI, is also too “heavy”. It has a slider to go back in time, but surrounded by a faux app window simulating traveling in time with your app state... sounds cool but is distracting and not so useful to find differences.
And yes, it works with iCloud, but only in apps that use auto save APIs: https://developer.apple.com/design/human-interface-guideline...
- FileMerge is a standalone diff tool inside the Xcode bundle.
It's pretty useful and impressive in its current state but isn't being actively developed from what I understand.
I avoided both for as long as possible, but we switched to office365 at work and the integration between the two, and teams, is pretty great tbh.