https://www.joelonsoftware.com/2008/02/19/why-are-the-micros...
The "zipped XMLs" you mention are basically XML dumps of the former binary format that evolved organically from the 1980s, when resources were scarce and they had to hack together a working office solution.
• Create a simple Excel document.
• Clone the document and change the text value of one cell.
• Unzip both .xlsx files into two different directories.
• Now launch Meld/WinMerge or similar and diff the directories.
Now tell me if you still think diffing this format is trivial.
(Have written streaming XLSX parser in the past.)
AFAIK there are no ready-made solutions for that so far. Would be very useful![2]
[1] It would be interesting to dive further in to this subject but personally I can’t currently find the time for that.
[2] Now that I think of it, this might be an interesting project for someone participating in Google Summer of Code. Not sure if the Git project will participate this year or not.
Sure, it's hard to diff and merge tree data structures, but it doesn't have to be perfect. Text diffing and merging is already imperfect anyway, yet it's very useful.
Not sure why we would need this at a file system level. You’d need diff tools for all sorts of file type.
If you would simply render each version to plain text and compare them (which is a solved problem), it would already be very useful.