Why are the Microsoft Office file formats so complicated? (And some workarounds)
joelonsoftware.com
joelonsoftware.com
And then he spends the rest of the article explaining why the file formats are so demented and impossible to read or create correctly.
Which isn't the same thing at all. Just because there is a rich history of it-seemed-like-a-good-idea-at-the-time decisions doesn't mean the end result is any good.
There is a serious, deep, and interesting problem of scaling and complexity management that could be discussed here. But Microsoft's approach seems to have been one of embracing complexity. And Joel's role, today, is just defending that approach.
I really enjoy Joel's blog, but sometimes I can't stand the clubbish-ness of the old-school Microsoft brigade. "We were a bunch of geniuses, doing the best we could..." Yeah, well, the software is usable, but it kinda sucks. Don't be surprised if the people MS victimized for a decade sound critical now that we can all see how the sausage gets made.
Here is what I think: nobody at Microsoft ever cared about making anything simpler than what they got after the first approximation, because complexity looks impressive and hence sells better. On the other hand, making things simpler requires more intellectual efforts and usually doesn't sell well, especially in the consumer product business.
Forward/backward compatibility does not require a lot of work or intellect, unless you're deliberately obfuscating your format.
Open XML is not designed for performance... it's XML, and today's computers are fast enough. It IS [supposed to be] designed for interoperability (somehow they managed to get ECMA to put their stamp of approval on it), but in reality it feels like it's a half-assed attempt to wrap all of Office's legacy formats in XML. For example, to import you still need WMF importing, because a lot of the graphics (including all clip art) are WMF.
Put another way, Microsoft didn't care about interop until they were convicted of illegal monopolistic practices and they couldn't get away with the kind of shit reflected in this spec anymore.
If a format's goal truly was to maintain forward/backward compatibility while remaining friendly toward low-end hardware, then it wouldn't be so hard to reverse engineer.
That's what most of us have been doing for years, simply to avoid the mess discussed in this article.
After reading the article, I, for one, will keep on doing it.
A,B,,,,,,,," " C,D,,,,,,,," "
That's newline there between " and "
I guess Excel loads that internally through a convertor, trims the rightmost unused columns, and then saves it that way.
Eventually, all software needs a rewrite. Not just MS's.
It's not a perfect implementation of the formats, but it's good enough for most things. And has the rather large advantage that you can run it under your favorite *nix.
It means you have to rewrite all of your date display and parsing code to handle both epochs. That would take several days to implement, I think.
That's just ridiculous. In my mind, you need a piece of code that reads the 1904 record and sets a flag in the code. Then, your date display and parsing code should all call one, or maybe two, functions which handle the conversion for you based on this flag. Thus, supporting two different epochs requires at most three components: one to read it and set a flag, one to convert a numerical argument to a date based on that flag, and one to convert a date back to a numerical argument (again, based on the flag). Should this really take "several days to implement"? It seems to me that an hour should be plenty of time.
It's going to be hard for me to continue to take him seriously if this is his view of how software should be built.