I wonder if this same pattern can be applied to physical engineering disciplines- e.g. a structural engineer assessing a bridge and finding numerous faults in the design, despite traffic still using the bridge as normal...
I wonder if this same pattern can be applied to physical engineering disciplines- e.g. a structural engineer assessing a bridge and finding numerous faults in the design, despite traffic still using the bridge as normal...
http://www.slate.com/blogs/the_eye/2014/04/17/the_citicorp_t...
An undergraduate doing a class project on the building uncovered the flaw.
Amazing story.
In theory, the .zip spec is broken. In practice, it's the most reliable format for transferring a group of files. (And don't even get me started on JPEG, where iirc the file format wasn't even specified until after JPEG files had been popular for years.)
> I wonder if this same pattern can be applied to physical engineering disciplines- e.g. a structural engineer assessing a bridge and finding numerous faults in the design, despite traffic still using the bridge as normal...
I would be amazed if this weren't the case. I know I've encountered a few cases where mechanical engineers had cocked up the design but the resulting machine still managed to limp along and mostly perform its function.
For those who speak only English, yes.
The encoding of filenames have always been a mess for zip files.
> Be liberal in what you accept, and conservative in what you send
There are several analyses of this maxim and whether it's a good choice for designing robust, secure systems. This particular Internet Draft doesn't agree [2].
[1] https://tools.ietf.org/html/rfc1122#section-1.2.2 [2] https://tools.ietf.org/html/draft-thomson-postel-was-wrong-0...
For M2M (machine to machine): Be strict in what you accept, and strict in what you send.
That's my take on it now.
For instance, a user will copy/paste URLs to his browser, there will probably be a space too much before or after. It's okay to clean it (and a better experience than a "site not found").
When a machine sends something "weird". Well, it's not possible to know if it's really wrong and it can't be corrected in any meaningful way. Just fail and throw an error so the developer can fix it.
I'm a fan of the strict solution as far as that goes, but there's a reason XHTML failed. Somehow asking people who write web sites to do it right or else it doesn't work is a big deal.
Well, when the job is to accept a hugely complex flexible poorly defined input format from 30 years ago that is written by millions of people who have no clue what they are doing, fixing common errors and getting anything to render is part of the spec ^^
And... oh wait!
I said "Be liberal in what you accept from the user" when it comes to user input. Web pages are user input, so yeah, browsers are no exception to the rules, they're actually a perfect example! :D
Obviously, your field of work dictates a lot here. And, if you are accepting something that has severe consequences on acting out, then yes, be more strict. However, the general principal holds. In general.
What you call a human mistake should be called a bug. The only acceptable way to handle a bug is to fix the code.
Not to be dismissive but the way you think is a classic beginner mistake. It's not the responsibility of other software to guess what bugs you'll put in yours. ^^
> Obviously, your field of work dictates a lot here.
large distributed systems, financial exchanges, trading systems, national government projects, aerospace, and even web stuff at times.
Some fields have low standards. That doesn't mean that the good practises don't apply, it just mean that people don't apply them ;)
You also have an odd misconception. It isn't having lower standards to be liberal in what you accept. It is actually harder. Much harder. And you should do it.
Easy example from finance. You shouldn't just take one currency. You should take in as many currencies as you can understand.
Does this mean to be magical? No. But you should ask, "how many different ways could this be given to me?" And you should instrument this with a marker for "unexpected."
This is a business feature request.
The "Being liberal about what you accept" is a technical guideline for protocol/format design and input processing/sanitation. Don't apply that rule to feature requests, it's not meant for that :(
It is fun to be pedantic with kids about "may I" versus "can I?" However, both work and if something can be understood, it should be understood.
So, should you build in all currencies? Not necessarily. However, you should definitely consider tagged inputs immediately, with the ability to extend the tags at some point. I'm not talking about big bang feature development, but ultimately, useful tools will have a plethora of ways they take in inputs. Because that is useful. And no design constraint should preclude this growth.
> "if you can understand it, you should act like you do"
For an unanticipated event, it's often less "if you can understand it", than "if you believe you can correctly guess the intention of an ambiguous request".HTML is a great example of that ambiguity. Take this example markup:
<h2 title="The <a href="https://en.wikipedia.org/wiki/Fall_of_the_Western_Roman_Empire">Fall of Rome</a>" class="roman_empire">The <a href="https://en.wikipedia.org/wiki/Fall_of_the_Western_Roman_Empire">Fall of Rome</a></h2>
This isn't too naive a markup error, and it's easy to see how a simple error in a CMS template could cause this output. So how should a parser treat it? It's non-conformant to the standard, but should it totally reject it? Recognise that the opening-bracket of the a tag in the title attribute is the start of the incorrect markup, and strip tags within that attribute? How should it appropriately recognise the misuse of double-quotes within a property attribute, that - according to the spec - should close the attribute?There is some merit in strict behaviour, rejecting non-conformant input. Especially when it comes to computer-to-computer API design, where rejecting non-conformant input may be better than attempting to interpret ambiguous input. Limitations can be an arrow to direct us back on the best path.
My argument here would be to do your best to not create ambiguity in what can be fed to you. The currency one is a great example. Never ever make something that takes in untagged integers as a value. It could mean too many things and you would have no method of knowing what was intended. But, once you start accepting strings, "$4" or "4 dollars", or "Four dollars" should all be on the table. (Though, yes, in this case probably best to stick with "4 USD" and complain about "dollars" being ambiguous without another qualifier.)
I'm not sure how this fits in with Postel's law, but your comment jogged my memory of it :)
IMHO it's a reflection of the software developers involved. The best tools, the ones we turn to time and time again, generally just work.
> Surprisingly, while this format is very common, it has never been formally documented. [1]
No. Excel is wrong when it comes to CSV.
Paste a Unicode string into Excel. e.g. Beijing in Simplified Chinese (北京市). Now Save As Windows CSV as beijing.csv. Close the file. Open beijing.csv. The cell now reads `___` (on Excel for Mac 2011 - maybe they deigned to fix it).
Excel just outputs bad data.
Once you add that, excel will open the file with utf8 encoding (if you use the utf8 byte order mark obviously). I haven't tried with other utf-* encodings.
Again don't know how to tell excel how to add that though :/ I've only had to deal with arabic in generated csv's.
All applications and operating systems should assume files WITHOUT a BOM are either ASCII, or the superset there of, UTF-8.
Please remember to say WHY you disagree if you do.
If you have a file without a BOM, you have to pick one.
As every 8 bit combination is an ASCII character of some kind, you can interpret every UTF-8 character as a combination of ASCII characters. And what you output will be different to what was input (unless you restrict yourself to single byte UTF-9 characters).
Without some other way of indicating the encoding format of a file, a BOM can be a tool to indicate "It's probably encoded using UTF-X".
However, the question "which?" can still apply. There are many encodings that are a superset of ASCII. UTF-8 is a superset of ASCII, but so are ISO-8859-X (for any "X"), Windows-1252, and many others.
When I've had problems in the past with this it's been around windows machines, which love their own encoding formats.
If it's merely ASCII, it doesn't matter. Nearly every charset contains all valid ASCII texts as a strict subset. UTF-7, UTF-16, UTF-32, and EBCDIC are the major counterexamples, and UTF-7 and EBCDIC aren't going to come up unless you're actually expecting them to. (Technically, ISO-2022 charsets can introduce non-ASCII characters without use of high bit characters, since they switch modes using the ASCII ESC character as part of the sequence. In practice, ESC isn't going to come up in most ASCII text and ISO-2022-JP (the only one still in major use) will frequently use 8-bit characters anyways).
The only useful purpose of a BOM is to distinguish between UTF-16LE and UTF-16BE, and even then it's discouraged in favor of actually maintaining labels (or not storing in UTF-16 in the first place). You can detect UTF-8 in practice without a BOM quite easily, and it's only Microsoft who feels obliged to need them.
I don't suppose I should get my hopes up too much that that option is going to be more prominent than the terrible default of saving as Mac OS Roman, right? (Whoever decided that Excel on OS X should export CSVs in an obsolete encoding for Mac OS Classic must have been trying to hurt Mac users.)
Unfortunately, to open the UTF-8 file again, it doesn't work ("北京市"). You have to make a new workbook and then import the csv in, specifying UTF-8 csv. It's pretty messed up but at least it's possible now.
The .zip spec is flawed, but…
I've never had an issue with extracting or using a .zip file
I'd suggest this is thanks to most software following (to some extent) the robustness principal: be conservative in what you do, be liberal in what you accept.Most of us will typically encounter fairly well packaged, conforming zip files. Occasionally we may come across something unanticipated - like this example, where HTML content is accidentally appended to the end of a zip file - and I suspect this is where we will find ambiguous behaviour: some package tools may crash, others might "extract" it as though it were content, others might ignore it.
It's this area of ambiguity that lends itself to vulnerabilities and attacks.
Re: physical engineering, I'd recommend a great book: "To Engineer Is Human". It talks about the evolution of engineering, which is a surprising amount of trial and error, with emphasis on the error.
Those dualisms can usually be resolved if you realize the vast and complex efforts that go into working around all the spec bugs - in this case, the various "find the magic number" heuristics.
A few things that I've hit in the wild that screwed up zips:
* Viruses/Worms that were rather primitive and start adding weird bytes in all kinds of places in all kinds of files
* Incomplete transmission over a modem. Depending on the protocol, you might even have most of the data, so you could read part of zip header, but not the actual archive or vice-versa. Normally, you knew right away the file was incomplete with a CRC check so not the worst problem, but the check itself was slow on old hardware.
* Weird things that added metadata and screwed with byte order, ends of file marks, etc. For instance there were some early attempts in the BBS scene to add metadata formats similar to what ID3 is to mp3s. Sometimes the writers would ruin the original file. I hit a few cases of strange attempts at steganography with software pirates trying to be "3l337" or whatever.
* Tools people tried to use to fix broken zips, that didn't quite fix them how they thought.
* Not really the fault of zip, but I've seen people rename a zip's extension to another archive format, causing the unarchiver to assume the format based on the extension. Never trust an extension if you can help it. (ex: rename a zip to rar, arc, lzz/lha, tar whatever)
* Floppy and hard disk repair programs. Sometimes these things would end up corrupting the bytes of files instead of fixing them when they tried to be more clever than moving things around. Sometimes moving things around also would result in things being ordered wrong for whatever reason in these programs. Some of the DOS Norton/Fastback/etc. ilk were especially frequent offenders.
Oh yes. Basically, EVERYTHING is flawed in INFINITE ways.
Or to put it differently, perfection doesn't exist. Luckily for us, the barrier for "being practical and useful" is a lot lower than perfection =)