How a load-balancing bug led to worldwide Chrome crashes
code.google.com
code.google.com
To my knowledge, Firefox doesn't do that. Safari doesn't do that. Internet browsers are probably the #1 most important app on a computer these days, browser reliability is vital.
I think it's a bit sensationalist to still refer to it "being crashed" at any time, rather than saying it "may crash due to a bug".
It is not a design flaw. Sure, this specific vulnerability would not be there if the remote sync feature wasn't there, but people like features.
Chrome has a pretty good security track record. I'm not worried.
try {
// Do syncing things.
} catch(...) {
// Continuing browsing.
}Trust me when I say you don't want to be writing code in a stack frame above someone else who catches and ignores exceptions using them.
There are actually some cases where windows will catch and discard segfaults if you have them in response to certain window messages. That bug was hard to track down, let me tell you.
2) Any proposed design choice must be implemented with code, and that code can have bugs and crash. That is what happened here.
It's a design choice that your product... which some people (myself included) pay money for, crash hard so that you can get better diagnostics? Sounds like misplaced priorities.
This sort of catchall and keep going error handling could leave your browser in a completely unknown state. It could start making bad requests, making the wrong requests, start leaking info, or more likely crash elsewhere but with a much less clean crash log.
When you don't know how to handle an error such that it bubbles all the way up to the top, often the best thing to do is crash. At least then you might get the logs that allow you to fix it and turn around the fix quickly and with confidence.
Crashes suck, but crashing and not knowing why sucks more.
(Incidentally, eliminating these kind of rarely executed branches is a bugaboo of mine. They frequently have problems.)
How do you pay for Chrome?
Well, first, the sync code might not need to be in C/C++. It could be in a sandboxed, safe language instead. Other browsers do that.
Or, at minimum, it could be in C/C++ but at least in a sandboxed side process, Chrome has the capabilities for that.
Of course, no one would argue otherwise. Every time you decide how much to reduce it, you define a tradeoff in terms of work vs. benefit.
Because it's entirely possible that Firefox or Safari, for example, could have been crashed by contacting the safebrowsing server, and the safebrowsing server returning an answer that crashes it.
Firefox also does remote firefox update checks and plugin update checks, etc.
None of the browsers you mention are "independent" of internet servers anymore. They are meant to function independently, as is Chrome, but exactly the right remote bug could likely crash all of them.
It wasn't a command that shut down Chrome, it was just poor handling of an edge-case which resulted in unexpected behavior (crash).
It's a sad truth that most programs will explode if you fling garbage at them. When push comes to shove, many development timelines don't have room to bulletproof against everything
We back up so much of our tools and data, but without a working browser, we're sunk. Especially non-technical people.
That a huge percentage of the internet clients in the world can be simultaneously removed from accessing the internet, either intentionally or accidentally, is troubling me this morning.
It's pretty ridiculous of you to point at a single mistake in implementation and blame the entire sync feature.
Remember that most JS is popular JS, with some small amount of custom lines. JS seems less crashy because people don't use as much "random" JS in general.
It isn't by design that syncing can affect the whole browser; it's a bug in the syncing code which should have been handled. There is no self destruct bug, and calling it that is incorrect. Are you aware of the fix?
http://src.chromium.org/viewvc/chrome/trunk/src/sync/engine/...
The fix is just checking if the model is valid before making the call which was throwing the out of bound exception.
It seems like just the other day when I was thinking about this very thing. http://rachelbythebay.com/w/2012/11/19/lb/
It might not be the greatest article, but it highlights a very real and new point of failure that cloud apps are introducing. I was hoping for a good discussion to check out later.
the bug report describes all that, roughly (if i've understood) but doesn't seem to be worried about the higher level issues - the inconsistent types and need for fragile human assertions about type logic.
(not java bashing - don't see why this couldn't be solved in java)
but i guess this is just a bug report. for an outage like this i suppose there's going to be a major review? is that all internal? would be interesting to watch.
Chromium is c++. The code diffs in the bug report are c++. Where is Java coming into picture?
(ps i was talking about the code on the server side; not chromium)
Oh I see.
> on the server side, it looks like the problem could have been avoided with better types - it seem that there was a confusion between status values that can include 0 and those that cannot (alternatively, perhaps better, there was no status for the case where the status was undefined?)
From what I understood, they are talking about protocol buffer types. The sync server sent message to chromium to throttle for all types say A, B, and C. Chromium didn't know about type C, some code returned 0 for unspecified type, another piece of code calculates index based on what is returned in the previous step, and then due to the 0, a negative index was accessed in the bitset leading to out of bound exception.
The issue is server sent all types to clients rather than sending only types known to the client, and the client didn't gracefully handle unknown types. I don't think a better type system would have helped the server - that looks like a logic bug, not a typing bug.
if you can do that then you can get the "infallible" compiler to provide the "this branch will not be executed" logic. but there may be efficiency trade-offs, or it may be so directly tied to low level protocols that it is impossible.
five ten years ago i would not have thought of the problem in this way - i would have agreed with you (and the bug report) that it is just logic. but slowly i am starting to learn to rely more on types. but i don't know enough here to suggest details...
or
Contribute to Chrome development.
wheres is the choice? look like these days using anything software or hardware from the tech giants means to be their pets
(It was because I'm worried about just how much information Google is collecting on me, so while I can feel smart for not having a crashing Chrome yesterday, really I was just lucky. There but for the grace of God.)
but sure it will respect your data, privacy and civil rights.. This Big Brother thing (under our backs) must stop.
Why the downvote?