Retrospective and technical details on the recent Firefox outage
hacks.mozilla.org
hacks.mozilla.org
It's hard to hate something once you truly understand it, I guess.
As to why, dunno, presumably it's extra effort for an unclear gain. Telemetry code wouldn't be parsing hostile input etc. And it doesn't stop bugs like this either.
In this case, that the problem was triggered by telemetry was a coincidence; it would've been triggered by other bits of coden in the future as they moved to Rust.
To me, this is a perfectly valid write-up with a good lessons learned. They have written it in a very diplomatic way, but to me, it is absolutely clear that Google screwed up here. How can you make such a change to a default behavior of critical infrastructure unannounced? That's just reckless towards your customers, and solidifies my belief to stay away from GCP.
If they had properly announced the change, even if the Firefox team hadn't then tested beforehand, at least the DevOps team would have put one and one together and just changed back to HTTP/2 and the outage would have lasted maybe 10 minutes. Instead, they frantically went through their git log to see what in the code base might have triggered this bug. Everyone who has been in such a position knows how incredibly stressful this is. I'd be absolutely livid at Google in their position. That it took two hours to fix this is clearly their fault.
It isn't. The bug was in the networking stack, and it just happened to be triggered by a GCP change which effected the telemetry service. Firefox having telemetry has nothing to do with the issue here.
"The amount of blame that is assigned to the Firefox team is staggering"
Telemetry is different to user traffic - it's less important! - but of course any in-process QoS would still create a point of interaction with user traffic.
> It just so happens that Telemetry is currently the only Rust-based component in Firefox Desktop that uses the [viaduct/Necko] network stack and adds a Content-Length header. This is why users who disabled Telemetry would see this problem resolved ...
The article contradicts your conclusion. If Firefox did not have telemetry, the bug would have had no impact, and users would not have suffered an outage.
> ...even though the problem is not related to Telemetry functionality itself and could have been triggered otherwise.
And then the article contradicts itself and agrees with you using some heavy-duty doublethink. Sure, if there were hypothetically other Rust services using the buggy network stack, they'd also have hit the bug: BUT THERE ARE NONE. The bug was in code which is only running because it's used by the telemetry services, so even though it might be in a different semantic layer it's the fault of the browser trying to send telemetry.
As a user, I place very low (often negative) importance on the tools I use collecting telemetry data, or on protecting DRM content, or on checking licensing status. They should focus on doing the job I'm trying to do with them on my computerr, serving the uses of the user, rather than doing something that someone else wants them to do. Sure, I understand that debugging and quality monitoring are easier with logs and maybe with telemetry, so I can understand using a few resources in the background to serve some of that data, but it must never get in the way of actual work getting done.
This is your mistake: as explained in the article, it could have affected any component. Telemetry happened to hit it first but anything using HTTP/3 with that path would have been affected.
“This is why users who disabled Telemetry would see this problem resolved even though the problem is not related to Telemetry functionality itself and could have been triggered otherwise.”
Is this really relevant, though? To the users who were unable to use their browsers normally it doesn't matter that this problem could have occurred elsewhere as well, but rather that it did occur here in particular.
If particular sites would break, then that could be debugged separately, but as it stands even people who'd be perfectly fine with browsing regular HTTP/1.1 or HTTP/2 sites were also now impacted, not even due to opening a site that they wanted to visit themselves, but rather some background piece of functionality.
That's not to say that i think there shouldn't be telemetry in place, just that the poster is correct in saying that this wouldn't be such a high visibility issue if there was no telemetry in place and thus no HTTP/3 apart from sites the user visits.
Surely one could adopt a "shared nothing" approach, or something close to it - a separate process for the telemetry functionality which only reads things from either shared memory or from the disk, where the main browser processes could put what's relevant/needed for it.
If a browser process fails to work with HTTP/3, i don't think the entire OS would suddenly find itself not having any network connectivity. For example, a Nextcloud client would still continue working and synchronizing files. If there was some critical bug in curl, surely that wouldn't necessarily bring down web browsers, like Chromium, either!
Why couldn't telemetry be implemented in a similarly decoupled way and thus eliminate the possibility of the "core browser" breaking due to something like this? Let the telemetry break in all the ways you're not aware of but let the browser continue working until it hits similar circumstances (if it at all will, HTTP/3 isn't all that common yet).
I don't care much for flame wars or "going full Stallman", but surely there is an argument to be made about increasing resiliency against situations like this one. Claiming that the current implementation of this telemetry is blameless doesn't feel adequate.
Which is exactly how the code was intended to work. Firefox did not design their software to hang in the event of telementry losing internet access.
I don't know firefox's internal architecture or its development, what follows is pure conjecture.
Their intention seems to be to slowly migrate the codebase from C++ to Rust. That telemetry is the only function to so far rely on their new rust networking library viaduct (and thus trigger the bug) could be because they wanted use their least important fucntionality as a test bed. In which case, if there wasn't any telemetry, a different piece of code would have been migrated to rust first and triggered this same bug. Without the telemtry, it would have presumably taken them longer to realise that things had broken, let alone resolve it.
So you're saying that Firefox did not on fact have an outage due to a change in their telemetry servers? That's not what the article said.
I understand that you mean to say that it isn't intended for networking to be taken down by telemetry. That's nonetheless what happened, and it could have been prevented by treating telemetry as a different class of traffic (not collocating it with normal requests), or by not having it, as others point out.
“This is why users who disabled Telemetry would see this problem resolved even though the problem is not related to Telemetry functionality itself and could have been triggered otherwise.”
> So you're saying that Firefox did not on fact have an outage due to a change in their telemetry servers?
Not the telemetry code. Not the fact that it "could" happen elsewhere. But rather the fact that it was in place and in this instance happened because of it.
Not that it matters that much. Regardless of the particular cause, a browser failing to work because of something changing externally is crazy (at least to me), no matter how you look at it.
Edit: this is now largely a duplicate of the other comment, hmm: https://news.ycombinator.com/item?id=30179023
It's natural for all the network stuff that goes in inside a browser to share code. You can say what you want about telemetry (I'm not a huge fan, personally), but this was a dumb bug and it is completely unreasonable to expect some kind of adversarial design "just in case a freak bug triggers on telemetry network requests".
I absolutely agree that this a dumb bug having little to nothing to do with telemetry. It is not even the first case-sensitivity HTTP/3 bug I’m personally encountering in the course of completely casual use[1]. Probably not the last, either, those joints ain’t gonna oil themselves.
At the same time, you know what? I’m glad you suggested this, because I certainly didn’t think of it. Yes, in an ideal world, telemetry absolutely should be a separate process (or thread, or at least not share an event loop—a separate “hang domain”, a vat[2] if you want). And so should everything else off the critical path.
I’m not saying Firefox is bad for doing it differently. I’m saying it’s silly that Firefox is forced to play OS to such an extent because the actual one isn’t up to its demands.
I read that at first glance as
> Probably not the last, either, those joints ain’t gonna roll themselves.
and thought, hm, I need to remember this debugging technique next time I'm stumped.
> In the coming weeks, we’ll bring HTTP/3 to more users when it's enabled by default for all Cloud CDN and HTTPS Load Balancing customers: you won't need to lift a finger for your end users to start enjoying improved performance.
In their blog on June 22, 2021. [1]. It probably should have been it's own standalone message sent to users (a "this should be a no-op" email), bit to claim that it was unannounced is misleading.
1. https://cloud.google.com/blog/products/networking/cloud-cdn-...
The clone failed with a mysterious error. After some minutes I checked the accompanying web site. The web site failed too, but, on refresh, this time I got a holding page explaining that the service was down. So I check the overall ticket system, and I find a change ticket, for the git system, saying there is planned maintenance, at 8am for one hour. Unadvertised because hey, it's 8am, most people aren't at work at 8am and this is a regular (Wednesday 8am) maintenance slot.
And I scroll down and I find that nobody remembered to actually do the task. They wrote it up, submitted, got it OK'd and then, eh, never did it. By the time the people who were supposed to do it were reminded it was 9am already. So, astoundingly, the service owner OK'd just doing it after lunch instead.
That failed 8am change was actually a re-run, of a re-run, of a re-run, of an upgrade that keeps failing and definitely takes over an hour to complete.
So instead of "It's fine to do this when nobody is at work and it's low risk" suddenly "It's fine to do this for 2 hours in the middle of the working day, though it'll probably fail and we have no roll back plan".
That's pretty shoddy. Glad to know an "Enterprise" cloud offering is hardly better.
From the article:
> This is why users who disabled Telemetry would see this problem resolved even though the problem is not related to Telemetry functionality itself and could have been triggered otherwise.
Don't spread FUD.
Of course all network access is shared for a machine so it's not possible to not have a single point of failure, but there are different ways of slicing up the access.
That was my thought after reading the start of it. Like "Oh no, Firefox has fallen into that void where their need for telemetry trumps users". Another product falling down at doing its primary function. But after reading the entire report that's just not fair at all. A bug relating to telemetry and their network stack caused failure in that networking code which affected everything. That is entirely different than software depending on telemetry to function properly. It wasn't by design that failing to phone home broke the software, it really was just a bug - a fairly obscure one. Sounds like if someone wanted they could just as easily blame the use of Rust in Firefox since some of the code involved was written in Rust. But that's not a fair or accurate conclusion either.
This seems like you're embellishing this part to tell a story? It's not supported by the linked post, and from the bugzilla bugs it seems like it was known almost immediately that the ESR builds were affected as well and so it almost had to be an external service, they just weren't sure which one at first.
Mozilla has opened themselves up here, as they market as a privacy and user respecting alternative, so when they fail to live up to their own marketing people are more annoyed while they expect random startup #456 to not care about their users privacy and have telemetry out the wahoo.
The way Mozilla does telemetry is different from how most places do.
I think that the biggest issue with these discussions is that there always seems to be this assumption that there is only one way to do telemetry, it always contains super invasive PII, Mozilla's telemetry must do the same, and therefore Mozilla's telemetry is just as evil as anybody else's.
Mozilla is remarkably open about how its telemetry works, beyond just being open source. Maybe this is more a problem of that information not being surfaced well, I dunno. I get that some people are philosophically opposed to telemetry no matter what, but I have seen enough cases of, "Wow, I didn't know it worked that way, and I'm actually okay with this," to know that informed users are not universally opposed to it.
If you're running Firefox, you can go to `about:telemetry` and see what data is there. Note that some of that data might be populated even if you have telemetry turned off. Don't reach for your pitchfork quite yet: I assure you that the data isn't being sent.
You can go to https://telemetry.mozilla.org to see various analyses involving telemetry data.
In the source tree, telemetry probes must be defined in one of a few places. Here are their locations:
https://searchfox.org/mozilla-central/source/toolkit/compone...
https://searchfox.org/mozilla-central/source/toolkit/compone...
https://searchfox.org/mozilla-central/source/toolkit/compone...
https://searchfox.org/mozilla-central/source/toolkit/compone...
> Don't reach for your pitchfork quite yet
No need to assume people are out of their minds or even critical. I am curious about how it's done on a technical level, with the old local ad system in mind (which I thought was a brilliant solution to Internet commerce and privacy). I've supported and contributed to Mozilla since before Firefox.
As for high level docs about how it works, I haven't been involved in quite a few years, so I'm not 100% sure about the best source, but this link looks like a good place to start:
https://rewis.io/urteile/urteil/lhm-20-01-2022-3-o-1749320/
HN discussion: https://news.ycombinator.com/item?id=30135264
More importantly Mozilla knows that there are people who do not want them to upload this information yet they continue to do it anyway by default without ever asking for consent. Worse, Mozilla keeps adding new leaks that concerned users will have to watch out for and disable after each update. This is of course by no means a problem unique to Mozilla - the software industry as a whole has not yet learned that no means no - but it is also a Mozilla problem and as long as they want to use privacy to market their software they will rightfully receive the loudest criticism. Thankfully laws are beginning to catch up with the digital age and people will have better recourse than asking software vendors nicely to not be mistreated.
They set themselves a higher standard by marketing as the good guys who fight for the user, and then made any number of moves that said users viewed as not being in their interests. Of course they get more blame. Like, Chrome has issues, but they're issues in line with being made by an adtech company; we might be unhappy at Google breaking adblockers (https://www.eff.org/deeplinks/2021/12/chrome-users-beware-ma...), but it's not out of character. Mozilla can say "More power to you. Mozilla puts people before profit, creating products, technologies and programs that make the internet healthier for everyone." (https://www.mozilla.org/en-US/) or they can, say, make Google the default engine ($), bake in a proprietary service (Pocket), rip out features (RIP compact theme), overrule user autonomy (Want to install an extension? Better upload it to Mozilla to get signed so they permit you to run it on your own computer!), ship a marketing extension through the "experiments" feature (https://blog.mozilla.org/en/products/firefox/update-looking-...).... but not both. Either empower the user, or don't, but don't pretend to empower the user while ripping away their control.
You can't please all people all of the time, and I agree the pocket integration, and the looking glass add were mistakes, but the other items were directly related to sustainability of the project ($, eng cycles), or user safety.
You can disagree with them as much as you like, but Firefox continues to support the ultimate in user control by releasing their product as open source. Roll your own build that doesn't require those features, sideload your add-ons, and/or fork the product.
As a user, the average Firefox user has far more control over the browser than Chrome, Edge, or Safari users do, and have the flexibility to use one of many Firefox forks that have the same beef as you.
> You can disagree with them as much as you like, but Firefox continues to support the ultimate in user control by releasing their product as open source. Roll your own build that doesn't require those features, sideload your add-ons, and/or fork the product.
By that standard Chrome is a paragon of user control. Firefox, as it actually exists, in the thing that Mozilla offers users to download, claims to care about user empowerment while constantly reducing users' power.
Another third party request is Firefox's phishing/malware protection. It periodically downloads their own bad site collection and if you visit a site, it check if it's on the list. And if it isn't, it checks with Google if the site is okay.
https://support.mozilla.org/en-US/kb/how-does-phishing-and-m...
P.s. Tor Browser turns this Safe Browsing feature off. Looks like I'm on the same page as them on the implications of it.
But even without the privacy issue, you should turn off save browsing because google should not be in control of what users can and cannot donwload. They clearly do not care about keeping that list free of false positives, for example: http://dege.freeweb.hu/dgVoodoo2/
> It does NOT contain any malware. Use a browser that is free of Google Shit Browsing security service crap (which is based on tons of noname antivirus "engines", look at VirusTotal if interested).
I have also experienced Googles disregard for false positives on that list myself. While they may "remove" false listings after you bug them, those entries will just be re-added the next week and of course because this is Google there is no way to get an actual human to look into it. It is insane that all browsers allow a private company to maintain such a list without complete transparency and publicly visible reasoning for why each entry is in it as well as well defined procedures to contest false postives with agagain, publicly visible reasons for denial.
That's not how Firefox does safe browsing. https://feeding.cloud.geek.nz/posts/how-safe-browsing-works-...
"One of the most persistent misunderstandings about Safe Browsing is the idea that the browser needs to send all visited URLs to Google in order to verify whether or not they are safe."
It also does not adress the problem with making Google the gatekeeper deciding what you can and cannot download - and don't tell me its just a warning, its set up in such a way that regular users will often not even know that they can bypass it.
"One of the most persistent misunderstandings about Safe Browsing is the idea that the browser needs to send all visited URLs to Google in order to verify whether or not they are safe."
1. Firefox checks it own list locally, which it updates every 30 mins
2. If not found, it chops off query params from the url, hashes it, chops the hash to a smaller size
3. Creates several bogus hashes, then checks them all with Google's Safe Browsing service.
YouTube turned off flash player as the default in 2015, and VP9 was supported in Firefox at the same time.
YouTube still serves h.264, vp9, and av1.
I was trying to figure out what this could be referring to.
https://news.ycombinator.com/item?id=19815348 https://news.ycombinator.com/item?id=28495546
Does anybody here remember enough keywords to find it out ?
Edit: I guess it was this Twitter thread https://mobile.twitter.com/johnath/status/111687123179245568...
Edit 2: The associated HN thread https://news.ycombinator.com/item?id=19662852
The problem wasn't some web server, it's the Firefox backend services running on GCP.
It might (as mentioned in TFA) have made them think to run some extra tests, which could have caught the bug. But it also would have made the response faster, as they would have known what changed far sooner.
When you do Operations for a living, the only thing you can expect is the unexpected. That's why even after you think you've tested a change, you carefully and slowly roll it out a bit at a time, monitoring golden metrics so you can detect a problem, stop the roll-out, and roll back.
It sounds like somebody just flipped a giant switch and never checked error rates, connection metrics, anything. Check out this graph: https://hacks.mozilla.org/files/2022/01/crashes-foxstuck2-20... Think maybe that would indicate somebody needs to roll back the last change?
The problem here, as usual, is a disconnect between stakeholders. Google has this service (it seems like the load balancer for their customer?) it wants to change for one reason or another. The customers may or may not have planned for the change Google is making. Google makes the change, but it isn't a stakeholder of the customer (they basically don't care what happens to the customer). So there is no direct feedback loop for the customer to tell Google something is wrong.
If Google was at risk of losing business from its customers going down, it would have a strong relationship with those customers and have a way to quickly help diagnose problems and roll back changes if needed. This is a great lesson for all customers to take away: don't depend on people who you don't have a close relationship with.
I don't think that's as bad as you try to make it... if the client says it supports something then it breaks when it uses it, it's the fault of the client, not the server.
Google literally wrote the books on SRE. For them to not know better is absurd.
Lay with the dogs, wake up with the fleas.
Google is a shitty company producing shitty products. When you select to do business with Google you select to do business with a shitty company producing shitty products and treating its customers like shit. Hence I fail to understand the Surprised Pikachu face when something like this happens.
We all make mistakes, but don't try to make it sound more grandiose ("a combination of multiple factors blah balh blah") than it is.
> As in HTTP/2, characters in field names MUST be converted to lowercase prior to their encoding. A request or response containing uppercase characters in field names MUST be treated as malformed
https://quicwg.org/base-drafts/draft-ietf-quic-http.html#sec...
Only if you haven't updated your knowledge since HTTP/1.1 days...
"This is why users who disabled Telemetry would see this problem resolved even though the problem is not related to Telemetry functionality itself and could have been triggered otherwise."
When expected anomalies happen, like telemetry being down or taking a long time to respond. Firefox is certainly already designed like that.
This was not that. This was a bug. There is no magical design that avoids bugs.
> This is almost akin to me if Tesla rolls out an update and the car decides to pull over to the curb to do the update, while you are driving to your job or worse to hospital with an emergency.
Firefox is not a car. You're going to have to get Mozilla a lot more funding if you think the browser should be designed with extreme resilience in mind as required for life-critical applications. If a health enterprise is using Firefox in a life-critical role, that's kind of their responsibility, not Mozilla's.
It does not need this big of an explanation. It's just a silly bug, they can do better by improving their testing. The lady doth protest too much.
It is critically important that we not introduce new single points of failure into the systems that our civilization depends on, and that we remove the ones that already exist.
If it can happen by accident, it can happen on purpose.
There is a why here, and it includes Telemetry mixed traffic as a potential culprit. There are reasons to unify traffic (proxy support, QoS and whatnot) but unification of the user and Telemetry streams isn't without risk, as has been shown.
A constant refrain over the last 10 years or so of Mozilla's descent while trying to justify the removal of features from Firefox has been that not doing so unnecessarily bloats the surface area of the codebase, and specifically that this increases the chance of vulnerabilities and defects.
Will the same argument be applied here, now with a case in hand, to justify the removal of telemetry, too?
What do you mean that's not what happened? Was there a "Firefox outage" or not? Are you disputing claims that the engineering team made about Firefox becoming unusable for users "for close to two hours"?
That's my point. A "telemetry bug" didn't make Firefox unusable, a networking bug that was triggered by a telemetry bug did. But it could just as easily have been triggered by anything else.
Secondly, "cause" doesn't automatically mean "root cause". (That's the entire reason we distinguish between the two by qualifying the latter to begin with.) It's perfectly reasonable to say "A caused B" even if the root cause lies elsewhere, with C.
Thirdly, none of this matters. It has no impact on the point being made by the person you responded to, which—to repeat—is that:
> It should be impossible for the phrase "the recent Firefox outage" to make sense.
It makes perfect sense a world where half the internet is going through Google / Cloudflare / Amazon /Akamai servers or some combination of the above, and they decide to roll out brand-spanking-new protocols to half of the internet at once. Sometimes that's going to break clients.
I don't like that world very much, but it's the one we live in.
There was no Telemetry outage. An HTTP3 response header’s case was changed by a third party without notice. Telemetry continued working, other than the case change causing a bug.
There was an http3 infinite loop bug in Firefox that crashed all networking. Many different things could have triggered the bug once it was introduced. Telemetry happened to be the first thing to do so, but not due to any faults in Telemetry’s code or implementation.
I completely agree. We shouldn't encourage this term with Firefox too. It is my personal client, not a global service.
Looks like one factor, to my eyes: telemetry.
I have telemetry disabled. But if you're going to default to "telemetry on", and then silently send data to sites that aren't in the address-bar, then it's your responsibility not to "break the web". You can't blame it on rust, or necko, or viaduct, or google.
It's pretty clear from the article that this was a bug, and telemetry requests failing was not intended to break the rest of the browser.
If Google had rolled out this change a year from now, it could probably have broken something other than just telemetry (e.g. maybe update checks, or certificate management) and your browser would still have been broken even with telemetry disabled.
If telemetry didn't exist as core part of the browser, nothing would've broken.
Therefore, the telemetry itself is the direct cause of the bug. It was, at best, poorly handled and too deeply integrated into the browser's core function.
You can't find a lighter straw man than that. Find me somebody who said that Firefox intended to break the browser if telemetry failed.
Lesson learned: do not opt your users into services without their consent.
So yes, it wasn't limited to Telemetry, but no I had not seen the bug in practice until that very moment.
You know I moved from Netscape 3.0 Gold to later versions, to Mozilla, to Phoenix, Firebird, and then Firefox. I tried other browsers but it's always a subpar experience for me. My only gripe is that they kept changing the UI.
Use an older version? Security problems. Use a simpler browser? Many sites will stop working.
Maybe it's for the best to avoid using sites that use complicated JS/CSS/HTML. But will still need it for say, government sites to pay taxes.
There's a lot of FUD and paranoia out there; 99% of exploits need JS and even those which don't technically need it, are almost always obfuscated using JS.
Leave JS off by default (there are extensions to do that) and don't turn it on unless you really do trust the site to run arbitrary code on your computer, and you're unlikely to encounter any problems.
Looks like the crucial issue to me. The SPOF which enables something as absurd as "Firefox Outage".
Man, I love Firefox and used it since it was called Firebird (with a small gap when Chrome was shiny and new and Firefox a slow RAM hog). But I really resent the Mozilla Foundation, they seem to be interested in everything but browser development. To be fair, (ab)using the browser as application runtime brought us so much complexity that developing and maintaining a secure browser as free software spare time project isn't feasible anymore.
If you read the whole post, the connection was explicitly for telemetry (and so you could avoid the issue by turning off telemetry), and it blocked other connections because the request went into an infinite loop rather than failing outright.
Speaking of reading the whole post:
>> users who disabled Telemetry would see this problem resolved even though the problem is not related to Telemetry functionality itself and could have been triggered otherwise.
This does not make sense to me:
>> Without the header, the request was determined by the Necko code to be complete,
This is written as if it makes sense to treat a request as "complete" when it's missing a content length header. Huh?!
All they mean by that is that it was a bug in their HTTP/3.0 code that could have been triggered by any HTTP/3.0 connection that was using that codepath. But the reason for that particular connection (which had the conditions to trigger the bug) was for telemetry.
> updates, telemetry, certificate management, crash reporting and other similar functionality
We can discuss about the importance of telemetry, but the others seem quite important to me.
The key is just to gracefully fail when something goes wrong (and there it didn't).
Does it mean that the same blocking bug could happen while browse a local website on an air-gapped network? Or while opening local HTML files while offline?
So yes, you can use Firefox in any offline environment.
> the client was hanging inside a network request to one of the Firefox internal services.
Presuming that while offline, there is no pending network request to hang on, yes, it would have worked if you were offline.
This is the part that really gets me. For an average user, they trust the certificate that is bundled with the browser vendor (yes you can do certificate pinning). It just seems like something like certificate for encryption, ought to be split up away from browser vendor rather managed by a open public repo mange by a non-profit. Or have it on a block chain type of ledger. Any thoughts on that HN?
None of them are; disconnect from internet and start Firefox. It will work.
It was just a bug in the Firefox HTTP 3 implementation that caused it to be rendered unusable; it just so happens that connecting to these services triggered it, but it also could have been triggered by another HTTP 3 service (as I understand it, anyway).
If you're making an HTTP/3.0 request formed "correctly" then yes, it too would cause the infinite loop. It's not in any way specific to the internal service.
The outage could have been worse, though, because the bug that was introduced in the code for the Firefox update was in a feature of Firefox itself that allows users to block certain updates until a later time. In the case of Firefox, that feature was used to block the update that caused the outage. If a similar situation occurred in the future, users would not have a way to block the update that causes the outage. That feature is available in other web browsers, but it is not as advanced or robust.
As others pointed out, why does Firefox even need to communicate with Mozilla services? Sure, telemetry needs to feed data back, if enabled, but if that fails why does it need to stop the browser from working?
Shouldn't the lesson learned be: The telemetry functionality in Firefox has a bug, where an infrastructure outage at Mozilla can "break" the browser. The fix isn't in infrastructure, the fix has to be in the code that communicates back, it should fail gracefully. It's not a problem if telemetry fail, either cache locally and just drop the data, it's honestly not important.
I'm sorry, I get that it's interesting how and why all this failed, but Mozilla makes it seem like they don't get what the root of the problem is.
Don't make your telemetry backend more reliable. Instead break it on purpose several times a day. That way a similar bug in the browser will not make it past dev channel.
My thinking is that list should be reordered as ultimate culprit to blame is firefox itself.(make the third point on lessons learned as the first one)
Of course they'll fix so that this problem doesn't occur in the future. But as said, it had nothing to do with telemetry. Just that telemetry happened to trigger the bug.
(That's a lot more graffiti than "8.8.8.8".)
They lay out a bunch of things to fix about a system that could bring down your browser. What if that system that communication with Google about GCP messed up for them isn't even the IPs you are contacting during unrest in Khazakhstan or an election in Uganda? What if it is a semi-intentionally confused transparent proxy?
Testing a few more things that a friendly proxy may do as it improves your connection is hardly the same as assuming the worst about your network in proper paranoia mode.
They are fixing that thing and all those circumstances and try to make sure that circumstances of that characteristic won't happen again.
The what-ifs you are talking about could just as well be any homepage on the internet.
... and that page could also be MITMed.
So your point is to not have bugs?
I'm the first to agree that it absolutely should be OPT IN and not OPT OUT as it is now, but even so my biggest concerns would not be trying to evade hostile countries. If that is your bar you can't just install a mainstream OS and mainstream software and assume that it is a good idea.
The bug that caused the hang was in the network stack itself. There was no way the calling code could have prevented this in any way. You can see this by taking a look at the linked HTTP3 code. It's not that the higher-level code kept retrying over and over causing the hang, that was not the problem here.
Under "Lessons learned" you can also read "investigating action points both to make the browser more resilient towards such problems". I agree that this is broadly spoken, but it covers ideas that would have made this technically recoverable (e.g. can network requests be compartmentalized to not block on a single network thread?).
It could have prevented it by not making the call in the first place.
“This is why users who disabled Telemetry would see this problem resolved even though the problem is not related to Telemetry functionality itself and could have been triggered otherwise.”
Since a browser's job is to make HTTP requests, a bug in the network stack would almost certainly have been hit in other places. This was highly-visible so it was quickly noticed but it's quite possible that a less frequent trigger could have plagued Firefox users for a much longer period of time as HTTP/3 adoption increases.
The code _does_ work the way you describe, _except_ for the latent bug that caused the networking thread to get stuck in an infinite loop, which it was never supposed to do, even when errors occur.
It was never supposed to work that way, and the fact it did was because of a bug they'd never seen before.
So it wasn't that "oops, we shouldn't have built the system to get stuck forever when it fails" but rather "this bug triggered that bug which combined to cause a far worse result than 1 bug alone could have".
The only "lesson learnt" there is either that they need better ways to find bugs, quadruple up their thread count just so that different subsystems can't coexist on the same threads to avoid a theoretical problem that shouldn't ever happen again, or they just come up with infrastructural changes to minimise the negative results of the next "2 bugs reacted together and caught fire" scenario, which is the one they went with, and the only sane one.
Hint: It doesn't.
And then when the regular users are all piling up on support because they can't learn how to configure the product (or indeed ask their power user friends for help, since the features no longer exist), what happens then?
Nothing. Nothing happens then. And the shitshow continues.
Feature usage is a poor proxy for usefulness, importance, usability, visibility; it confounds them all.
No. The benefit you're describing is not a direct benefit. It's an exemplary instance of indirect benefit even under the most generous evaluation criteria/process.
I know of changes for the worse that Mozilla has made, using telemetry as a justification.
I don't know if I've heard of any changes for the better that have been prompted by telemetry data. It could be... maybe there are bug fixes or UI refinements. I just haven't heard of any.
* Noticed performance regressions not caught by our testing, and therefore been able to fix the regression.
* Noticed an unexpected number of users with hardware acceleration disabled, and therefore been able to find and fix the bug that was causing them to have acceleration switched off
* Figure out which device in a category is most commonly used by our users, so that I can dogfood my work on a representative device
Those are just a few examples off of the top of my head. It's not about removing features because telemetry says nobody uses them. People Mozilla use telemetry to answer all sorts of important questions. We also have to jump through hoops to add any new data collection, justifying why it's needed and ensuring the data is not personal. As is right, because we take user privacy very seriously
The problem here was that something that is known to fail for all sorts of reasons (network IO) was happening without such a timeout. Or with a timeout with a failure mode that it never happens (yikes). That's a design problem and even something with a very small chance of happening is extremely likely to actually happen at some point with a product that is this widely used.
This stuff is hard of course and I end up addressing issues related to his once in a while. The fix is usually to surround such code with defensive measures such as timeouts, retry mechanisms, telemetry, logging, etc.
The additional question/learning is why they never noticed this happening before. Because it probably did; they just never noticed because the very thing that would have told them was actually hanging. People killing an application for whatever reason is something that you'd want to know however.
There was no way for the calling code to do this. This was literally an infinite loop inside the network stack. Imagine the network stack itself going `while(1) {}` on you, without checking if the request was canceled.
Even if you detect that this happens, there is nothing you can do as the caller. You can't even properly stop the thread, as it is not cooperating. So recovering from this type of failure is hard.
Like what happened in a comment that I called out yesterday, you're silently inserting extra qualifiers that aren't in the original; the person you're responding to didn't say anything about calling code.
If the network stack can end up doing the equivalent of `while(1) { /.../ }`, then that's the bug, no matter what's in the ellided part. There's not "no way" to deal with this. (In the specific case of `while(1)`—which I recognize is a metaphor and not a case study, so onlookers should please spare us the sophomoric retort—it's as simple as changing to `while(i < MAX_TRIES)` with some failover checks.) In some industries, this sort of thing is mandatory.
There is no algorithm that will determine the "halting status" of an arbitrary (program, input) pair, but that does not prevent a team of programmers from working in a subset of the set of all programs in which every program halts. Restricting themselves to that subset might make the team less productive (i.e., raise the cost of implementing things), but it probably does not materially limit what the team can accomplish (i.e., what functionality the team can implement) provided they're not developing a "language processor" (a program that takes another program as input).
Any code that can end up blocking forever under normal circumstances already has a timeout and recovers from that.
This wasn't a normal circumstance, this was a logic bug.
> The problem here was that something that is known to fail for all sorts of reasons (network IO) was happening without such a timeout.
No. Read the article. It was an infinite loop. Equivalent to while(1);. Not a network timeout. Not a network error. An infinite loop. A logic problem.
I am appalled at how many people replying in the comments here cannot grasp this basic fact. This isn't about some dumb telemetry design where telemetry requests block everything else. This was a logic bug in the network stack that wedged the entire thing eating 100% CPU. There's no miracle fix for infinite loop bugs.
The suggested learning was that more testing should have been done, but as a solution, more testing is a cop out. A real solution is to develop code in a language that supports a robust type system, and then using that type system effectively in development.
Not an easy solution, so in the short term, we'll have more testing, and more bugs.
You know what happens when you put a while(1); in the middle of the nginx codebase? The whole server process hangs. This is normal in an async design. We don't write software to be magically resilient against freak bugs, especially not something like a browser that is not intended to be used in life-critical applications.