A large crash spike affecting Firefox users on Linux
fosstodon.org
fosstodon.org
I used to be part of a team developing a popular browser WYSIWYG editor. Every release of any of the supported browsers was a coin toss regarding introducing new bugs.
From this perspective developing for the still supported back then IE8 was easier, because there was no chance for it to ever change.
Some features really are worth it.
The problem is more that they prioritise the glossy fluff ones. Like time-limited "inspirational" colour schemes.
https://www.mozilla.org/en-US/firefox/114.0/releasenotes/
> Users on macOS, Linux, and Windows 7 can now use FIDO2 / WebAuthn authenticators over USB. Some advanced features, such as fully passwordless logins, require a PIN to be set on the authenticator.
Webauthen was supported for quite some while, _but only a subset of it_.
With CTAP you can insert the stick, enter its pincode, touch it and you're logged in. It replaces even the username.. This never worked.
What did work was using the token as a 2FA token only (FIDO1 method). But that doesn't replace passwords.
But I also only have been using a subset of the Webauthn standard, specifically the subset generally used when using it for the 2nd factor in 2FA.
But the standard provides other usage methods, too. E.g. like using it as main factor + a PIN. And this methods had not yet been supported in the past.
2FA (FIDO1) has worked for a while yes. But it still requires a username and password and the token is only used for 2FA. But this is not what webauthn is. It's only a small subset that existed under the FIDO name before webauthn was designed and was basically grandfathered in. But it's not really what webauthn is about.
In full passwordless mode you insert the token, enter its pincode and touch it to login. No username nor password needed. It's a bit like a bank card.
Not many sites support this method, for example Paypal only supports normal old FIDO1 2FA (and only one token which is ridiculous). But this support is also needed to finally enable full passkeys in the future. This support is also needed to finally enable full passkeys in the future (I believe 1 or 2 things for that are still coming in a near-future version).
https://github.com/torvalds/linux/blob/84df9525b0c27f3ebc2eb...
The stack isn't for bulk storage, I thought that was common knowledge.
Basically the way a kernel detects stack growth is via a guard page that causes a fault on memory access. If the allocation is bigger than the offset to the guard page, and the start of the allocation is accessed before the end, you get this. An attacker might even be able to generate a non-faulting memory access pointing at somebody else's buffer.
I think it's on the language implementation to understand what that offset might be and generate benign memory access on large stack allocations, allowing the kernel to fault and intervene. Consider it part of the ABI where you're running.
But IMO a compiler should generate extra benign reads on a large stack allocation. You mention avoiding VLAs in the kernel, but even user mode code has this problem.
Certainly a browser JITing random JS from the internet should be able to work around such a problem.
How can we get Mastodon links like this to open in the Mastodon app?
On android this is decided by the url (apps can request certain urls to be forwarded to them) but since there are many different servers in fediverse that becomes impractical
Thanks for the heads up. I wonder at what android version I'll decide it's no longer worth the breakage to upgrade. I was still happiest with the options of Cyanogenmod 4.4, the only new thing I can remember being happy about since then is "allow permission while using" (and even that is not working well)
Mastodon had a custom URI for interacting with it, but it wasn't an opt-in feature and the constant prompt to register the handler annoyed people. That's why the feature and the protocol were removed and they haven't been added back.
I think it's rather silly to first ignore the calls to opt into the protocol popup, then remove it completely because there were too many complaints about the protocol popup, and now refuse to add it back because of "bad UX".
not a beta tester thou
So it may be a large function with lots of temporary values visible in scope. If I understand correctly, local lambdas would also count and they have their own captured context. I'm sure it's possible to find a pathological-but-not-unreasonable way to reproduce it.
(I didn't dig into the specifics of the Google code because, as weird as it is to have 20000 stack values, we really should be able to handle it. This was, at the end of the day, a bug in our stack probing code.)
That's pretty wild, I need to look deeper into how to disable telemetry reporting in Mozilla. I'm pretty sure even Microsoft is sanitizing their crash reports to exclude as much information as possible that could identify the user.
Also, does anybody known why Firefox still depends on the deprecated libdbus-glib-1-2 [1] in Debian and based distros?
For example, try to uninstall the package. Then download the latest Firefox from their website [2]. Extract the archive, launch the executable inside it from a terminal. You will see an error message that it is unable to load the DBus library.
I really blame myself though. It only happens after I've had the browser open for weeks, with at least 3 windows open at once, and literally hundreds of open tabs with hundreds more come and gone since Firefox was restarted. I have about:memory bookmarked, and hitting everything in the Free Memory box seems to delay the inevitable a bit while allowing Javascript seems to make it worse. I doubt many others have the problem and I'm impressed Firefox holds up as well as it does!
Regarding libdbus-glib-1-2, you may want to open a bug. It looks like [1] it's mainly used by the ~12 year old UPowerClient, and more recently, for wake/sleep and timezone change notifications (nsAppShell).
[1] https://searchfox.org/mozilla-central/search?q=DBusG.*&path=...
I went from Yahoo to Google and really valued their product but they became so poor overtime. So focused on spiking their results with ads and then, much worse to me, hiding news and other content based on US centric biased narratives that apparently we all need to follow.
It was very apparent and annoying.
Don't get me wrong. I don't think assigning blame is the most important thing to do when troubleshooting. I'd rather not, but when that process starts it should be factual.
So, sorry, its not Google's fault... Then we also throw Linux under the bus. "it's not our code, it's Linus" here is the code. But, that Linux kernel code that kills a process if it accesses too far from it's stack pointer has the following comment:
"Accessing the stack below %sp is always a bug (...)"
I haven't got time to look at the history how it was changed and why in the Linux kernel and it became "not a bug". If someone knows more, please do explain.
So is it a "bug in Firefox" Or "bug in old Linux"? I can't say with absolute certainty without researching how exactly the stack allocation in old Linux kernel works, how is it documented etc.
So if anything I'd thank Google for exposing the bug ;-)
On a side note, I've recently experienced a similar JS/firefox/web site bug. There is this open-source ecommerce software called shopware. They use symfony (yes, PHP i know...) and the most recent major version simply freezes Firefox when one goes to the admin interface and looses connectivity. Not just freezing one tab, no, freezes entire Firefox, multiple open windows. This is on up to date Arch with a new Linux kernel so it's definitely not this issue,but it does happen in Firefox and not in Chrome.
JavaScript bugs like this are hard to find. I think AI may be one tool that will help us find them faster (intentionally or not).
As for testing. Really, you expect them to test against Linux Kernel 4? If so, how about 2 as well?
Just in case you didn't realise. Kernel 4.20 (the one that changes this behaviour so it wouldn't show the Firefox bug) was released in December 2018. That's 4.5 years ago.
Once the user sees this crash constantly, switches to Chrome and it doesn't crash, then they blame the crashing browser which is Firefox. Rightly so.
Really says a lot about these Firefox developers and users immediately blaming Google for their JS code when Firefox was supposed to protect or handle against cases like this without crashing.
I am still in the dark as to what the bug was here. Did Firefox stop doing probes for JIT code? Not do them at all, because most JS stack frames are small?
https://phabricator.services.mozilla.com/rMOZILLACENTRAL304d...
Stack probing is kind of a weird thing though. I'm kind of surprised C++ compiler isn't doing that properly. Are they using inline assembly for speed?
For the probing, according to the code:
// Can't push large frames blindly on windows, so we must touch frame memory
// incrementally, with no more than 4096 - 1 bytes between touches.
//
// This is used across all platforms for simplicity.
https://searchfox.org/mozilla-central/rev/c936f47f3a629ae49a...Edit: Or perhaps that is not the stack of the C++ program but rather than stack of the Javascript program.
So the code in this function is not performing the stack-probing, it generates code to perform it instead.
The one exception about sharing the same stack is that the first few iterations of a function run in an interpreter implemented in C++, and that interpreter has its own stack that we heap-allocate. This particular bug occurred during the transition from the C++ interpreter to JIT code.
Morally it's Google. Google is like Bitcoin. A huge natural ressource hog for a questionable benefit. Here the benefit is tracking users to make advertising billions. For that goal a billion of smart phones need several extra GB of memory an significantly larger batteries. What is the ecological footprint of that?
Typing on a seven year old smart phone with 2 GB (SailfishOS, so yes it is maintained. Maybe not perfectly, but better than many Androids half as old). It works quite well on reasonable pages even without add blocking. Of course super heavy pages won't work and for Google search I haven't even consented. Occasionally I get reminded of that when some site embeds it. Well, a good reminder for me not to use them.
It's very likely to have affected users on other websites, but Google is a common denominator for debugging.
My uneducated guess: they're using the Google Closure Compiler to make smaller JavaScript bundles. It saves some bandwidth and allows for better optimizations. It seems like a reasonable engineering decision to ensure product decisions don't affect the user experience too negatively, something a lot of us are familiar with..
But first building insanely heavy functionality and then optimizing it is not sustainable approach.
The background is of course that increasingly heavy function is moved into the client. Technically and especially economically this makes sense for Google. But the environmental footprint is outsourced to client devices. Greenwashing in a way. Avoidance of computing is a better solution than all types of optimizations. Not everything that is technically feasible is good for the planet and mankind.
That's bundling, not tree shaking. Tree shaking is an optional additional process during bundling where unreachable code is automatically removed.
This example had me remembering the discussion from earlier today about why modern code is so slow even though the machines are so fast. I doubt there were many Win32 programs that attempted to pass in 20,000 parameters to a function.
The C++ compiler may or may not implement stack probing correctly, but this bug is in the Javascript JIT compiler. The fix is to make the JIT compiler update the stack pointer register each time through the loop so that if a page fault happens the kernel won’t see a huge amount of stack get allocated all in one go. Instead it’ll see several page faults each asking for a smaller amount which it is happy to grant.
See https://phabricator.services.mozilla.com/rMOZILLACENTRAL304d...