Speedometer 3.0: A shared browser benchmark for web application responsiveness
browserbench.org
browserbench.org
I'm looking forward to sucking at this, and then slowly and systematically improving. :^)
Improving Performance in Firefox and Across the Web with Speedometer 3
https://hacks.mozilla.org/2024/03/improving-performance-in-f...
> This is the first time the Speedometer benchmark, or any major browser benchmark, has been developed through a cross-industry collaboration supported by each major browser engine: Blink/V8, Gecko/SpiderMonkey, and WebKit/JavaScriptCore.
Very unscientific results using a Mac Studio - Chrome: 20.4, Safari: 17.9, Firefox: 20.1.
Safari on an iPhone 13 Pro Max - 16.5.
Safari 17.4. : 5.52
Chrome 122 : 6.25
Firefox 123 : 7.26
The results pretty much confirms my general feeling about how the browser behaves on my machine as well. Where Firefox being the fastest.
And 22.5 on my iPhone 14 on iOS 17.4
That is my smartphone is ~4x faster than my laptop.
22.7 with Content Blocker
28.7 without Content BlockeriPhone 15 Pro (safari) - 23.1
MacBook Air M1 (chrome) - 25.9
MacBook Air M1 (safari) - 18.0
REGULAR MODE
Safari 17.3.1: 27.1
Safari 17.3.1 (Private): 26.2
Firefox 123.0.1 (uBlock Disabled, Enhanced Tracking Protection Disabled): 29.5
Firefox 123.0.1 (w/ uBlock Enabled, ETP Enabled): 27.1
LOW POWER MODE
Safari 17.3.1: 17.37
Safari 17.3.1 (Private): 16.99
Firefox 123.0.1 (uBlock Disabled, Enhanced Tracking Protection Disabled): 20.0
Firefox 123.0.1 (w/ uBlock Enabled, ETP Enabled): 17.8
Safari has 0 extensions installed, and Firefox 0 extensions installed besides uBlock Origin. Benchmarks were run with each browser as the sole application open and plugged in to a power supply.
Speedometer 3.0: Arc: 22.6, Orion: 19.6, Safari: 19.0, Chrome: 22.6, Firefox: 20.7
Speedometer 2.1: Arc: 408, Orion: 467, Safari: 481, Chrome: 404, Firefox: 478
No changes beyond stock browser. No extensions beyond stock install. Battery.
But the title of future performance king is up for grabs! And now we have a de-facto standard for browser performance benchmarking.
It's not a physical speed, just a benchmark number. Think of it as arbitrary units, which allows you to compare different version of browsers on the same machine.
On the other hand, premature optimization is the root of all evil.
> Think of it as arbitrary units, which allows you to compare different version of browsers on the same machine.
That's precisely the problem. It's arbitrary, meaningless. Without any physical units, I don't know what's good or bad, fast or slow. And why do the scores go from 0 to 140 when the web browsers are all getting approximately 20?
The web ecosystem is extremely mature and widely used. The workloads are fairly well understood. It is a magic unit, but the factors that go into it have a lot of thought from real-world scenarios. Bringing up "premature optimization" is completely irrelevant because that's not what this is, it's about as far as you can get from that.
I don't know what it is. How exactly does the score relate to the experience of the web browser user?
I'm a browser extension developer, and I've occasionally had people ask me about Speedometer scores, but I have no idea what they're supposed to mean or what to tell these people.
The score is a rescaled version of inverse time - if it goes up, that implies the browser can handle more user operations per second, or alternately, it takes fewer milliseconds to complete a user operation in a complex web app.
We know that, but you haven't said anything specific about scores other than higher scores are faster, in an abstract sense, which has already been established.
If you run all the tests in half the time, your Speedometer score will double. If your score improves by 1%, it implies that you are 1% faster on the subtests.
(There are probably some subtleties here because we're using the geometric mean to avoid putting too much weight on any individual subtest, but the rough intuition should still hold.)
Benchmarking is hard. It is very easy to write a benchmark where improving your score does not improve real-world performance, and over time even a good benchmark will become less useful as the important improvements are all made. This V8 blog post about Octane is a good description of some of the issues: https://v8.dev/blog/retiring-octane
Speedometer 3, in my experience, is the least bad browser benchmark. It hits code that we know from independent evidence is important for real-world performance. We've been targeting our performance work at Speedometer 3 for the last year, and we've seen good results. My favourite example: a few years ago, we decided that initial pageload performance was our performance priority for the year, and we spent some time trying to optimize for that. Speedometer 3 is not primarily a pageload benchmark. Nevertheless, our pageload telemetry improved more from targeting Speedometer 3 than it did when we were deliberately targeting pageload. (See the pretty graphs here: https://hacks.mozilla.org/2023/10/down-and-to-the-right-fire...) This is the advantage of having a good benchmark; it speeds up the iterative cycle of identifying a potential issue, writing a patch, and evaluating the results.
21 is apparently better than 20, but how much better? You could say "1 better", tautologically, but how does that relate to the real world?
Driving a car 1 mile per hour faster may be better, in a sense, but even if you drove 24 hours straight, it would only gain you 24 total miles, which is almost negligible on such a long trip. Nobody would be impressed by that difference.
A 5% raise for someone who makes $20k per year is $1k, whereas a 5% raise for someone who makes $200k is $10k, which would be a 50% raise for the former.
> "The score is a rescaled version of inverse time" is the key here.
> If you run all the tests in half the time, your Speedometer score will double. If your score improves by 1%, it implies that you are 1% faster on the subtests.
> (There are probably some subtleties here because we're using the geometric mean to avoid putting too much weight on any individual subtest, but the rough intuition should still hold.)
To directly answer your original question: a reading of 21 is 5% better than a reading of 20 because 21 is 5% greater than 20, and this means that a 21 speed browser should do things 5% faster than a 20 speed browser.
TL;DR: They behave like you'd expect.
To what?
I talked about driving a car. Miles and hours are an absolute reference. We still have no absolute reference for Speedometer.
If you scratch out the labels of your car speedometer and forget which is which, it still measures speed. 80 is still 33% faster than 60, regardless of the units.
I suspect your questions would be answered better by playing around with the tool in question for a few minutes anyway, as you seem to be asking about capabilities the tool does not purport to have.
Exactly.
The speedometer graphic was inherited from Speedometer 2. When Speedometer 2 was released, scores were in a reasonable car-speed range. The combination of hardware and software improvements meant that early versions of Speedometer 3 (which includes a subset of Speedometer 2 tests) were consistently scoring above 140, so we adjusted the scaling factor (IIRC, by ~20x) to give plenty of room for future improvements.
Still incredible that a gameboy or an 80's computer with a CRT feels more responsive than most devices these days.
Bring back tactility. I'm convinced the choppiness and weird waits are actually psychologically stressing us out. That's why good keyboards + old low latency OS'es or typewriters are so soothing to use.
I don’t see noticeable lag on pressing/tapping buttons or other ui components in day to day browsing, even on my quite old iPhone.
There are obviously ways to make delays in web content anyway (user action->synchronous network request being the canonical one), but assuming there’s nothing silly like that lag isn’t an issue I’ve noticed.
Actual execution latency is something I worked on for many years in JSC, and so there are a lot of engine optimizations to reduce that latency as much as possible (the interpreter itself, the interpreter performance, byte code caches, hilarious amounts of lazy parsing and source skipping, etc) so even the first time you have a ui element trigger code there shouldn’t be any significant delay.
Obviously if a developer makes poor choices there’s only so much you can do, but by and large there aren’t that many bad things a web developer can do that a native dev can’t also do (and devs in both environments frequently do :-/).
Meaning, the 300ms delay (if the site developer does no optimization) with mobile browsers?
Samsung S24 Galaxy - 13.8 (very new, with the good processor)
iPhone 13 Mini - 22.3
iPhone 15 Pro - 23.1
MBA M2 - 24.2 (Safari)
Win 11 i9-13950HX - 26.8 (Edge)
I figured the iPhones would be faster but not 2x.
I think that matters less for benchmark wars between tribes, fun as those always are, than as a stark reminder that any of us building for the public should remember than an S24 is _really_ fast for an Android phone and your median user probably bought whatever was on sale a couple years ago. That means that if you’re one of the many developers using an iPhone which isn’t roughly a decade old, you have no idea how your app feels to the median user because your device can run so much more code before it feels sluggish.
Curious: Is it the Exynos version (Rest of the world) or the Snapdragon version (US, China)?
As with any other benchmark its results will be interpreted incorrectly and will have little effect on real world.
Google already has vast amounts of real-world data. The end result? "Oh, you should aim for a Largest Contentful Paint of 2.5 seconds or lower" (emphasis mine): https://blog.chromium.org/2020/05/the-science-behind-web-vit... Why? Because in real world the vast majority of sites is worse.
Browsers are already optimised beyond any reasonable expectation. Benchmarks like these focus on all the wrong things with little to no benefit to the actual performance of real-life web.
Make all benchmarks you want, but then Google's own Youtube will load 2.5 MB of CSS and 12 MB of Javascript to display a grid of images, and Google's own Lighthouse will scream at you for the hundreds of errors and warnings Youtube embed triggers.
Edit:
Optimise all you want, and run any benchmarks you want for the "real world", but performance inequality gap will still be there: https://infrequently.org/2024/01/performance-inequality-gap-...
Optimise all you want, and run any benchmarks you want for the "real world", but Lighthouse will warn you when you have over 800 DOM nodes, and will show an error for more than 1400 DOM nodes (which are laughably small numbers) for a reason: https://developer.chrome.com/docs/lighthouse/performance/dom...
Why wouldn’t this?
Bad developers (or management dictates) will be bad no matter what. That’s not a reason to give up.
1. These companies themselves don't practice what they preach. No matter how fast Speedometer 3 is, Google's own web.dev takes three seconds to load a list of articles, and breaks client-side navigation. Google's own Lighthouse screams at you for embedding Youtube and suggests third-party alternatives [1]
2. The DOM is a horrendously bad, no-good, insanely slow system for anything dynamic. And apps are dynamic.
There's only so much you can optimise in it, or hack around it, until you run into its limitations. The mere fact that a ToDo app with a measely 6000 nodes is called a complex app in these tests is telling.
And the authors of these tests don't even understand the problem. Here's Edge team: "the complexity of the DOM and CSS rules is an important driver of end-user perceived latency. Inefficient patterns, often encouraged by popular frameworks, have exacerbated the problem, creating new performance cliffs within modern web applications".
The popular frameworks go to extreme lengths to not touch the DOM more than it is necessary. The reason the DOM and CSS end up being complicated is precisely because apps are complex, and the DOM is ill-equipped to deal with that.
This only goes to further show that browser developers have very little understanding of actual web development. And this is on top of the existing problem that web developers have very little understanding of how fast modern machines are and how inefficient web tech is.
This brings us neatly to point number 3:
3. Much of the complexity on the web in the modern web apps is due to the fact that the web has next to no building blocks suitable for anything complex.
https://open-ui.org was started 3(4?) years ago by devs from Microsoft and you can see just from the sheer number of elements and controls just how lacking the web is.
So what do you do when you need a proper stylable control for you app? Oh, you "use inefficient patterns by modern frameworks" because there's literally no other way.
And even if all of those controls do end up being implemented in browsers, it will still not be enough because all the other things will still be unavailable: from DOM efficiency to ability to do proper animations to ability to override control rendering to...
[1] I'm not kidding. Here's the help page it links: https://developer.chrome.com/docs/lighthouse/performance/thi...
Of course website authors should also do their part :-)
As an example "NewsSite-Next" has 8 out of 9 repetitions between 215 and 236 ms, but 1 out of 9 (iteration 5) is 1884 ms. This is such a radical outlier I have trouble believing it could be a browser bug. Visually, the interface seems to get "hung" when switching between tests sometimes. I don't have a great explanation for that.
This specific issue in this case in with NewsSite-Next/NavigateToUS which reports a 144.5% variance as a result of this one outlier.
I see several others like this in the results, although none quite as extreme.
So while Firefox is now super fast the performance of webapps might still be very bad on mobile.
And I think this applies to all modern browsers: they are fast at rendering very slow webapps and websites.
E: I was wrong/extremely out-of-date - it does have JIT but relies on the Safari/Webkit implementation. In ancient versions of iOS, the WebView widget that third-party browsers were forced to use had JIT disabled, but that’s long since changed.
In this case, I think the 3 score must be either very old/low-end Android hardware or a measurement error. I don’t think any iOS browser gets 3.x scores, on even remotely modern hardware.
What do these values represent? I can guess the last. Unsure of the first 2.
96.84 ±(5.0%) 4.87 msSounds like old hardware or some other issue
I have never felt performance has ever lacked, outside of a few outlier sites (youtube, facebook, twitch). But those are tightly coupled with their (crappy) implementations.
The scale goes up to 140 so that there's some space for software improvements as well as future hardware.
Chromium (122.0.6261.112) without any extensions:16.7 ± 0.34
Brave Beta (122.0.6261.111): 12.8 ± 0.62
Brave Beta (122.0.6261.111) Private: 15.9 ± 0.58
Floorp (11.10.5 based on Firefox ESR 115): 7.40 ± 0.20
Floorp (11.10.5 based on Firefox ESR 115) Private: 7.91 ± 0.19
Librewolf (123.0-1): 8.41 ± 0.19
Librewolf (123.0-1) with uBlock Origin disabled: 8.86 ± 0.17
I haven't run into anything like this; you may be an outlier. I interact with ~7 Firefox instances (Win) each week. Each has diff configs and plugins.
Mac mini M2, macOS 14.4, Chrome 122: 30.2 ± 1.6
Good ol' Apple Silicon.
I vaguely remember there's a privacy protection that rounds timer information. eg. all timers get rounded to the nearest 100ms. If you have a bunch of tests that take less than 100ms to complete, those tests might seemingly complete at the same time they start, which causes them to have infinite score.
There is only one warning in the Console: Ignoring ‘preventDefault()’ call on event of type ‘wheel’ from a listener registered as ‘passive’. react-dom.production.min.js:29:112
Firefox 123.0: 12.1 ± 0.62
Edge 122: 14.8 ± 0.68
Google Pixel 7a, GrapheneOS, on battery:
Fennec 123 with Dark Reader: 2.91 ± 0.066
Fennec 123 without Dark Reader: 5.28 ± 0.089
Vanadium 122: 6.96 ± 0.39
This kind of thing drives me nuts. And it’s just text, it’s not like it’s rocket science.
Pixel 8 Pro; Android 14 QPR2; Chrome: 9.47 in an incognito tab
Vivaldi:12.2
Brave:18.1
Safari:18.2
Chrome:19
Firefox Focus:21
MacBook Pro, M2 Pro, 16GB, plugged in, external display: Safari=31.2 Chrome=29.4
iPhone 12 mini, plugged in: Safari=19.4
HP Z2 mini (i7): Edge=15.9
Panasonic Toughbook CF19 (win 10): Edge=4.7 Chrome=5.6
Galaxy Tab S5e: Chrome=2.2
Oculus Quest 2: browser crashed
Tizen TV: displayed, wouldn't run
Nintendo 2DS: displayed, no css, wouldn't run
You brave fool. I love that you tried it.
Machine: AMD 5950X, 32GB RAM, 3080 GPU, Windows 11 Pro 23H2
Firefox v123.0.1 Edge v122.0.2365.80
EDIT: Interesting, I tried both in private windows later to bypass extensions, and Edge got 10.8 this time, Firefox got 16.9. I now have more questions.
Ran it in each browser one at a time while the music played in Chrome.
Chrome 122.0.6261.112: 21.3 +/- 0.64
Edge 122.0.2365.80: 20.1 +/- 0.78
Firefox 121.0.1: 18.5 +/- 0.75
Machine specs: Intel Core i9 12900k (24 core) / 64GB RAM / 3080Ti / Windows 11 Pro 23H2
After finishing the tests, I played that same song on Firefox and Edge. Both Firefox and Edge played it perfectly.
> audio players on Edge stutter a lot while Firefox plays them smoothly
I'm curious about what could be leading to this inconsistency as I use Web Audio for a number of projects, so I have a bit of a vested interest. It is notoriously easy to do WebAudio wrong or to do just a bit too much computation which leads to buffer underruns (pops). It also may have a lot to do with specific tracks on DeepSID, could you share some tracks that perform inconsistently for you?
I think Edge's problems come from some kind of power efficiency setting, not necessarily performance-related. (Like a low-granularity JS timer, or something like that)
EDIT: Turning off all efficiency settings on Edge didn't make any difference: 11.0
Definitely would be a bummer if web audio didn't work reliably on the default power plan (which is Balanced iirc)
- Firefox on WSL2 on the same machine gets 10.1 despite that the rendering must be horribly slow since it goes through Remote Desktop layer.
- Firefox gets a 25% speed boost when I disable 1Password extension. Disabling it on Edge makes no difference.
I don't think that AV stuff would be tested by Speedometer.
It probably isn't, but fwiw yes web audio is controlled by JavaScript. Doing it right means using web audio worklets, which is a special purpose JS context that has no access to your main page context.
Desktop Firefox: 25
Desktop Chrome: 26
Laptop Firefox: 16
Laptop Chrome: 20
Laptop Safari: 21
Phone Firefox: 12
Phone Chrome: 10
---
Desktop: 5900X, 3090, Linux
Laptop: M1 Pro 14"
Phone: S24 Ultra
Ran all tests in private window to avoid extensions, and gave a minute to cool between tests. Laptop/phone was plugged in.
Safari: 24.1
Firefox: 21.7
Extensions can really slow things down:
Safari w/ Ghostery: 6.83
Safari w/ AdBlock: 13.0
firefox 26.3
desktop: 7800xd3, 2060, linux
Highest scores across tests: Firefox: 6.34 +- 0.31 Brave: 11.3 +- 0.37 on Ryzen 9 7940HS + RTX 3060 mobile
Which really sucks since I highly prefer Firefox but this past week I've been trying out Brave and I think its noticeably faster and smoother to me. Even with the reduced speed I'm still swayed toward Firefox for the customization factor you can achieve with userchrome.css file.
The browsers have the same or very similar APIs for the extensions but that is just the interface; each browser executes the extension's instructions differently (a lot or a little - I don't know the browsers' code). The same extension will impact Brave's performance differently than it will impact Firefox's. In other words, the same extension is not, in this sense, the 'same' on each browser.
In this sense, an extension is part of the user experience, like a website. The Speedometer test suite doesn't include those extensions (I assume) and that is the experience the browsers are optimized for.
The parent's test doesn't represent that; it does represent their desired experience, of course.
Interesting because I have only 5 extensions. The heaviest extension seems to be Dark Reader which causes over 5 point changes.
Not sure if it's a coding issue in the Firefox version of Dark Reader, or it's hitting some slow path in Firefox itself.