Digging for performance gold: finding hidden performance wins
blog.chromium.org
blog.chromium.org
I'd nitpick a little bit and say it's possible that an optimization in one case causes a slowdown in another case- or worse, a bug. Benchmarks can also be inconsistent on the "same" case due to caches, etc, some of which may live outside of the code you actually control. Even the simplest program will vary a bit when you re-run it on the same system due to the state of that system (other processes, temperature, CPU cache, etc).
Some optimizations are clear wins, but many of them involve trade-offs and can have some mystery. Thorough testing/benchmarking helps a lot, but it can only get you so far.
Integration/end-to-end tests are terrible for performance benchmarks. They are designed to capture functional issues at the seams. They are usually heavy and not very diverse (because very fragile and/or expensive to maintain), usually focused on a few key critical functional paths, e.g. "can I put something in my shopping cart and pay".
That's pretty much the opposite of real-world". The "real-world" is the distribution of what the end-user experiences. With proper tracing, one can identify the key real-world hotspots (including the data associated to it) and then focus on optimizing that.
Actually, I often enjoy it a bit too much... it's frequently the case that I'll realise I've just spent an entire day reducing memory allocations that didn't really need reducing, rather than building features :(
Imagine it was a system for detectives to track complex murder* investigations. They only need to conduct approx 10 such investigations per year, but quality/correctness is crucial.
The full end to end tests are performed thousands and thousands of time per year, and if they are slow the entire development effort is slow. So you end up needing to make the whole system performant - even though end users aren’t experiencing the pain, devs/QA are.
(* this is not the actual domain, but similar rarity + criticality)
// TODO: Remove this hack if we drop Windows XP
assert(min_win < win7)
... simple. I recall finding a function deep in Google search that had been "optimized" in x86 assembly, but way back when the cache lines were 32 bytes. On Opteron and later the "optimized" code was slower than idiomatic C++. That's when I decided any kind of performance decision needs to be recorded, somehow. Either something like `assert cache_bytes==32` or just a FIXME($date) that forces someone to revisit the decision every year.Edit: Looking through that thread seems like some plugins caused the issue.
It’s amazing to me how inefficiently a lot of browser extensions are written - eg last I checked, metamask pulls in web3, which is a clown car of javascript that takes hundreds of milliseconds to parse. That code needs to be parsed every time you navigate to a new website. You might not notice a single extension like that, but with a few bad extensions it’s easy for your browser to slow to a crawl. The obvious response is to blame the browser for stuff like this, but it’s usually the extensions that are causing your problems.
screenshot: https://dl3.pushbulletusercontent.com/dRiiaqbW844ZN3QHGcdWNe...
(Disclaimer: I used to work on Chrome many years ago.)
[1] https://www.chromium.org/developers/how-tos/trace-event-prof... [2] https://ui.perfetto.dev/
>Depth vs. breadth.
Ah yes, which direction do you look at your program? do you look at which functions consume the most resources bottom up (probably some string or memory function in libc) or top down?
If you're the person writing the system libraries for enormous platforms, probably bottom-up, but if you're an application developer, top down. Sometimes though, especially with the performance issue described in the article you're in the middle -- those are tough to spot!
Then, yes, spend time bottom up, what's were you're more likely to find consistent gains, usually by finding ways to call those low level functions less often.
Is that actually true? Doesn't this just mean that once every 50 minutes the system has been janky for >= 60s? Anything less & you're below the Nyquist frequency & are unlikely to be actually sampling it, no? My knowledge of signal analysis is just what I recall from some intro university classes so there could be more involved here in this claim so happy to learn if I'm misremembering (+ it might be made more complicated because their also sampling across a population of users).
(Speaking as somehow who regularly has to shut down Chrome because it's making my entire machine janky).
You either need more RAM, or a browser that uses less RAM...
Clearing the cache, or even the entire Chrome profile will fix it if its the case.
What you say probably is an issue with CPU scheduling though.
I do think it causes issues with CPU scheduling but it could be any number of other issues. I don’t think kernel developers are looking at improving the overall perf of the system with a large number of chrome tabs.
[1]: https://source.chromium.org/chromium/chromium/src/+/master:u...
EDIT: Filed the suggestion upstream: https://bugs.chromium.org/p/chromium/issues/detail?id=120214...
That's the point of the article: `GetLinkedFonts` is the "obvious culprit", but it's the fallback to the fallback, it should not be getting called in the first place. It doesn't really matter that it's slow because it should almost never be called.
And then, I assume they fixed (or will fix) the cache so that it'd cache failures, so GetLinkedFonts would only be called once per failure instead of being called over and over again, only for failure (as successes would get cached after the first one).
Also, this article isn't talking about UI blockage. This is talking about the time delta between user input & the result hitting the eyeball, presumably even across any asynchronous threads/IPC.
> Also, this article isn't talking about UI blockage. This is talking about the time delta between user input & the result hitting the eyeball, presumably even across any asynchronous threads/IPC.
Aren't those the same things?
> Aren’t those the same things?
Depends how you define it. Typically I think of “UI blockage” interpreted as “main thread doing CPU work or blocked on something and not processing events”. That’s a subset of the problems described (and maybe not even a perfect subset since you may have some kinds of UI blockage that’s not directly tracked to a user action). A user action might cause a repaint of the cursor/text. That repaint actually gets to the user through the compositor which is an external process (for security reasons). That’s all asynchronous and means you have to actually plumb through all your time stamps and metadata about the source event in all dependent work in a meaningful enough way to come up with an answer.
You have to decide how to divide a continuous measurement into discrete instances of jank. Assuming that interpretation is right, they have decided to lump together up to 30 seconds of bad behavior into a single instance. Since the hit rate is still low, that seems like a pretty accurate way to get a count.
If they wanted to measure dropped frames, they could do that too, but it's a less useful number all by itself because you have no idea how they're distributed. Lumping every 30 seconds together gives you a much better idea of distribution.
Where is that setting? I'm pretty sure it asks on install, but what about after that?
Update: Seems to be in settings -> Sync and Google Services -> Help improve Chrome's features and performance
I ran very much into the problem that there are not really "unicode" fonts but rather the web browser is patching together characters from different fonts when you use Chinese, Arabic, Emoji(s), etc.
I want something that looks like the card at the art museum that introduces a piece so I have just a nice serif en font and a Japanese font I like because I have a lot of Japanese subject matter.
If I wanted to print some math character or arabic I would have to register that typeface in my system but it is a hassle at the moment.
What I get for this (as compared to HTML) is that the system understands the border of the card, which is big for a 6x4 card and I can align multiple printings on both sides to the limits of the hardware.
Jank is the least of my problems.