The web at maximum FPS: How WebRender gets rid of jank
hacks.mozilla.org
hacks.mozilla.org
The reasoning was somewhat different, web pages were essentially static (we didn't do "DHTML"), if the page rendering process could generate an efficient display list, then the page source could be discarded, and only the display list needed to be held in memory, this rendering could then be pipelined with reading the page over the network, so the entire page was never in memory.
Full Disclosure: while I later wrote significant components of this browser (EcmaScript, WmlScript, SSL, WTLS, JPEG, PNG), the work I'm describing was entirely done by other people!
[1] - I joined in 97, the first public demo was at GSM World Congress Feb 98
The FORTH code then downloaded PostScript code into the NeWS server, where it would be executed in the server to draw the page.
It even had an Emacs interface written in Mocklisp!
http://www.donhopkins.com/home/archive/HyperTIES/ties.doc.tx...
http://www.donhopkins.com/home/images/HyperTIESDiagram.jpg http://www.donhopkins.com/home/images/HyperTIESAuthoring.jpg
http://www.donhopkins.com/drupal/node/101 http://www.donhopkins.com/drupal/node/102
http://www.donhopkins.com/home/ties/ http://www.donhopkins.com/home/ties/fmt.f http://www.donhopkins.com/home/ties/fmt.c http://www.donhopkins.com/home/ties/fmt.cps http://www.donhopkins.com/home/ties/fmt.ps http://www.donhopkins.com/home/ties/ties-2.ml
My understanding is that the demo at GSM World in 98 was the first time anyone had demonstrated a graphical browser on a cellphone that implemented common web standards (HTML/HTTP/TCP) rather than using a transcoding gateway (up.link or WAP style — the approach taken by Unwired Planet)
The smallest device that we targeted (IIRC) had a screen that was 100x64 pixels, we had 64KiB of RAM at runtime (`unsigned char gHeap[0x10000];`), and our ROM footprint was about 350KiB
https://m.gsmarena.com/benefon_q-42.php
It even supported black&white animated GIFs in HTML pages.
Have there been any measurements on what the end result is on a typical modern laptop?
GPUs are much more power efficient per FLOP. E.g. in my desktop PC, theoretical limit for the CPU is 32 FLOP/cycle * 4 cores * 3.2 GHz = 400 GFLOPS, for the GPU the theoretical limit is 2.3 TFLOPS. TDP for them is 84W CPU, 120W GPU.
A GPU has vast majority of transistors actually doing math, while in a CPU core, large percentage of these transistors are doing something else. Cache synchronization/invalidation, instructions reordering, branch prediction, indirect branch prediction (GPU has none of that), instruction fetch and decode (for GPU that’s shared between a group of cores who execute same instructions in lockstep).
When discussing existing browsers:
> But often the things on these layers didn’t change from frame to frame. For example, think of a traditional animation. The background doesn’t change, even if the characters in the foreground do. It’s a lot more efficient to keep that background layer around and just reuse it.
> So that’s what browsers did. They retained the layers. Then the browser could just repaint layers that had changed.
Then, later in the article, when discussing WebRender:
> What if we removed this boundary between painting and compositing and just went back to painting every pixel on every frame?
(Emphasis mine.)
So is WebRender less efficient in terms of power usage? Or is there some other factor that offsets the cost of this extra work?
A particularly relevant bit:
> For the case where the CPU would have painted a single pixel, WR will certainly use a bit more power than CPU rendering with a compositor. For the case where a large portion of the screen changes, CPU renderers might miss their frame budget spending > 16ms drawing a frame, where WR will complete that task in 4ms.
> We might light up part of a GPU for longer than compositing would have, but we just saved >12ms of CPU compute. The power hypothesis is that the saved CPU compute consumes more power budget than the extra GPU compute we added. If there was no further JS code to run (app is idle) then we can go back to idle state after 4ms, instead of after 16+ms.
Edit: To clarify, I realize that the screen gets refreshed ~60 times per second (depending on refresh rate), but I don't think any rendering actually happens if it doesn't need to.
Why not?
Painting a page is mostly just copying pixels from textures too.
Benchmark it yourself if you don't believe me!
This article has some diagrams and explanations: http://www.hardwaresecrets.com/introducing-the-panel-self-re...
That said, GPUs are pretty clever about this. A big chunk of power consumption comes from IO. i.e., moving data from the GPU to the off-chip DRAM. Mobile GPUs optimize for static scenes where nothing is changing by keeping hashes of small blocks (say 32x32 pixels) of the frambuffer in on-chip memory. If the hash of a block hasn't changed, they don't bother re-writing that block into the framebuffer. But the GPU is still ends up running the shaders every frame.
Mention Quantum Flow and "performance" in it. It worked for me in the past that performance issues I had were solved within a few weeks or days.
Make sure to attach a performance profile with https://perf-html.io/ using the linked add-on there.
I have the same issues on Linux. It's really fast, but often it hogs one core completely and the CPU does a lot of work, sending the fans in overdrive. Especially when videos are involved.
Let's hope things get ironed until a stable release in november
Mention Quantum Flow and "performance" in it. Make sure to attach a performance profile with https://perf-html.io/ using the linked add-on there.
I had many cases where the mozilla devs were glad about such bug reporst (especially if you are on Linux - I am too :) - which is not so often used and reported like Windows or Mac)
Edit: Not sure why I'm being downvoted for stating a simple fact. The top-of-the-line AMD and NVidia cards have TDPs of 200-250w.
The important information to answer the question is how much power the GPU will use, not how much it can use.
I would expect that in the long term, GPU rendering would be more efficient. In the short term the fact that the CPU is having to do a lot of work to manage the GPU may make it less efficient.
These UIs and Game engines typically burn CPU and GPU even for an idle menu on the screen.
I suspect that Mozilla will only re-render and composite on actual changes. Otherwise the power draw will be quite noticeable. The benefit of this new arch will be that they can guarantee 60 fps during that time.
Just to give more perspective, invalidation is not free, and there are cases where invalidation itself takes longer than your frame budget. Also, in the case where you have to repaint everything or almost everything anyway, invalidation just made your problem worse since you spent cycles figuring out you can't skip any work.
I presume, though, that things are buggier then and the potentially introduced performance drops might actually make it feel slower for now. I don't know, though, I haven't tested it with just gfx.webrender.enabled.
You can find the current full list of steps to enable WebRender here: https://mozillagfx.wordpress.com/2017/09/25/webrender-newsle...
For example, I use a four finger swipe down to hit CMD-W, which closes the window in focus for most apps.
Speaking of rendering text glyphs on the GPU, there's a really clever trick(commonly called loop-blinn, from the two authors): https://developer.nvidia.com/gpugems/GPUGems3/gpugems3_ch25....
You can pretty much just use the existing bezier control points from TTF as-is which is really nice.
Additionally, the original Loop-Blinn technique uses a constrained Delaunay triangulation to produce the mesh, which is too expensive (O(n^3) IIRC) to compute in real time. You need a faster technique, which is really tricky because it has to preserve curves (splitting when convex hulls intersect) and deal with self-intersection. Most of the work in Pathfinder 2 has gone into optimizing this step. In practice people usually use the stencil buffer to compute the fill rule, which hurts performance as it effectively computes the winding number from scratch for each pixel.
The good news is that it's quite possible to render glyphs quickly and with excellent antialiasing on the GPU using other techniques. There's lots of miscellaneous engineering work to do, but I'm pretty confident in Pathfinder's approach these days.
Reminds me of Mark Kilgard's old NV_path_rendering extension, which (as so often happened for interesting NV stuff) never made the jump to portability. One of its touted benefits as an in-driver implementation was the ability to share things like glyph caches between multiple apps with separate GL contexts, but with "apps" increasingly becoming "browser tabs", maybe a browser-managed cross-tab cache is almost as good.
BTW, is the WebGL demo available online anywhere, for people who don't want to install npm?
Not yet. There are enough known issues that I'd like to fix first.
Do you think, once the implementation is more complete, that it would make it possible to render fullscreen scenes of vector graphics (say in SVG) with hundreds of moving and morphing shapes on a mid-range phone/tablet/pc?
I know it's very difficult to predict these things, but I thought you may have already seen enough performance characteristics to make an educated guess :)
I assume you've seen other things like Slug, glyphy, etc., which use a combination of CPU pre-processing and GPU processing to make the bezier intersection as efficient as possible...
Slug is asymptotically slower than Pathfinder at fragment shading; every new path on a scanline increases the work that has to be done for every pixel. (Of course, that's not to say Slug is always slower in practice; constant factors matter a lot!) GLyphy is an SDF approach based on arc segment approximation that does not do Bézier intersection at all.
For example, when using SVG shape-rendering: optimizeSpeed ? I truly hope that SVG is going to be part of this new magic, and that the shape-rendering presentation attribute is utilized. I don't think current SVG implementations get much of a speed boost by optimizeSpeed.
Speaking of which, to what extent will SVG benefit from this massive rewrite?
BTW, I see you implemented PathFinder in TypeScript as well. What do you think of it, especially since I assume a lot of your recent programming has been in Rust?
That said, Rust on the server and TypeScript on the client go really well together. Strong typing everywhere!
Just throwing it out there as a neat thing, the GPU space has all sorts of fun stuff like that.
I'll have to take a look at Pathfinder, have any links handy that I can dig into?
The name "WebRender" is unfortunate though. Things with a "Web" prefix - "Web Animations", "WebAssembly", "WebVR" - are typically cross-browser standards. This is just a new approach Firefox is using for rendering. It doesn't appear to be part of any standard.
So, it might actually turn into somewhat of a pseudo-standard.
This is in stark contrast to the style engine in Servo which relies on memory representation to be fast. So integrating it requires very tight coupling of data structures.
I'd largely forgotten what pixel shaders actually were, so it was nice to get a high level understanding through this article, especially with the drawings!
UPDATE: Ah, I see it's mentioned in the future work: https://github.com/servo/webrender/wiki#future-work
Vulkan?
This could possibly make some of the serial
steps above able to be parallelized further.
So it will be using OpenGL then?It's actually OpenGL which fits less into the architecture, but it's still easier to just bundle WebRender's pipelines all together and then throw that into OpenGL.
Don't PC games use thousands of draw calls per frame?
That said, even Intel GPUs can often deal with large numbers of draw calls just fine. It's mobile where they become a real issue.
Aggressive batching is still important to take maximum advantage of parallelism. If you're switching shaders for every rect you draw, then you frequently lose to the CPU.
A web page may not be able to reuse as much image data, I know. But smart game engines frequently look for ways to better batch sprites.
And technically speaking, if you're using a Z-buffer, you don't need to sort opaque layers at all. You can draw the layers in back and then draw more layers in front. Yes, you get overdraw in that case, but if you're using Z values for layering, you could potentially get better batches by drawing in arbitrary order (i.e., relying on Z-depth to enforce opaque object order), and in my experience larger batches gives you a bigger advantage than reducing overdraw.
Win 10, latest nightly, WebRender enabled
I had to get to hundreds of thousands (maybe 250k?) and it started bouncing between 40 and 60.
Very neat.
For video games you often see minimum/recommended hardware specifications to run a game.
This feels a bit like cheating. Not all devices have a GPU. Would Firefox be slow on those devices?
Also, pages can become arbitrarily complicated. This means that an approach where compositing is used can still be faster in certain circumstances.
That is certainly true, but a) the cases where you can do everything as a compositor optimization are very few (transform and opacity mostly) so aside from a few fast paths you'd miss your frame budget all the time there too, and b) we have a lot of examples of web pages that are slow on CPU renderers and very fast on WebRender and very few examples of the opposite aside from constructed edge case benchmarks. Those we have found had solutions and I suspect the other cases will too.
As resolution and framerate scale, CPUs cannot keep up. GPUs are the only practical path forward.
I'm not even sure. The frame rate is important for smoothness, but the regularity is also important : reading a video with at a consistent 30fps rate is more pleasant than running a 60fps animation and dropping a frame every 2 seconds.
I think there's a good argument for preventing security-critical apps from raw GPU access, because graphics cards and drivers are a huge amount of attack surface.
I still think WebRender is the way forward, but I hope they get it working with something like llvmpipe.
If anything it'll push the GPU makers to have better drivers support.
EDIT: As for the no-GPU case, it is an edge-case, in which we could for example switch to the classical renderer. If you're running FF from a VM as your daily driver, there's something not right somewhere I think. (I'm thinking of C&C servers for satellites still running on WinXP and stuff)
I am. QubesOS is a similar setup that runs applications in Xen VMs for security. I think you can do GPU passthrough but it's not recommended. At this point I don't think I would ever go back to running a browser "bare", even if there were no bugs, being able to have complete control over it (e.g. pausing, blocking network access, etc.) is a godsend.
I think the idea that "it'll push GPU makers to be secure" is weak. We're so far from that point it's not even on the horizon.
In general we've not smoothed out this stuff so that you can fall back cleanly when a GPU doesn't exist.
PS: And my second point?
Maybe 15 years ago.
I think 200 - 1000 GFLOPs is common modern era CPU performance. A single core can do 32 floating point operations per clock cycle. Half that, if you don't count multiply accumulate as two ops. Drop another half if you want 64-bit floats.
Also, what (time and energy) savings could compositing bring if implemented on the GPU? (versus repainting everything on the GPU).
Another savings (memory, time, unsure about energy) comes from not having to upload giant pixel buffers to the GPU.
In practice LRU caches work pretty well.
I think the best approach would be some hybrid between invalidation and redrawing everything.
So does that mean that it is known not to work with that GPU? Can you override the blocklist to see what happens?
Edit: It also says:
> Direct2D: Blocked for your graphics driver version mismatch between registry and DLL.
and
> CP+[GFX1-]: Mismatched driver versions between the registry 8.15.10.2697 and DLL(s) 8.14.10.2697, reported.
Indeed that is correct, the driver is marked as version 8.15.10.2697 but the fileversion of the dlls are 8.14.10.2697, this seems to be intentional by Microsoft or Intel, note that the build numbers are still the same. Firefox is quite naive if it thinks it can just try to match those.
Firefox, especially the new Quantum version is awesome. But Rust as a side product might be the best thing Mozilla brought us. I'm truly thankful for that.
I'm coming from a a Python background and I'm too spoiled by Pycharm's insane tooling and autocomplete and stuff.. I wanna trail de Rust path too!
thanks in advance.
For writing code, I've been using IntelliJ with the Rust plugin for the past year.
Autocomplete works pretty well, there is Cargo integration so you can edit your dependency file in IntelliJ and things will automatically update so that autocomplete and documentation (go to definition) for the newly added dependencies start working immediately.
There is also simple build/check/run support, with build and test output appearing in the lower panel inside the IDE. There are some quirks, but they have been very minimal.
I've never used IntelliJ to debug Rust itself, because in most of the projects I've worked on, the Rust code has been designed as a completely separate library with a semi-public API, intended for embedding in another language like Swift or C#.
So the debugger has almost always been Visual Studio or Xcode, but it works exactly as I would expect even without using any sort of Rust plugin for either one.
The other day I was stepping through the call stack to find the source of a crash (one that I caused by failing to keep the P/Invoke signatures up to date with our Rust functions), and Visual Studio switched right from C# to Rust from one frame to the next, showing the source code and the line where the crash occurred in Rust.
Really I can't think of anything that has been even a moderately significant problem, it's been great all around.
If you're interested in trying out a Rust IDE, those are the two I'd try and see how you find them.
Personally I tend to use Vim out of habit, and debug in Visual Studio. I might switch at some point for RLS (or try to hook it up to Vim).
I use rr for debugging (with gdb's command-line UI; I wish a GUI like Eclipse supported rr).
I can't emphasize enough that trying rr is worthwhile if you are developing in a language that gdb supports (not just Rust but C and C++, too).
I need to do fast consuming and for Kafka the situation is quite good, for rabbitmq I have stuff that works but I needed to hack with a knife to get stuff working the way I wanted.
Everything's quite low level in the end and there hasn't been a task I couldn't solve either by myself or finding a library and doing a couple of pull requests to get things running. The hardest was to implement a working HTTP-ECE for web push notifications by reading and understanding all the RFC drafts, but the project[3] teached me a lot. Basically all my consumers deal with millions of events every hour, using 0.1 cpu's and about 50 megabytes of RAM. Uptimes reaching 2-3 months if I don't need to update the software.
[0] https://github.com/hyperium/hyper
Net/http is still less mature, but this an absolutey massive golang ecosystem. You can certainly write http clients and servers today.
Unicode, string processing, regex—all of these have performant, stable implementations, either in std or in crates. Recently i’ve been doing audio processing in rust; the code there is at least as good as the equivalent in go. Overall I’d say I haven’t had issues finding a package for something in a couple years, though the quality varies from “has a full support community” to “DIY if you need something”.
However, rust really shines with datastructures. Heap? Btree? Doubly linked lists? It’s all high quality, performant, and type-safe code (though the internals are unsafe as hell), which was my major pain point in go. Doing anything with a typed datastructure feels a lot like copy/paste coding in go, though apparently templates formalize this.
I've been reading more and more about Rust and while I've fallen in love with its premises I still fail to tackle the real world task of starting to do stuff with it. Would you give me some pointers?
I work with Python and C professionally and have more than 15 years coding backend stuff for un*x systems, to give you some bg.
thanks in advance!
But, rust has everything you need: find-grained control over memory and ownership, inline assembly and intrinsics (I haven’t attempted autovectorization yet), bindings to common audio formats, and excellent parallelization. The datastructures are extremely expressive, especially if you come from a C background; so far I really just want better VLA support, which is mostly annoying to work around (you have to manually poke the values into memory at the correct offset in an unsafe block).
I did find that tokio-rs wasn’t suitable as I had hoped for realtime work with an async/future api and fell back onto ringbufs locked with mutexes. The good news is you can wrap that in futures itself and have a great async api that does its work eith shared memory, minimizing the viable races.
Rust will definitely be a player in the audio workstation world; right now you’ll be implementing the bindings yourself, but you can go out and write plugins today so long as the ABI is C.
Our REST stuff with quite minimal traffic is done with Clojure.
Update: about:support says not ready for Android
To enable, toggle 'layers.acceleration.force-enabled' as well as 'gfx.webrender.enabled'
edit: It's also working through my Nvidia 950m (through bumblebee), although subjectively it seems to have a little more lag this way.
[1] http://blogs.valvesoftware.com/abrash/down-the-vr-rabbit-hol...
"People talk about 60fps like it's the end game, but VR needs 90fps, and Apple is at 120. Resolution also increasing. GPUs are the only way. Servo can't just speed up today's web for today's machines. We have to build scalable solutions that can solve tomorrow's problems."
Here's for example an early demo showing Wikipedia at ridiculous frames per second (starts at 0:26:00): https://air.mozilla.org/bay-area-rust-meetup-february-2016/
In the video, he says 500 FPS, but assuming there's no more complicated formula behind this, I think it would actually be 2174 FPS. (0.46 ms GPU time per frame -> 1/0.00046s = 2173.913 FPS)
Baby steps.
Also go read everything else Lin Clark has done. They're pretty great.
Is this going to work at all on Linux?
[1] https://www.rust-lang.org/en-US/install.html Select 2, for custom install, and change to "nightly". My host triple was x86_64-unknown-linux-gnu. [2] https://github.com/pcwalton/pathfinder [3] TextDemoView.initContext (view.ts:324) Uncaught (in promise) TypeError: Cannot read property 'createQueryEXT' of null .
EDIT: The three demos are: some text; the SVG tiger; some planar text in 3-space.
That is, I have to be lucky for this benchmark [1] to not crash my "regular" Firefox Nightly. With WebRender on the other hand, the benchmark becomes entirely unimpressive, as if you were just playing a pre-rendered video.
My system has an Intel i5-3220M with HD Graphics 4000 (was midrange for a laptop in 2012) and 4GiB RAM. OS is openSUSE Tumbleweed with a KDE Plasma + bspwm combination as desktop environment, so no desktop compositor (no idea if that makes a difference).
Also fonts look slightly different, like I am browsing through using my Linux machine.
Sadly, seeing the state of the industry, people will just use this as an excuse to continue write more and more sloppy code that would perform terrible even on the newest and faster WebRender version.
In other words, this version might run current websites at 60 fps. But wait a few years, and it will become the norm that a lot of websites render at 10fps or less.
There should also be a way to punish developers who fail to run their sites at 60fps on WebRender, similar to how browsers will start to punish sites that run without https.
For example, if a site fails to run at 60fps for a few seconds, show some kind of alarm on the address bar that this site is very slow and might crash the browser.