Incidentally, one thing I noticed when I was trying to port Linux GPU drivers to Windows some time ago is what appeared to be an excessive amount of indirection; there are so many layers and places where things could be simpler.
Incidentally, one thing I noticed when I was trying to port Linux GPU drivers to Windows some time ago is what appeared to be an excessive amount of indirection; there are so many layers and places where things could be simpler.
On the client side, I think most Linux apps still draw their UIs on CPU, usually accelerated with SIMD. Firefox and Chrome (I think SkiaGL is enabled on Linux?) are exceptions; they use OpenGL and/or Vulkan to draw their UI. Video playback is a different beast and in theory relies on vendor-specific extensions to decode the video in hardware. However, the last time I looked at Linux video decoding (which was years ago), the drivers were awful and interfacing with each vendor's APIs was a huge pain, and so most apps just did video decoding on CPU. (Besides, the Linux ecosystem prefers open codecs, and hardware has only recently gotten support for non-patent-encumbered video formats.)
Firefox has Web Render running on top of ANGLE which is a generic OpenGL layer that converts the OpenGL calls into native platform calls. ANGLE is a Google project and it is the base library for Skia which is used by Chromium to render everything. IIRC Qt / QML also uses ANGLE for Windows.
Rasterizing small glyphs is super fast anyway, there's not much of a need to accelerate it if you can just cache the glyph bitmaps.
Nowadays VA-API is near universally supported, and any half-decent video player uses it to do hardware decoding.
(No, OpenVG is not viable. No, Xrender is not viable. cairo and Skia both use the 3D hardware in combination with a CPU render engine.)
The scanout-time hardware was often less useful that you might think - only in dynamic scenes where the GPU is otherwise idle (like playing video possibly with a static UI overlay was the premier use case).
For static scenes it's more efficient to render out to a buffer (using the GPU as the scanout overlay pipes often had limited feedback capability) and just output that using overlays disabled. It didn't take many frames for that to be worth it.
For apps that were animating or otherwise updating it's window, most UI toolkits used the GPU for widget rendering. And often the scanout pipes didn't hook into the (relatively large) system caches like the GPU did, so there were times it was again faster to composite the screen on the GPU to a single scanout buffer than flush already cached data, the get the scanout hardware to read it back from the memory bus.
And there weren't as cheap as people thought - one stat I remember was that the total area of the GPU on the omap4 platform was smaller than the display pipes. Though that is now a pretty old chip, and always had a bit of focus on "multimedia".
The overlay capabilities of the modern Snapdragons are also quite absurd. They support like upwards of a dozen overlays now and even have FP16 extended sRGB support. Some HWCs (like the one in the steam deck) even have per plane 3D LUTs for HDR tone mapping (ex https://github.com/ValveSoftware/gamescope/blob/master/src/d... )
The composition is bandwidth heavy of course, but for static scenes there's a cache after the HWC in the form of panel self refresh.
Personally, I don't see it as a "2D drawing API", it doesn't accelerate anything special about 2D, only blits and transforms, which a 3D API will eat for breakfast.
Edit: But actually, I couldn't find references to anything similar besides VBE/AF which even when current got almost no support directly in hardware, so folks had to resort to hardware-specific DOS TSR's. I'm not sure if there's anything newer than that.
The bigger issue is that there's little reason to farm vector graphics rendering out to the window server in the first place. The main reason would be to avoid a window blit on HiDPI displays. But the tradeoff is that the XRENDER API is all you get, and usually apps have more sophisticated needs than what it can provide. For instance, browsers can't really use XRENDER nowadays because there's no way to describe CSS 3D transforms in it. And if you use it you're at the mercy of the window server to implement it reasonably, which is not a safe assumption. (A lot of the reason Chrome on Linux was faster than Firefox in the early days is that Firefox used XRENDER, while Chrome rendered on CPU. I remember at least one engineer at Mozilla who was bitter about that, after putting in all the work to make Firefox use it only to have it be a net loss.) In any case, you can avoid the window blit by simply using scanout compositing, as detailed in my other reply, so there is really is no compelling reason to reinvent XRENDER.
The semantics of Xrender simply don't match with what modern GPUs give you, even ones with 2D pipelines.
[0] https://gitlab.freedesktop.org/search?search=sna_pixmap_move...
But if there was an API for 2D acceleration that was actually supported (and could be used simultaneously with desktop composition), then it could be added in to something like SDL then suddenly applications would support it.