Fast 2D Rendering on GPU
raphlinus.github.io
raphlinus.github.io
With that being said, I'm looking to delve deeper into this subject. Do you have any recommendations on where to start? Whenever I look beyond simple drawing API's, the focus seems to be entirely on 3D rendering (which is interesting but not my main focus right now).
In the meantime, antigrain.com is one good (if old) source. The original PostScript "red book" was extremely influential in its time (it's where I learned a lot of this stuff) but is quite dated now. Best of luck, and I'm also happy to field requests for more specific areas. For example, for color theory (an important aspect of 2D graphics!), handprint.com is quite a remarkable resource.
such an unfortunate state of affairs!
i am currently learning how to render graphics using the GPU on my mac using apple metal. what i am getting is that the GPU has been optimized for 3D rendering?! GPUs make no provision or easy way for rendering 2D graphics?
it makes no sense to me... that's where you start...
as far as i understand, 2d and 3d have literally zero to do with each other in how they are rendered. one is a bunch of triangles. the other is lines, curves, thickness, gradients, and fonts (which are essentially little programs)
You can reduce all these to drawing triangles.
Raph is describing an architecture where path evaluation happens on the GPU, without being baked to triangles.
You can, people have tried this, and it sucks. The main problem is that the conversion of Bézier paths to triangles is a hard problem with lots of conditional branching. Even when you do it, there is the other problem of rendering triangles with really good antialiasing, MSAA forces a compromise between performance and quality. By contrast, piet-gpu does an exact-area calculation for antialiasing.
So it's not a question of whether you can do it, but whether it works well, and approaches like piet-gpu absolutely stomp triangles.
Easier than you think. Here's couple lines of pixel shader that does that, with really good antialiasing and without MSAA:
https://github.com/Const-me/Vrmac/blob/master/Vrmac/Draw/Sha...
And: historically they've been computed mostly on CPU, but I think it's time for that to change.
It would be great to wait a bit for OS & GPU power management to evolve before biting the bullet on that. My laptop goes from 6 to 2.something hours of battery as soon as I have a GL context opening somewhere, likely because it seems to power on its discrete GPU automatically in that case.
maybe ? the computer on which this happens is a 1070. But please be aware that series 10 are a very small percentage of people. The average laptop of non-tech people around me is easily 8 years old, often on their 2nd or 3rd battery... and these people won't be able to complain easily to anyone when their new battery's life suddenly is halved because of $SOFTWARE.
At any rate, the 2D graphics we expect now are a lot more complex than the unantialiased lines, blits, and fills of old.
The various libraries such as freetype for font rastrization only works on CPU.
Plenty of work should be done to research and implementation is left to be done in order to use the GPU more widely.
Of course CPU rendering is always more straightforward than GPU, the higher performance comes at a significant cost in complexity.
1) tricky to shoehorn 2D graphics onto the APIs that GPUs provide and
2) really not needed. I can easily render eg: a world map with hundreds of thousands of lines at hundreds of frames/second with one core.Thanks for publishing this, it's awesome work! I'm looking forward to progression to wgpu hinted in the Github README.
It is likely that CPU-side encoding can be made more efficient, though, by just filling in quantities to a template, rather than encoding from scratch.
https://github.com/servo/pathfinder/pull/350
hopefully you guys aren't doing identical work in parallel (no pun intended) :)
- Don't libraries like Skia, Qt, Cairo use GPU rendering? I've always assumed so. I mean, this is 2020, GPUs have been around for decades.
The problem is that primitives artists use are different. 3D rendering tends to all consist of polygon meshes, which are relatively easy to render. 2D rendering (basically) consists of Bezier paths, which are harder. The equivalent in 3D, which is adaptive subdivision, is not really a solved problem in real-time either.
Additionally, 2D rendering quality tends to be more important than 3D rendering quality. Whereas you can get away with 4xMSAA or hacks like FXAA in 3D, true 16xAA (without hacks) is the absolute minimum for 2D rendering quality nowadays, and even it isn't considered great for some tasks like font rendering (Pathfinder and piet-gpu both use analytic AA which is effectively 256xAA).
> - Don't libraries like Skia, Qt, Cairo use GPU rendering? I've always assumed so. I mean, this is 2020, GPUs have been around for decades.
There's a difference between renderers with GPU support and renderers that are oriented around using the GPU efficiently. In many cases this results in an order-of-magnitude speedup. On the GPU, state changes are expensive, and many such renderers that have GPU support don't really go out of their way to avoid them. There are also occlusion culling optimizations that most renderers don't do, but piet-gpu and Pathfinder do.
The easy way is to just tesselate the font into polygons -- but this tesselation often depends on the zoom level. The same thing can be said about implicit curves etc.
Most libraries do use GPU's for basic draw operations (e.g rendering a gradient), but to build something like Photoshop -- you need much more complexity.
Also, I see a lot of variations of this question, but I should state this more clearly. There's been accelerated graphics in one form or another for a long time, but what I'm doing is a completely different type of thing. In my world, on the CPU you just encode the scene into a binary representation that's optimized for GPU but in many ways is like flatbuffers, and then the GPU runs a highly parallel program to render the whole thing. In previous approaches, the CPU is deeply involved in taking the scene apart and putting it back together in a form that's well suited to relatively dumb pixel pipes. Now that GPUs are really fast, that approach runs into limitations.
It also depends what you're trying to do. I'm focusing here on dynamic paths (and thus font rendering), while most of the libraries optimized for UI put text into texture atlases and then use the GPU to composite quads to the final surface, something they can do well.
https://blog.mecheye.net/2019/05/why-is-2d-graphics-is-harde...
As a Qt user / developer - you can use the GPU for your app, e.g. with Qt Quick or QGraphicsView, but there are sometimes good reasons to stick to CPU & software rendering, e.g. it is somewhat common to want to have intertwined "native" OS widgets (which are all CPU-rendered raster things) and custom GPU-drawn scene - this is a case where things fall a bit apart.
Another thing is that pretty, freetype-like font rendering is super expensive when you have a lot of text to show, and can't really be done (at least I have definitely not seen infinality-level beauty from the state of the art) on the GPU yet... next to that filling some rects (read: 95% of UI) with SSE/AVX/AVX2 as Qt does is stupidly fast.
This is exactly what Pathfinder does. pcwalton is on this thread and is the main author of that.
https://github.com/servo/pathfinder#features says:
> Advanced font rendering. Pathfinder can render fonts with slight hinting and can perform subpixel antialiasing on LCD screens. It can do stem darkening/font dilation like macOS and FreeType in order to make text easier to read at small sizes. The library also has support for gamma correction.
from the screenshots I saw so far, pretty much not.
> It can do stem darkening/font dilation like macOS and FreeType in order to make text easier to read at small sizes.
there are tons of different ways to do that. Even freetype has a few different algorithms to do it, some not even merged if I'm not mistaken, which give wildly different results
I rather like the demonstration of rendering including subpixel rendering at https://twitter.com/pcwalton/status/971475785616797698, as well.
I don't intend at all to cast doubt on pcwalton's abilities - the work is brilliant without any hesitation.
But I wonder how that is possible given that "platform rendering" pretty much has changed every other macOS version and every Windows version ("ClearType" from WinXP is definitely not "ClearType" from Win10) ; and let's not talk about the customization abilities of freetype which makes rendering on any two linux boxes also entirely distinct.
Those libraries use GPU rendering using textures, like distance fields(Qt) or simply calculate fonts in 2D and draw it on textures using the GPU, like Apple or cairo usually do.
Distance fields are blurry when you have small fonts on display.
Things like calculating the exact area order the 2D curve is something you could easily do in CPU, but it is extremely difficult to do on the GPU. You need decoupled data in order to parallelize it.
It bothers me how little progress has been done on "shading" languages front compared to overall many-core computation models and capabilities over the years. And, that is despite the fact that shaders are very often where the most time is spent in modern workloads.
Compute with Vulkan is another story. It offers some nice abstractions, but it shows that it's mostly intended for async-compute/work-offloading for rendering, IMO. Too much fruction.
The "impedance mismatch" is that you (generally) have to write in a style to extract lots of parallelism. This tends to be very different than the way you'd write scalar CPU code, but not completely alien to me as it has a lot of similarity with the way you'd write SIMD. I've pretty much gotten the hang of it now. I'm thinking of a blog post of redoing path_coarse.comp from its current basically scalar style to a more parallel version, as that would I think illuminate the issues.
[1] http://comparch.gatech.edu/hparch/papers/gera_ispass18.pdf
Suppose I was to make an ebook reader designed to be modelled more closely after paper, on a device like the Surface Book’s 13″ 3000×2000 display, showing two pages side-by-side. Each page might contain something like 2,000–2,500 letters. I want to be able to flip through pages like I might with a paper book, so that I might be roughly completely rendering several pages at once, and perhaps parts of several more pages; ideally it might render the page like a real 3D page, but if that’s too troublesome I’d settle for an affine transformation while flipping. Assume that the layout of the pages, with all the shaping, is all done ahead of time and is in memory.
I’ve never seen anyone attempt anything like this before. In the old way of doing things, I think anyone attempting this would render each page to a bitmap and use that as a GPU texture, and I think that could provide acceptable performance (unless you were flipping rapidly through hundreds of pages, because that’d take a lot of GPU memory to keep), but I imagine that the quality of the rendering mid-turn would be fairly atrocious—it could be the sort of thing where the appearance of the text subtly changes half a second after you finish turning the page, as it switches from one renderer to a slightly different one.
Would the performance of this new approach be sufficient to render what I describe at 60fps, while rendering each frame perfectly?
I was talking about the Surface Book; its Intel Core i7-6600U has Intel HD Graphics 520 as its GPU; probably not too far off your 630’s results. (And I’m interested in what integrated graphics can do, more than a dedicated GPU.)
My guess based upon your paper-1 results is that this is probably just barely possible with integrated graphics, so long as you employ a few tricks to reduce the amount of work required.
Doing a warp transformation in the element processing kernel would probably work just fine, and give you realistic movement and razor-sharp rendering.
I'd love to see such a thing and would very much like to encourage people to build it based on my results :)
Look at phone interfaces; even if you scroll fast, they don't add motion blur, and on high framerate and/or low-persistence displays, you can read things while it scrolls.
I work on software that is used to create visual effects in real-time. It's used in live events such as concerts, clubs, corporate presentations. There you want to be able to render high-quality text in high resolution, even higher than 4K, and still allow a performer to add effects on text such as zooming, scrolling, doing 3D transition effects...
Today the best way to do this is via a font atlas but it does not look nice when zooming. Also if you need non-European characters it may heavy preparing the font atlas.
So I am looking for a library that would allow this kind of rendering manipulation for text rendering. Do you have any suggestions?
I have nothing against licensing and it looks like things are moving there. Do you have a recommendation for Metal / DirectX support for that kind of rendering facilities? What matters more for me is performance.
The reason you don't see this is because no one runs a 3d engine in their text readers. Fonts are not usually shared as sprite sheets. You also lose subpixel font rendering.
I expect a screen that supports or is designed for pen drawing to be somewhat more likely to be above 60Hz. All of these things are niche things that not many care about, but in any case, it’d be nice to be able to do better. And like with Formula 1 race cars, benefits from high-end techniques tend to trickle down to other more mainstream targets in time.
Bottleneck is in the input layer. While using a desktop computer look at the mouse cursor, now move the mouse - hardware accelerated 2D graphics, imperceptible latency.
Low latency thru orthogonal multiplexing https://www.youtube.com/watch?v=t1VcC9_yhc0
But that pretty much what we have now even on modern systems.
For desktop UI purposes (but not for games) GPU acceleration started to become critical only relatively recently - on high-DPI screens. Number of pixels jumped 9 times between 96 ppi and 320 ppi (Retina) screens. CPUs haven't changed that much in the time frame so you may experience slow rendering on otherwise perfect screen picture.
What profile? They did RAM to RAM DMA transfer exactly for video acceleration with the Agnus chip, which would access memory through DMA channels to perform 2D blitting.
What I think he was thinking about was Blitter in general, and not particular implementations using DMA controller. First Blitter accelerated 2D graphics I read about were done on Xerox Alto using microcode.
There's also a 3D fixed pipeline that was all that the first 3D GPUs could do and it is also removed from modern GPUs.
A game I still play to this day is called rFactor, and it runs much faster in DX9 mode than in DX8 or DX7 mode in my hardware. At the very least, it behaves that way in this week's test. This means my hardware doesn't have that old fixed pipeline.
Some people with old hardware can only run it with decent frame rate in DX7.
So the answer to your question is: yes, 2D has been accelerated since forever, but now we have general purpose 3D programmable video cards, and they are so good that now we only have general purpose 3D programmable video cards and the other forms of acceleration are obsolete.
Except for things like video codecs and DRM.
It's just that my Win3.1 computer was too slow for me to remember the acceleration.
Also remember the glorious Fury3 game.
[1] https://en.wikipedia.org/wiki/X.Org_Server#2D_graphics_drive...
[2] https://en.wikipedia.org/wiki/Direct2D
[3] https://docs.microsoft.com/en-us/windows/win32/direct2d/comp...
I recently had to do a proof of concept displaying a huge gantt chart (I really mean huge!).
Normally, 2D drawing has some disadvantages with textures and scaling/rotating.
This proof of concept included testing out various methods on the <canvas>, including both using WebGL and standard 2D calls.
To my surprise, I was able to get the standard 2D calls way faster, even with texturing and rotations (which I did not expect). All browsers (Chrome, Firefox, Edge, IE) do GPU optimizations with standard 2D canvas drawings.
Only when I was putting a shitload of different things on there (gantt chart looked more like a barcode at this point), the WebGL implementation started to outperform the 2D calls.
Remark that I didn't go into programming shaders. In the end, it wasn't necesarry since the normal canvas calls already supported plenty of performance.
Also remark that I'm not just some random dude with no experience. I'm creating a game dev tool https://rpgplayground.com, which runs fully in the browser using Haxe and the Kha graphics library, and made plenty of games during my career.
The article and most discussion here however concerns rasterization, specifically path rendering, which is useful when rendering things like fonts and image formats like SVG. You use primitives such as lines and curves to form outlines of objects and then fill them. It's a problem that does not very easily translate into rendering textured triangles (which is what the GPU is best at) or copying/moving screen regions, so there's some work in doing it in a performant way still, and a lot of it typically happens on the CPU.
I've been finding Patrick's pathfinder work fascinating so it's nice to have yet another project to follow.
In general I've found NanoVG to have comparable performance to the master branch of Pathfinder. NanoVG currently performs better at text (PF's implementation is designed to get international text right first and optimization hasn't really been done yet), while PF generally does better at complex vector workloads like the tiger.
https://github.com/styluslabs/nanovgXC
If you give it a try, let me know how it works for you.
Pathfinder in there is actually using regular old GPU rasterization with triangles and so forth, and I'm fairly confident it's about as fast as you can go at D3D10 level (i.e. no compute shaders) without sacrificing quality. Note that the numbers can vary wildly depending on hardware. On my MacBook Pro, with a powerful CPU and limiting myself to the Intel integrated GPU only, Pathfinder is actually about equal to the GPU compute approach on a lot of scenes like the tiger, though it uses a lot of CPU.
From my intuition, this seems pretty specialized for vector-like rendering, with a lot of small bezier shapes.
That's fixable, take a look: https://github.com/Const-me/Vrmac#vector-graphics-engine
My CPU process is moderately expensive, and I use quite a few tricks to reduce overdraw, e.g. both draw calls (the complete vector image, however complex, takes 2 draw calls to render) use hardware depth buffer and early Z rejection.
Problem is that any practical 2D rendering solution shall support as GPU as CPU rendering unfortunately.
It would be interesting to see any GPU equivalent of something like AGG by Max Shemanarev, RIP.
Ideally GPU and CPU rendering backends should have pixel perfect match that makes "adding new 2D backend" tricky at best.
[1]https://gitlab.com/gnachman/iterm2/-/wikis/Metal-Renderer
I’ve looked before and although OpenGL is good for windowing, widgets etc, it’s not great for sub pixel rendered/anti aliased text.
I don’t use compute shaders, tessellating input splines into polylines, and building triangular meshes from these.
Sciter (https://sciter.com) on MacOS uses Skia/OpenGL by default with fallback to CoreGraphics.
It is possible to configure Sciter to use CoreGraphics on MacOS to compare these two on the same UI by using
SciterSetOption(NULL, SCITER_SET_GFX_LAYER, GFX_LAYER_CG);
I think it is safe to say that Skia/OpenGL is 5-10 times more performant than CG on typical UI tasks.Some discussion here https://arstechnica.com/civis/viewtopic.php?t=1285571 , though a lot of guessing.
In any case Acrylic/Vibrancy effects or "blur-behind background" (like on screenshots here: https://sciter.com/sciter-4-2-support-of-acrylic-theming/) are achievable only on GPU and on DWM level.
They might use it for other stuff like glyph compositing (I think this is one reason they got rid of RGB subpixeling, to make it more amenable to GPU), but last I profiled it, it was still doing a lot of the pixels on CPU, as others have stated.
A lot of the academic literature (Massively Parallel Vector Graphics, the Li et al scanline work) is dependent on CUDA, but that's just because tools for doing compute on general purpose graphics APIs are so primitive. I have a talk and a bunch of blog posts on exactly this topic, as I had to explore deeply to figure it out. See https://news.ycombinator.com/item?id=22880502 and https://raphlinus.github.io/gpu/2020/04/30/prefix-sum.html for more breadcrumbs on that.