Once support for hardware planes becomes more common in Wayland compositors, this can be tied to ultimately allow no-copy rendering to the display for non-fullscreen applications, which for video playback (incl. likes of Youtube) equals to reduced CPU & GPU usage and less power draw, as well as reduced latency.
What.
The original design of X actually encouraged a separate surface / Window for each single widget on your UI. This was actually removed in Gtk+3 ("windowless widgets"). And now they are bringing it back just for wayland ("subsurfaces"). As far as I can read, it is practically the same concept.
GTK3 got rid of windowed widgets because Keith Packard introduced the Xrender extension, which basically added 2D compositing to X, which was the last remaining use for subwindows for every widget.
The excuse that was used when introducing windowless widgets is to reduce tearing/noise during resizing, as Gtk+ had trouble synchronizing the resizing of all the windows at the same time.
Yes, that's the point. When you can tell Xrender to efficiently composite some pixmaps then there's really no reason to use sub-windows ever.
>You make your toolkit's programmer's life more complicated, not less, by having windowless widgets (at the very minimum you now have to complicate your code with offsets and clip regions and the like).
No, you still had to have offsets and clip regions before too because the client still had to set and update those. And it was more complicated because when you made a sub-window every single bit of state like that had to be synchronized with the X server and repeatedly copied over the wire. With client-side rendering everything is simply stored in the client and never has to deal with that problem.
There is, or we would not be having subsurfaces on Wayland or this entire discussion in the first place.
Are you seriously arguing that the only reason to using windows in Xorg is to have composition? People were using Xshape/Xmisc and the like to handle the lack of alpha channels in the core protocol? This is not what I remember. I would be surprised if Xshape even worked on non-top level windows. heck, even MOTIF had windowless widgets (called gadgets iirc), and the purpose most definitely was not composition-related.
No. The subsurfaces in Wayland are only designed for two things:
1. Direct scan-out, as in TFA. (Because a subsurface can be directly translated to a dmabuf)
2. Embedding content from one toolkit/library into another. (Because without it, lots of glue code would be needed)
It's discouraged to use them otherwise, they would complicate things for no benefit.
>Are you seriously arguing that the only reason to using windows in Xorg is to have composition?
If you mean sub-windows, yes, that and consequently because of the way that XRDB worked. I don't see why you would ever use them just for input events within the same toolkit, they don't do anything special there.
https://www.donhopkins.com/home/catalog/unix-haters/x-window...
That's what happened when I simply resized stock xcalc several times in a row to different sizes and aspect ratios. I wonder if they've finally figured out how to fix that bug in 35 years?
That was why I was compelled to hack an X11 window manager to take a command line argument telling it which window id to treat as the root window, then ran xcalc, discovered its window id with "xwininfo" or some such utility, then ran the window manager on the calculator, putting window frames around each of the calculator's buttons, so you could resize them, move them around, open and close them to icons, etc! That was a truly customizable calculator.
Sure it lets you do fast 2d acceleration but we don't use 2d accel infrastructure anywhere anymore.
Subsurfaces have been in Wayland since the beginning of the protocol.
This is simply getting them to work on demand so we can do something X (or Xv) could never do since Drawables would get moved to new memory (that may not even be mappable on the CPU-side) on every frame.
And that's to actually use the scanout plane correctly to avoid powering up the 3d part of the GPU when doing video playback on composited systems.
how about Graphics Offload?
Yes, you're technically right that this would have been possible years ago but it wasn't actually ever done, because X11 never had the ability to do it at the same time as using compositing.
You really need to add context to these statements, because _right now_ I am using through Xorg a program which uses a frigging colormap, which is as non-RGB as it gets. The entire reason Xlib has this "WhitePixel" and XGetPixel and XYPixmap and other useless functions which normally fetch a lot of ire is because it tries to go out of its way to support practically other-wordly color visuals and image formats. If anything, I'd say it is precisely RGB which has the most problems with X11, specially when you go more than 24bpp.
> there for attaching GPU buffers
None of this is about the GPU, but about about directly presenting images for _hardware_ composition using direct scan-out, hardware layers or not. Exactly what Xv is about, and the reason Xv supports formats like YUV.
> There isn't any incentive to implement this in X11 either because X11 is supposed to work over the network and none of this stuff would
As if that prevented any of the extensions done to X11 in the last three decades, including Xv.
That doesn't change what I said. The colormap is mapping indexes to RGB palette entries and technically doesn't even support any other color spaces. Nothing about that is "non-RGB". The other visuals are just various other ways to do RGB. If this isn't making sense to you, think about how this would be implemented in the driver.
>None of this is about the GPU, but about about directly presenting images for _hardware_ composition using direct scan-out
I'm sorry? What do you suppose is doing the direct scan-out to the monitor if not the GPU?
>As if that prevented any of the extensions done to X11 in the last three decades, including Xv.
That's not relevant, XV actually does support running over the network as long as you don't use the SHM functions. But regardless, yes, it actually did. The main example being indirect GLX which hasn't been updated in decades because it's not feasible to run it over the network anymore.
You need to dynamically change stacking of subsurfaces on a per-frame basis when doing the CRTC.
I agree though it would require a lot of changes to the server and no one is in the mood (like, dynamically decide whether I composite this window or push it to a Xv port or hardware plane? practically inconceivable in the current graphics stack, albeit it is not a technical X limitation per-se). This entire feature is also going to be pretty pointless in Wayland desktop space either way because no one is in the mood either -- your dmabufs are going to end up in the GPU anyway for the foreseeable future, just because of the complexity of liftoff, variability of GPUs, and the like.
You'll need API to remap the Drawable to the scanout plane from a compositor on a per-frame basis (so when submitting the CRTC) and the compositor isn't in control of the CRTC. So...
Nothing in Xorg is a minor change because it has to be tested with every window manager and compositor before you can even think about merging.
https://gitlab.freedesktop.org/wayland/wayland-protocols/-/t...
"This global is a factory interface, allowing clients to inform which type of presentation the content of their surfaces is suitable for."
Note that "global" refers to the interface, not the setting. Which Wayland compositor has the equivalent feature of "xrandr --output [name] --set TearFree off"?
Here you are saying that latency of 250ms is unnoticeable, which is utter nonsense.
Yeah, that is utter nonsense. Just try it for yourself, instead of pulling statements like that out of thin air. You would notice a difference in tens of ms even when editing text. Why do you think people cannot stand writing with VSCode for example?
Re: moving/evolving shapes, I did not think I had to clarify that the brain is a massively parallel system with multiple modes of operation. Editing text does not require you to reprocess all visual signals from scratch, because that is not how the visual cortex works. The perceived latency when editing text is between pressing a key and your brain telling you "my eyes have detected a change on the screen. I will assume that it is the result of me pressing a key". It does NOT take 250ms to make this type of assumption, and is basically how our vision operates. It's a prediction engine, not a CCD sensor.
Which people? Every recent study I've seen shows VSCode as the most popular code editor by a large margin. Maybe latency isn't as important as you think?
>Are you saying that latency in the order of 250ms when editing text is unnoticeable?
No. Sorry for the info dump here but I'm going to make it absolutely clear so there's no confusion. The latency of the entire system is the latency of the human operator plus the latency of the computer. My statement is that, assuming you have a magical computer that computes frames and displays pixels faster than the speed of light, the absolute minimum bound of this system for the average person is 250ms. You only see lower response time averages in extreme situations like with pro athletes: so basically, not computer programmers who actually spend much more time thinking about problems, and going to meetings, than they actually spend typing.
Now let's go back to reality: with a standard 60Hz monitor, the theoretical latency added by display synchronization is a maximum of about 16.67ms. That's the theoretical MAXIMUM assuming the software is fully optimized and performs rendering as fast as possible, and your OS has realtime guarantees so it doesn't preempt the rendering thread, and the display hardware doesn't add any latency. So at most, you could reduce the total system latency by about 6% just by optimizing the software. You can't go any higher than that.
However, none of those things are true in practice. Making the renderer use damage tracking everywhere significantly complicates the code and may not even be usable in some situations like syntax highlighting where the entire document state may need to be recomputed after typing a single character. All PC operating systems may have significant unpredictable lag caused by the driver. All display hardware using a scanline-based protocol also still has significant vblank periods. Adding these up you may be able to sometimes get a measurement of around 1ms of savings by doing things this way, in exchange for massively complicating your renderer, and with a high standard deviation. Meaning that you likely will perceive the total latency as being HIGHER because of all the stuttering. This is less than 1% of the total latency in the system and it's not even going to be consistent or perceptible.
Now instead consider you've got a 360Hz monitor. The theoretical maximum you can save here is about 2.78ms. This can give you a CONSISTENT 5% latency reduction against the old monitor as long as the software can keep up with it. Optimizing your software for this improves it in every other situation too, versus the other solution which could make it worse. If it doesn't make it worse, it could only save another theoretical 1% and ONLY in a badly perceptible way. It just doesn't make sense to optimize for this less than 1% when it's mostly just caused by the hardware limitations and nobody actually cares about it and they're happy to use VSCode anyway without all this.
So again, you can avoid these accusations of "utter nonsense" when it's clear you're arguing against something that I never said.
>The perceived latency when editing text is between pressing a key and your brain telling you "my eyes have detected a change on the screen.
Your brain needs to actually process what was typed. Prediction isn't helping you type at all, if it did then the latency wouldn't matter anyway. If you're not just writing boilerplate code then you may have to stop to think many many times while you're coding too.
Only if the DE or application can't keep up. Oh, it overran the frame time, so let's just dump half of it to screen...
I'm still baffled that this is a problem. I tried ray tracing a hierarchy of rectangles decades ago and it was pretty fast. Should be better than real time today, and that's doing it the expensive way. I suppose the challenge is composting stuff from different processes every frame.
> I suppose the challenge is composting stuff from different processes every frame.
This is a significant part of the issue, though the fact that synchronization essentially means waiting is also another one (without synchronizing you get partial frames but those partial frames are still feedback for the user that makes the desktop feel more responsive).
And you're not limited to two images. As the frame rate increases, the number of images increases and the tear becomes less noticeable. Blur Busters explains:
https://blurbusters.com/faq/benefits-of-frame-rate-above-ref...
When touch typing, fingers work decoupled from the eyes anyway, unless you are waiting for the intellisense or copilot prompt, that is usually constrained by language servers anyway, not the framerate.
Most humans will actually prefer the "slightly faster" option. (Obviously if you can do both, then they'd prefer that; but given the trade-off...)
For example, with a 60 Hz display and vsync, game actions might be shown up to 16 ms later than without vsync, which is ages in FPS.
I've seen no game consoles that allow you to turn vsync off, because it would be awful. No idea why this placebo persists in PC gaming.
Also only two of the four changes mentioned in the mail are about CVEs.