> Let me state up front that I have no idea why Wayland would have this additional latency
It's not really hard to guess at - it's probably caused by abstraction layers resulting in multiple layers of buffering. Of course, it's all just speculation until proven by experiment, so good on TFA for doing some of that.
Retro hardware from the C64 or NES era was single-tasking with such predictable timing you could change the video resolution partway through the screen refresh and have two resolutions on screen at once. If you want, you can check the player input and update Mario's position right before he gets scanned out - the minimum possible latency by design is zero frames. (Of course, in practice the whole screen data is buffered up during vblank, which is the same latency as using a framebuffer. Also the NES didn't allow changing video registers outside of blanking periods. But the C64 did.)
X11 isn't that close to the metal, but it's designed with a single level of buffering (assuming a compositing window manager isn't used). There is a single framebuffer covering the whole screen. Framebuffer updates are completely asynchronous to scanout, so there is around 0.5 frames of latency on average between when a pixel is changed and when it's scanned out. Apps borrow pixels from this big global framebuffer to draw into. An app sends a "draw rectangle" command relative to its own borrowed space, and the server calculates where the rectangle should go on the screen, and draws it into the framebuffer, and then with an average latency of half a frame, that part of the framebuffer is scanned out to the screen.
On Wayland, there are more layers. The app has to draw into its own pixel buffer (noting that it is probably double- or triple-buffered and has to redraw the entire window rather than relying on non-changing stuff to still be there) and then once the compositor receives the buffer from the app, it has to do the same thing again, copying all the per-app buffers into a big screen-size buffer, which is probably also double- or triple-buffered, and once the vblank hits, the most recent full frame is drawn to the screen. It's just more work with more layers, and there's no surprise that latency is higher. It's hardly the first or 100th time that an elegant software architecture resulted in the computer having to do more work for the sake of keeping the architecture pure.
Something important to note is that tearing reduces latency. If new data is available while the scanout is partway through a frame, you have the choice to draw it immediately, causing tearing, or wait until the next frame, increasing latency. Always doing the high-latency thing is an explicit design goal of Wayland. (A third theoretical choice is to always time things precisely so that the new frame is calculated right before the end of vblank, but... in a modular multitasking architecture with apps written by different organizations and unpredictable CPU load, good luck with that.)
Now the really big caveat: I'd be surprised if GNOME's X11 WM wasn't compositing. If it's compositing, then X11 behaves similarly to Wayland but with even more IPC and what I just said is no reason it should have one frame less latency on average. Still something to think about though.