Proposed Windows NT sync driver brings big Wine/Proton performance improvements
gamingonlinux.com
gamingonlinux.com
I’m surprised the kernel would be okay with this impl. It reminds me of when OpenVZ tried to upstream their big blob of container logic. The kernel maintainers forced them to part it out into individual patches so it could be reworked (and we ended up with cgroups and namespaces.)
If Linux is missing certain kernel-provided sync primitives, why not just implement them all as regular Linux syscalls for every userland app to use (similar to how e.g. futexes were introduced) rather then crowding them all into some compatibility chardev with an ABI arcane enough that it would only ever be talked to by one library?
The point that would convince the Linux kernel maintainers to actually accept the patchset, though, would likely be the introduction of generally-useful "species of sync primitives." In such a way that those primitives can be used to solve the "hard part" of WINE's virtualization of NT's versions of those primitives; but which wouldn't be constrained to guarantee an efficient zero-impedance call path between those implementations.
[0] https://lpc.events/event/17/contributions/1517/attachments/1... [1] https://lwn.net/Articles/824409/
And that's the part that made OpenVZ unacceptable to integrate into the kernel — the axiomatic framework being introduced by the patch series. No matter how you split and refactor and document and drip-feed in the OpenVZ patches, they'd still just be a set of patches each proposing one component of an already-externally-designed higher-level feature, that can only exist if all the patches are accepted. (See the relevance?)
My understanding of how things should go in Linux kernel development, is that when submitting a single patch to LKML, 1. the motivating user-story should first be explained, and 2. the patch that follows should then be the most idiomatic way to solve that single problem within the existing architecture the Linux kernel. This will then 3. start a discussion, where any "specialization" of the design from its normative "obvious" form may be proposed, and justifications for this argued.
When submitting a patch series, each patch must go through this process independently. If there is any common framework to the patches, it can only be because the "obvious way" (by the consensus of all participants) to solve patch 2's user-story, with patch 1 already in place in a dev branch, is to refactor the logic of patch 1 and patch 2 together.
In other words, it should be up to the core kernel maintainers — not the contributor — to propose an in-kernel architecture like "this set of operations introduced over the last few development patch series, would make a lot of sense to expose as messages passed into a single character device." And such an architecture should be proposed because it's what's best for the kernel, rather than because having it that way would make the contributor's life easier in implementing their library.
(The only exception to this being if the aim is to support direct interoperation with some existing portable software, by exposing a syscall or device or filesystem that other platforms already have, and that existing software already depends upon on these other platforms. Which doesn't apply here — NT itself doesn't have an ntsync device!)
I would expect, if this process were followed for NT's sync primitives, that:
• some of the resulting changes would introduce ioctls;
• some of the resulting changes would add flags for the io_uring or epoll syscall families;
• and some of the resulting changes — for the legacy NT sync-primitive objects specifically — would probably become their own syscalls (probably in such a way that they only exist with a certain kernel module loaded, so that the attack surface they represent can be disabled in hardened systems.)
These changes would likely all solve the problems as stated (i.e. making it possible to execute NT software that uses sync-object X, with only a userland translation library between the kernel and the software, no RPC daemon needed) ... but these changes likely wouldn't solve the problems in the most efficient way for WINE. Not if there's any other uses — in greenfield Linux software, in virtualization of other OSes, in accelerated kernel-mode hypervision, etc — for exposing such primitives.
And also, obviously, these primitives would all end up fairly different from one-another, and from what NT expects — and so WINE would still need a fairly complex translation layer to target these syscalls. Just like it does today. Just, one without inefficiant and apparently-untenable userland arbitration.
Most of what Wine wants added to the kernel should not have a scope of use beyond Wine, which is why the Wine devs have put so much effort into trying to solve these problems in userspace.
WaitForMultipleObjects and Nt(Set|Reset|Pulse)Event are used throughout Windows source code
[1] https://learn.microsoft.com/en-us/windows/win32/api/winbase/...
[2] https://devblogs.microsoft.com/oldnewthing/20050105-00/?p=36...
It’s a shame openVz pioneered container, yet docker ate their lunch
OpenVZ (like LXC) was more of an OS level VM alternative type of container. Docker was more of a developer oriented distribution/isolation system than OpenVZ/LXC/etc which were more of a sysadmin oriented hosting system. In the very early days Docker didn't make much sense to me until I realised that.
I'm very suspicious vale/proton/wine slept on 2-3x performance boosts. It sounds like this problem is already solved, but this implementation might be more elegant, but its the same performance out of edge cases.
Edit: ah, it's about wine upstream, not kernel upstream. they won't take esync/fsync because those have some missing features.
https://github.com/torvalds/linux/commit/a0eb2da92b715d0c97b...
https://www.phoronix.com/news/FUTEX2-futex_waitv-v2
...
https://www.phoronix.com/search/FUTEX2
...
https://www.phoronix.com/news/Futex2-System-Call-RFC
https://www.phoronix.com/news/FUTEX2-LPC-2020
https://www.phoronix.com/news/FUTEX2-2021-Still-WIP
https://www.phoronix.com/news/FUTEX2-Bits-In-Locking-Core
Though none of the articles so far have bothered looking up... 'NtPulseEvent() or the "wait-for-all" mode of NtWaitForMultipleObjects()'Not many great hits for NtPulseEvent() but the first result I got says "Function sets event to signaled state, releases all (or one - dependly of EVENT_TYPE) waiting threads, and resets event to non-signaled state. If they're no waiting threads, NtPulseEvent just clear event state."
NtWaitForMultipleObjects() https://learn.microsoft.com/en-us/windows/win32/api/synchapi...
After reading this, I'm still unclear about what's missing. A guess is the wakeup on timeout?
The funny part is that those are almost never used anyway, so it's more of a formality on the Wine part to be always correct. But at least it's moving somewhere now possibly.
I suppose a driver makes sense, but I'm curious how many recent programs actually use things like PulseEvent given the warnings Microsoft has put on it. People seem to have moved to the pthread-like APIs in recent code I've seen.
Any pointers?
That's actually one of the painful parts about video on Linux, is that if you've got something that needs GPU acceleration via libav, you've got three (or four) not so nice choices: nvidia, which has good hardware support but is hit or miss if it actually works outside of their proprietary nvenc, especially for libav and vaapi; AMD, which seems to lack hardware for this on low power cards, or it's poor where it exists; Intel, where the ARC cards don't work well (or at all) if you lack resizable BAR (which means you can't use anything with any age as the host). Intel's integrated graphics is probably the best choice, but still I'd really like a discrete GPU with good Linux h264, vp9 and av1 encoding and decoding support, as it doesn't seem to exist today.
The resizable BAR problem is annoying, but that's not limited to AMD, either. I'm not 100% sure about Intel, but if resizable BAR doesn't affect performance, that only makes me think they left some easy performance win on the table.
I have to agree on Intel QuickSync, though, that's hands down the best video encode/decode solution for consumers. AV1 encoders are very recent additions to the GPU landscape, though, so you won't be able to get those on a budget. AV1 decoders are easier to obtain, and H.264/VP9 especially so.
I get graphical glitches in gnome, it crashes GPU accelerated browser tabs, my logs are full of “FAULT_PTE ACCESS_TYPE_READ” errors, it freezes the displays for 30 seconds seemingly randomly. And that’s with the latest drivers available for Linux and the latest available 6.5 kernel.
I’m mostly on Windows these days tho bc it’s a better fit for my daily stuff
But running many other ML technologies or Whisper variants is impossible because of lack of Cuda.
So if you are any serious about ML, overpriced and undermemoried Nvidia is the only way to go at the moment. But if you don't care about that and only care about gaming, Amd will be the best choice for most people.
https://www.gamingonlinux.com/users/statistics/#GPUModel-top
As a rule of a thumb, use AMD.
I found it mostly worked, and then broke controller input about 3-4 months ago, and then partially fixed it last month.
What controller or stick/throttle are you using?
Not sure if you're running flatpak Steam but I hear controller recognition works better with standard Steam.
Other than the control contretemps, it's been a good experience absent the occasional Proton breakage. I stick to Experimental, FWIW.
Mac is a better experience for every single thing other than gaming (that may change soon too). My steam deck is an amazing mobile gaming device. I’m hoping one day I can actually just steamos on my gaming pc and dump my last vestigial windows install.
The only big problem left is kernel-aware anticheat. Ideally, we would just stop allowing game companies to hijack the windows kernel in the first place...
This is like when a guy said to RMS, "I have to use proprietary video editing software to do my job" and RMS was like "Then find another job."
IF LINUX CANNOT RUN THE SOFTWARE PEOPLE WANT TO RUN, IT IS A POOR OPERATING SYSTEM.
We've been telling open source folks this since the 90s. It's time for them to listen and learn.