Proposed futex2 allows Linux to mimic the NT kernel for better Wine performance
lkml.org
lkml.org
It's not clear that this is a win overall. But hey, if Windows games run better under wine...
[1] https://lwn.net/ml/linux-kernel/87o8thg031.fsf@nanos.tec.lin...
800 threads on similar WaitForMultipleObject is already a horrific design if you even remotely care about performance. The only viable case for 800 threads (and wine) would be blocking IO which is an extremely boring one.
For instance, in a previous job, a C++ codebase existed where the architect had horrendously misunderstood the actor model and had instated a "all classes must extend company::Thread" rule in the coding standard. So close, yet so incredibly far away.
At the end of the day, wine (like windows) still has to run these programs and do it's best with terrible choices that were made in the past.
I do believe there are tons of incredible horrid code (esp. when it comes to concurrency) - but most games tend to use already established engines, less likely to have glaring concurrency design blunders.
In case it's not clear to anyone reading along: if you find yourself doing anything like this, take a long, deep look in the mirror; you're almost certainly making the wrong choice. Best I can say is that the architect learned from some of their choices in the years since from the resulting consequences.
However, that code is still out there. The engineers have tried to rewrite since then, but it turns out that rewriting isn't adding new features. I don't agree with the company's calculus on the matter, but that's a tale as old as time.
And you'd be surprised how many games aren't written well. Writing a game engine is a bit of a rite of passage among young coders, and a surprising subset of those actually ship. In my experience, a lack of maturity in the field and (we'll be courteous and say) 'questionable' threading models (and as you point out inheritance models too) go hand in hand.
The point at the end of the day is that Wine users (like Windows users) expect the runtime to do the best it can even against terribly written applications. The users want to run that code, it isn't going to be fixed, so anything Wine can do better to paper over the (incredibly obvious to someone skilled in the art) deficiencies of that code is in their users' benefit.
In this particular example you’re probably right though, because performance does matter to game dev, and that would be a horrific architecture perf wise
Yup (64 though), it's an extra annoying limitation. It has been there since the 90s, WinNT (cant recall win95 version as I never used it with win95). It's limited by MAXIMUM_WAIT_OBJECTS which is 64.
For instance Java implements non-blocking IO (java.nio.Selector) by using multiple threads and sync between them to effectively achieve linux' epoll.
----
my point was mostly that the 'huge' 4% happens with massive amounts of threads and WaitForMultipleObjects won't see even that... so kind of click-baity. Flip note: sync. between threads can be done better than relying on OS primitives aside for real wait/notify stuff (which indeed it's implemented via futex)
But if they had meaningfully increased that, then they would run into the same scaling problems as poll(2) - since the return code gives you an index, you wouldn't have to loop through the handle set in user mode like pollfd requires, but every syscall would need an O(n) scan of the HANDLE[] array on the kernel side. You need something like iocp/epoll/kqueue where kernel mode has a list of interesting handles, that doesn't need to be scanned and copied at every syscall.
Red Hat says: <<FD_SETSIZE is hardcoded to 1024 in /usr/include/bits/typesizes.h and can not be changed.>>
Source: https://access.redhat.com/solutions/488623
If I misunderstand, please let me know.
If you can accumulate a series of 4% improvements, then even the computer consumer will start to see the significance.
A 4% improvement on the 4:01 would be a 3:51. A 3:51 wasn't achieved until 1966, about 20 years later.
So 4% can be quite a jump!
A 4% improvement on Bolt's 9.58s would be about 9.19s - probably not something we'll see for several decades, if ever.
Such dominance and grace.
He jogged over the line with that record too.
So no, this won’t give you 4% in general.
I would think that the linked patch is good evidence that work is being spent on this.
Yes, a 4% improvement is huge.
So if you get 4% for free here, that could mean hundreds of dollars savings compared to having to buy the next tier hardware. Multiplied by the number of Intel customers.
4% could mean the difference between "I can still use my old computer" and "I need to buy an upgrade".
(consumer chips are usually basically identical, just with functionality disabled)
It's likely the wine WMO implementation is getting a much more significant increase in performance than 4%, but the patch description doesn't provide numbers.
1. If someone is willing to take the time to develop and submit something to the kernel, it probably has some redeeming factors to it. The original idea or code might not get merged, but something similar to meet those needs likely can be if it's really something the kernel can't do.
2. The modularity of the Linux Kernel build system means that adding new features doesn't have to be that big of a deal, since anybody who doesn't want them can just choose not to compile them in. It's not always that simple, but for a lot of the added features it is.
A lot of the history and discussion can be found here: https://lkml.org/lkml/2021/4/27/1208
I personally feel like the kernel maintainers have been very conservative over the years regarding the introduction of new system calls, but then again I've maybe gotten used to the trauma of being familiar with windows internals.
Consider the slogan "composition not configuration", Linux is squarely the latter. All the hardware, all the different use-cases, that are just added to what's mostly a monolith. You can disable them, but not really pull things a part.
Now, Linux also has some of the most stringent code review in the industry, and they are well aware of issues of maintainability, but fundamentally C is not a sufficiently expressive language for them to architect the code base in a compositional manner.
What really worries me is that as the totality of the userland users an increasingly large interface to the kernel, it will very hard to ever replace linux withsomething else when we do have something more compositional. I am therefore interested, and somewhat working on, stuff that hopefully would end up with a multi-kernel NixOS (Think Debian GNU/kFreeBSD) to try to get some kernel competition going again, and incentivize some userland portability.
Note that I don't think portability <===> just use old system calls. I am no Unix purist. I would really like a kernel that does add new interfaces but also dispenses with the old cruft; think Rust vs C++.
Linux is a Monolithic Kernel, sounds like you are more interested in Microkernels. https://www.wikiwand.com/en/Microkernel
One microkernel alternative could be Redox-OS: https://www.wikiwand.com/en/Redox_(operating_system)
It seems pretty hard to write a microkernel without using composition, though, since the runtime requirements kind of dictate your code structure.
Plan 9 would like to have a word with you.
Some of the changes aren't compatible with the kernel direction as a whole and so they're a separate branch. Given enough time (and if the change is major AND ends up being widely used) it may get merged in.
Those who want a specific kernel with a defined feature set often will look to one of the BSDs as a base (as often they also have reduced hardware requirements, think routers).
Do you have an example of something that you consider part of the "dumping ground?"
As I understand it, this is exactly what Linus does. He's the final arbiter of what makes it in, and has a history of setting high level direction for how stuff should be (usually prefixed with "I'm rejecting this because").
Are there some specific evolutions that you think Linux adopted, and eg. FreeBSD didn't, that were/are crucial to Linux' now overwhelming dominance?
Both have several escape-hatches for arbitrary driver-userspace communication, which makes the number of syscalls pretty irrelevant. (E.g. Linux has ioctl, netlink, procfs, sysfs, ...)
Yes, we need a better metric for accounting for those things, which will paint a much more sobering picture of the complexity and growth of complexity of the userland <-> kernel interface.
NUMA awareness especially stood out as highly significant.
The new supported sizes seems nice and shiny - if well utilised, it could improve performance and capacity in general, capabilities in the low-end and scaling in the high-end.
I don't think anyone in their right mind would design the IndexedDB APIs the way they are with the background of modern JS. It would certainly be promise based and I think there would be a better distinction between an actual error with IndexedDB vs just an aborted transaction. It's also extremely confusing about how long you can have transactions ongoing. I suspect they would replace the nonsensical autocommit behavior of transactions currently with some sort of timeout mechanism. There is way too much asynchronous code out there these days to rely on the event loop for transactions committing.
Linux is honestly pretty clean as far as outdated abstractions go, but Windows is definitely very guilty of keeping around old outdated abstractions. Having execution policies for PowerShell while batch scripts get off scotch free is really strange. These restrictions would only make sense if Windows didn't let users run VBScript, Batch, etc. Can anyone tell me how running a PowerShell script is inherently more dangerous than VBScript or Batch scripts?
Is that why *BSD has lower market share for servers and other networking equipment than it did 15 years ago?
Even the larger world outside the field is somewhat aware of the the looming complexity problem: https://www.theatlantic.com/technology/archive/2017/09/savin...
Basically on our current trajectory we will end up with something like the biosphere that can do stuff but too complex to fully understand --- but also probably well more fragile than that?
So the codebase is actually not growing as much as it could.
It's not a microkernel but it does a pretty good job of being flexible for a monolithic kernel.
If you're curious, clone the kernel source and run `make menuconfig` and you can play with it yourself
E.g.: syscall doesn't accept multiple handles and other bells and whistles? Well duh, it's not supposed to—just make a new one that does.
Of course, it's probably too late to change this design choice now...
Another thing to remember is that in C, it's easier to make a call more specific, by hardcoding values than it is to make it more general, by picking between variants in a branch. I can see programmers wishing the API was more dynamic while writing a giant if-else tree to select the appropriate version of the syscall, but the opposite problem, when the same invocation is always needed but the API takes a lot of dynamic options, is not as bad because it is easy to write a bunch of literals in the line that calls the function. (If C had named arguments that approach would be almost flawless.)
I'm talking above about the first half of your previous comment, where you say that having operation names in the function names would somehow require multiple calls and break atomicity. I can only understand this if Linux syscalls are parameterized to do one thing and then another thing and something else, in one call—which sounds horrible, but still can be expressed in function names.
At the end of the day a function call is an interrupt with the syscall parameters in a bunch of registers. As you do not want to use a different interrupt (as there are only so many interrupts [1]) for each function call, the syscall id is passed as a parameter in a register. On the kernel side the id is used to index in a function pointer table to find and call the actual system function.
Turns out this table is actually limited in size so some related syscalls use another level of indirection with additional private tables (I think this is no longer a limitation and these syscall multiplexers are frowned upon for newer syscalls).
All of this to say that FUTEX_WAIT and FUTEX_WAKE are just implementation details. A proper syscall wrapper would just expose distinct functions for each operation.
But for a very long time glibc did not expose any wrapper for futexes as they were considered linux specific implementation details and one was supoosed to use the portable posix functions. So the futex man pages documented the raw kernel interface. That didn't stop anybody from using futexes of course and I think glibc now has proper wrappers.
[1] yes, these days you would use sysenter or similar, but same difference.
P.S. Come to think of it, I sorta knew about the interrupt business, and vaguely about the later but fundamentally similar method (but not that there are many interrupts for syscalls, though thinking about eight-bit registers clears this up quickly). Guess I just forgot that libc is kept separate from the kernel, in the Linux world.
I was addressing operation names in the function names in the second half of my comment, in the first half I was addressing breaking up specification into a state machine process like OpenGL does. State machines are another way of breaking up big functions with lots of options. They have pseudo default argument support as an advantage, and ease of implementation via static globals as a historical (before multithreading) advantage. State machines are kind of like builder pattern as a way to get named arguments and default arguments in a language that doesn't have either, except instead of the builder being a class it's a singleton (with all the attendant disadvantages.)
It's actually hard to understand for me that an API that was created for efficient computing doesn't specify/guarantee the relative speed/latency of operations.
Of course, I'm guessing C programmers would cringe at the prospect of going full COW for such use-cases.
With the difference in my proposal that the object is treated as a COW.
- Syscalls are not designed to be called directly from end-user programs. Instead lower level libs calls syscalls.
- Getting a syscall accepted is a long and tedious proccess, so they should as general and future proof as possible.Afaik, if locking a mutex succeeds, futexes avoid a context switch to and from the kernel. The fast path happens completely in userspace. Only when a mutex lock causes a thread to sleep, the kernel is needed. All of this improves mutex performance.
There are other locking primitives than mutexes, and the current futex interface is not optimal for implementing common windows ones. Hence this proposal.
This is a common misconception, a "context" switch is a lot more expensive than a "mode" switch that goes from user mode to the kernel mode (and back).
Of course, if the futex has to wait, it'd be a context switch for real.
>"Before the introduction of futexes, system calls were required for locking and unlocking shared resources (for example semop)."
>"A user-space program employs the futex() system call only when it is likely that the program has to block for a longer time until the condition becomes true"
The first indicates futexes are an improvement because they avoid the overhead of system call but the second states that futex() is a system call.
Is this similar to how a vDSO or vsyscall works where a page from kernel address space is mapped into all users processes? In other words is it similar to how gettimeofday works?[1]
[1] https://0xax.gitbooks.io/linux-insides/content/SysCall/linux...
The userspace fast path can in principle be implemented on top of any kernel only synchronization primitive (a pipe for example).
The nice thing about futexes is that they do not consume any kernel resource when uncontended, i.e there is no open or close futex syscalls. Any kernel state is created lazily when needed (i.e on the first wait) and destroyed automatically when not needed (i.e. when the last waiter is woken up). Coupled with the minimal userspace space requirements (just a 32 bit word) it means that futexes can be embedded very cheaply in a lot of places.
>The nice thing about futexes is that they do not consume any kernel resource when uncontended, i.e there is no open or close futex syscalls.
This is an extremely interesting and useful observation, thank you for making it.
Futexes are missing a lot of useful functionality - see the OP, and as another example it's very hard to integrate them with an event loop - so I've been dreading building anything with them, but I thought I needed it to get fast synchronization. But for my use cases, I already have setup and teardown calls. Maybe I can do fast userspace synchronization on top of some other kernel object - a pipe, or something, as you suggest.
Do you have any more to share about this observation, or pointers to any implementations of fast userspace synchronization on top of something other than futexes?
But there isn't really ever a reason to mix futexes and event loops. Futexes are not a synchronization primitive, just a waiting strategy. Decouple the two and you can integrate your synchronization primitives with an event loop while still using futexes for the non event-loop cases.
I'm not sure what you mean, can you elaborate? Suppose I had a mutex (synchronization) implemented with a futex, and a shared memory queue (waiting) implemented with a futex; suppose both are being used to coordinate between multiple processes. I guess you're suggesting that one of those two doesn't need to be integrated with an event loop? But why? Both of those seem useful to have as part of an event loop.
You have two options: you can use an eventfd to implement the queue full/empty event, or provide a generic notification interface (i.e. a callback). Any non event loop users wiuld simply have the callback signal a futex, but it allows for more complex use case: obviously you can wakeup the event loop from the callback, but you could resume a coroutine, send an async signal or whatever makes sense for the application.
Thank you for pointing this out. That makes sense. This is sort of like the idea that the fastest system call is the one that is never made. Can you say how does something in userland know if something "would" block without actually crossing the user/kernel boundary? Does the kernel expose a queue length counter or similar via a read-only page?
But in the contended case the flag is already set by somebody else, so you leave a sign (maybe you set it to 42) about the contention in the memory address, and you tell the kernel that you're going to sleep and it should wake you back up once whoever is blocking you releases. When whoever was in there is done, they try to put things back to zero, but they discover you've left that sign that you were waiting, so they tell the kernel they're done with the address and it's time to release you.
None of the conventional semantics are enforced by the operating system. If you write buggy code which just scribbles nonsense on the magic flag location, your program doesn't work properly. Too bad. The kernel does contain some code that helps do some tricks lots of people want to do using futexes, and so for that reason you should stick to the conventions, but nobody will force you to.
>"When whoever was in there is done, they try to put things back to zero, but they discover you've left that sign that you were waiting, so they tell the kernel they're done with the address and it's time to release you.
Did you mean to say "and it's time to release to you" here?
>"None of the conventional semantics are enforced by the operating system"
Ah OK, this is the part that I feel is maybe always left out of the things I've read on futex(). I guess this is just always implied then that some library implements these semantics correctly? And that library is generally going to glibc?
Yup. For example, the pthread_foo functions are futexes under the hood.
The POSIX threading API accommodates many possible implementations, and so in Linux many of the bits that look expensive (and might be on other systems) are more or less free thanks to futexes, e.g. setting up the mutex in Linux is just allocating an aligned 32-bit value on the heap and setting it to some initialising value, but on some systems it involves an OS system call to get a mutex handle.
This is the true benefit of the futex. The typical scheme for using it looks almost exactly like a well known trick for speeding up conventional mutual exclusion features that have low contention, which you'd see in say BeOS as the Benaphore, or described in several books about high performance programming. But those tricks all imply the expense of first obtaining a handle from the operating system for every such mutex whereas with a futex that part is free and you only talk to the operating system at all once there is contention for it to help you manage.
Do you have any links that document this trick in Linux? Is there name for it or search term I could use to find out more?
https://www.haiku-os.org/legacy-docs/benewsletter/Issue1-26....
Note that when I say "typical scheme for using it" I did not mean, that your program should implement such a scheme itself, instead futex is intended for the sort of people who implement the concurrency features in your language or APIs to use - and this is how they'd go about doing that to deliver the features you just use. So, for most of us this is interesting to know about but unlikely to impact our day-to-day practice.
Fuss, Futexes, and Furwoks: Fast Userlevel Locking in Linux
https://www.kernel.org/doc/ols/2002/ols2002-pages-479-495.pd...
Another great paper that came after some additional experience by glibc developers:
Futexes are Tricky, by Ulrich Drepper
I realize that one is userspace and the other kernel but would the above be the right mental model?
A mutex is a lock -- if multiple users attempt to access a resource simultaneously, they each attempt to acquire the lock serially, and those not first are excluded (blocked) until the current holder releases the lock.
Mutexes are typically implemented with semaphores:
First a more developer-oriented explanation (because I imagine that's what you really want):
A semaphore locks a shared resource by implementing two operations, "up" and "down". First, the semaphore is initialized to some positive integer, which represents how many concurrent "users" (typically threads) can use that resource. Then, each time a user wants to take advantage of that resource, it "down"s the semaphore , which atomically subtracts one from the current value -- but the key is that the semaphore's value can _never_ be negative, so if the value is currently zero, "down"ing a semaphore implies blocking until it's non-zero again. "Up"ing is just what you'd imagine: atomically increment the semaphore value, and unblock someone waiting, if necessary.
Semaphores are generally seen as a fundamental primitive (one reason being that locks/mutexes can be implemented trivially as a semaphore initialized to one), but they also have broad use as a general synchronization mechanism between threads.
For a true ELI5 of semaphores (I enjoy writing these):
Imagine everyone in class needs to write something, but there are only three pencils available, so you devise a system (a system of mutual exclusion, per se) for everyone to follow while the pencils are used: First, you note how many pencils are available -- three. Then, each time someone asks you for one, (which we call "down"ing the semaphore) you subtract one from the number in your head and give them a pencil. But if that number is zero, you don't have any pencils left, so you ask the person to wait. When someone returns and is done with their pencil, you hand it off to someone waiting, or, if nobody is waiting, you add one to that number in your head and take the pencil back ("up"ing the semaphore). It's important that you decide to only handle one request at a time (atomicity) -- if you tried to do too many things at once, you might lose track of your pencil count and accidentally promise more pencils than you have.
Isn't that a semaphore rather than a mutex?
For ELI5 on semaphores, it's like going at the train station, airport, or Mc Donald's: There is a single queue for multiple counters. Everyone is given a paper with a number, and once a counter is free, the one with the lowest number can walk to it. A semaphore is just the number of free counters, it decreases when someone goes to one, and increases when one leaves. If the semaphore is zero, there is no free spot, and people will have to queue.
It's quite close to another synchronization primitive, barriers: if you have to wait for n tasks to finish, you make one with n spots, and release everyone when all the spots are filled.
To me, mutexes are semaphores, but with only one counter: there is only exactly one or zero resources available. If you take the mutex, others will have to wait for you to be done before they can take it. They can queue in any fashion, depending on the implementation: first come, first served; fastest first; most important first; random; etc.
Please excuse my ignorance...
This sounds easy, but is in practice wrought eith peril. You need high performance, not waste too much cpu time on sleepers, some notion of fairness when you have lots of waiters, minimal memory usage, etc...
"Detailed approach of the futex"
This can't only be a win for emulating Windows, surely (future) Linux native code could benefit from this too?
It would be nice if io_uring polling had a futex-like feature that spinned in userspace rather than having to syscall io_uring_wait every time. I feel like that feature would be more universal than futex2.
Adding an atomic int somewhere in the userspace would add ~100ns in the kernel path but save many microseconds in the userspace-spin path during high contention so it seems like a win to me.
Of course a low latency option should also exist for those cases that require it and currently that pretty much means spinning.
Can we take a moment to appreciate how mind-blowing astonishing a 99th percentile latency in the us range is given a general-purpose, multi-user OS with support for virtual memory is?
That would certainly help porting Windows software to Linux.
They decided to rename it since the kernel maintainers don't want NT interfaces and having NT in the name doesn't help
https://docs.darlinghq.org/internals/basics/system-call-emul...
I have been using bpf in seccomp for non-root for years.
Maybe your distro patched that out. The stock kernel allows bpf for non-root.
They would probably remove those too if the balance of Linux didn't put "don't break user space" higher than "defense in depth security".
I don't really know what the use cases for unprivileged BPF are though.
it's actually pretty nice how generic they made it.
> (Note that this was formerly called the Win32 API. The name Windows API more accurately reflects its roots in 16-bit Windows and its support on 64-bit Windows.)
Though Win32 is going to stick around for a while as can be seen in the URL for the page that's quoted from: https://docs.microsoft.com/en-us/windows/win32/apiindex/wind...
regarding win32 vs. the nt api. i'm not super familiar with the pedantic distinctions there, but i do know, much to the programmer's chagrin, the thousands of windows api functions have varying levels of kernel entry, where i always have appreciated the (relatively) small number of unix style syscalls and the clear delineation between kernel and userspace api. (although i suppose futexes blur that, in name at least!)
https://lore.kernel.org/lkml/20210427231248.220501-1-andreal...
the reason why linux isn't gaining shares on steam anymore is because valve openly said linux has no ecosystem and is forced to emulate windows -- proton
Before Wine actually got good, the few "native" ports we had of major games on Linux were poorly optimised and often worse than just running the Windows versions in Wine.
Sometimes you gotta take what you can get. Valve's approach to invest in Wine means I don't need to dual boot any more. Native games are the cherry on top.
>the reason why linux isn't gaining shares on steam anymore
Did anyone ever think Linux had any kind of comparable gaming ecosystem? lol
Valve invests in Linux as a safety net in case Microsoft ever flicks a switch and forces all game distribution through the Windows Store or similar.
Valve decided to seek short term gains instead of thinking of a long term strategy by working with linux people to develop a nice ecosystem with libraries / drivers / app distribution
They could collaborate with Epic/Unity to facilitate linux builds
Valve went selfish and the result is here, linux share isn't growing, despite massive efforts from microsoft to kill themselves with vista, windows 8, and window 10 massive spyware drama and now the updates that break gaming
What if windows 11 is again a free upgrade and completely breaks wine? by introducing massive architecture changes
That'll be the death of wine and linux gaming, because no thought went into securing a nice and flourish ecosystem
The part that you suggest they should do "instead" is exactly what they have been doing. As far as I know, every game by Valve is native on Linux, along with their source engine.
> They could collaborate with Epic/Unity to facilitate linux builds
Both Unreal Engine and Unity can export games for Linux. Both seem to be able to be run on Linux for development as well. What more is needed?
> What if windows 11 is again a free upgrade and completely breaks wine?
There is literally nothing Microsoft could do to break wine; it isn't their software. If Microsoft breaks backwards compatibility, no existing software for Windows would run on the new version. Linux + WINE would be the best option for running the majority of Windows software with the most recent updates (one could of course still use older versions of Windows instead, but that's a bad idea long term).
By looking at steam stats it doesn't seems to work, but i'll trust your judgement since it seems to.. not work at all.. but that might be the goal, to waste time, effort, money, and they gotta justify the 30% tax somehow, hey we are working on things nobody care about!
* Most nontrivial Windows program would get very confused if exposed to a trivially mapped Unix filesystem. It's not so bad if every program would use environment variables to locate stuff, but you can't discount the possibility that some of them hardcode paths. Also, application data in Windows is usually kept in a common folder, while in Unix filesystems it is usually split across `/usr/bin`, `/usr/share`, `/usr/lib` and so on.
* Many Windows applications store their configuration in the registry. That part is not too bad though; it should be completely doable in userspace.
* The application might issue system calls that have different semantics, or that don't even exist on Linux. TA is significant because it paves the way to an efficient implementation of one of these system calls.
The application data in windows would not be the same as `/usr/` paths. It would be the equivalent of `~/.config` and `~/.local`. The last thing anyone in Linux would want is for Windows to start populating their system folders, let alone probably creating hundreds of top-level folders in their `~/.config` or `~/.local` directory.
> Many Windows applications store their configuration in the registry. That part is not too bad though; it should be completely doable in userspace.
The Windows registry is an atrocious monster that should have died already. You can rarely ever truly delete an app in Windows, as there will be sometimes hundreds of registry keys leftover.
> TA is significant because it paves the way to an efficient implementation of one of these system calls.
If there is mutual benefit outside of Windows, then sure why not. If it is built just for better Wine compatibility with Windows then I hope to see a rant by Linus telling them how ridiculous they are.
I don't like the registry either. My `not too bad` refers to the effort required to support it. Wine should make it possible though give each application its own copy of the registry. This copy could be nuked along with the application when it is uninstalled.
Linus doesn't like a lot of things. ZFS for example. It certainly sucks to maintain out-of-kernel patches, but if there is enough benefit, someone will do it. And it would be actually pretty sweet and ironic if Windows applications run faster on the Linux desktop than on Windows. Also, I am optimistic that futex2 will find many applications outside of Wine.
Taking "emulate" at its dictionary-definition, which is to imitate or copy, it's easy to argue that Wine is an emulator, because it copies and imitates the Windows ABI and API (the earliest versions of Wine even called itself a "WINdows Emulator"). They just happen to draw a hard line against emulating x86.
define people by what they do
that works with software too