https://github.com/spurious/SDL-mirror/commit/59728f9802c786...
https://github.com/spurious/SDL-mirror/commit/59728f9802c786...
So it would be the game developer that would have to update the version of the SDL library. The binary patching done seems like a good-enough alternative in the meantime.
I have the feeling such bundling of dependencies is fairly common when porting games for Linux.
[1] https://old.reddit.com/r/linux_gaming/comments/1upn39/sdl2_a...
Edit: (what I believe to be) This freezing bug was only added 3 years ago, so it might actually have it.
What should not be on your main thread is any long blocking compute, which is why rendering/game logic often goes to another thread - although simple games could easily be single-threaded.
> What should not be on your main thread is any long blocking compute
Isn't that contradicting yourself? I'm pretty sure open() can block.
No, blocking compute would be you doing something for a long period of time.
"open/close can block" means very little. You only need a mitigation if it ends up blocking long enough to be a problem in reasonable setups.
You do that have to care about what happens if someone runs the code on an intentionally terrible/horribly slow but technically in spec toy filesystem. No need to prematurely optimize for this scenario.
And especially with devfs you cannot just listen to POSIX and must know what the kernel is providing you - knowing how long operations take on such fds is normal design input.
That criteria is established here - the OP is about an issue affecting paying end-users! Premature optimisation is not relevant - the software is failing.
It is unsafe to make fair-weather assumptions about customer systems.
Consider a common software failure: where the user is saving data to a SMB or NFS partition, and then there is a loss of connectivity to that file-server, and the developer has done whatever I/O they need on the main thread. This causes data loss.
You /should/ be able to assume (1) reliably fast return from core async-coordination syscalls (e.g. select, poll), and (2) that you will not suffer process or thread starvation caused by someone else. Respecting those constraints, it is good practice to isolate sync calls to a non-main thread in order to catch when they are not returning. This robustly covers both common scenarios (like the missing-filesystem) and obscure scenarios like the one in blog post.
The software is not failing, it is experiencing performance degradation: a 500ms pause whenever excessive and entirely unnecessary work is done.
The solution is to not do the work. Moving open/close to a different thread is premature optimization, as no necessary call has been profiled to cause issues on any known system.
Performance 101, do not do things that you do not need done. Even if you want live hotplug and input reconfiguration during gameplay without touching any menus, you only open a device when it appears.
> It is unsafe to make fair-weather assumptions about customer systems.
It is more pointless to optimize for worst-case scenarios - experiencing performance degradation on a faulty system is fine.
All applications have a minimum performance requirement to remain responsive, which is equivalent to always making a certain degree of "fair-weather assumptions".
> Consider a common software failure: where the user is saving data to a SMB or NFS partition, and then there is a loss of connectivity to that file-server, and the developer has done whatever I/O they need on the main thread. This causes data loss.
This is non-sequitur - doing something on the main thread does not cause data-loss. Losing connectivity causes data-loss.
Heck, as main-thread I/O with an event loop implies non-blocking fds, you would not even be blocked by this unless you call fsync(2) to explicitly block until flush is complete, which a normal application does not need to do. The other (horrible) side-effects of network filesystems will cause problems for your application no matter how you interact with the fd.
Furthermore, you cannot use device files without reasoning about their exact implementation. They are not basic files.
> Respecting those constraints, it is good practice to isolate sync calls to a non-main thread in order to catch when they are not returning.
That's a hack, and is not even a solution. What are you going to do when they don't return? Accumulate dead threads and inconsistent shared application state?
I don't think the open and close calls respect non-blocking on linux when operating on files (specifically files, not sockets - for sockets the return is always quick as far as I know). From man open(2), "I/O operations will (briefly) block when device activity is required, regardless of whether O_NONBLOCK is set". And my recollection is that they will non-briefly block if the problem is a hanging NFS mount.
"What are you going to do when they don't return? Accumulate dead threads and inconsistent shared application state?"
Good point. My practice is to use child processes. These can be killed, and so I do not run into this. But subprocesses ratchets up the amount of work to be done, because you then need to do async IPC. So it's now a lot of extra work. It's even worse for multiplatform stuff because now you are exposed to platform differences (e.g. select is suboptimal on Linux, but poll is not available on Windows).
As you say, using threads within same proc would lead to stale threads. In some contexts this would be tolerable but it is not nearly as simple+clean as I presented.
Thinking hard about addressing this on Linux brings me down, every time. I hope io_uring will make pure async practical within a single process. Even if it does, the multiplatform story will remain complex. I am not fond of your disregard for user data in the (tangential) discussion about disappearing filesystems, but you have won me over to embracing the main thread in this context.
Are there more distros that allow you to do this?
Gentoo also has the concept of "Slots", so you could have multiple versions of the same libary installed and packages will choose their version to build against accordingly.
https://github.com/systemd/systemd/blob/main/src/libudev/lib...