Most UI applications are broken real-time applications
thelig.ht
thelig.ht
If you're controlling a bandsaw, you've got a hard real-time application. You can't miss your window for the next instruction.
Most user interfaces are soft real-time. Occasionally missing your window is fine. And so it is OK to do things like hash lookups whose average performance is O(1), and whose worst case performance is O(n). Ditto for dynamically resizing arrays. As long as you're back in time, most of the time, it can be OK.
The problem isn't that we code UIs as soft real-time. It is that the soft bit gets squishier with each layer of abstraction. And there is a feedback loop where people get more and more used to slow applications, so nobody is concerned if their application is unnecessarily slow as well.
Compounding this is the fact that we often measure performance in terms of throughput, instead of latency. Therefore, every single device and peripheral is willing to lose a bit of latency. It happens at all levels. The classic I like to quote is http://www.stuartcheshire.org/rants/latency.html.
And the result is that modern applications on modern hardware are less responsive than older applications on older hardware. Fixing it is a question of fixing a lot of little problems. And, as we used to joke about Microsoft, what's the point of developing fault-tolerant software when we've already developed fault-tolerant users?
So I've pressed a button to load a file that's usually really fast and this time nothing happens because the async call is taking it's time. Is that being represented to the user in some meaningful way? Is the button disabled until the operation completes? Is it stuck in the down position to show I can't press it again, or does it spring back up and show a loading dialog? Is that any more clear in a real time system then it would be in a blocking situation?
The problem is not blocking vs realtime, it's about programmers not understanding all the state their code can run in, and it's not clear that a realtime system would save the user when they fall into that unknown state.
Personally I set up i3 to open most windows asynchronously, so my flow isn't interrupted. It's great, but takes a bit getting used to windows not randomly stealing focus. It's not for everyone though.
You cannot allow editing - the existing buffer will be replaced once load completes. You can show the menu, but most options should be disabled. And it will be pretty confusing for user to see existing file remain in read-only mode after "open" command.
The most common solution if you expect loads to be slow is a modal status box which blocks entire UI, but maybe shows progress + cancel button. This definitely helps, but also a lot of extra code, which may not be warranted if usual loads are very fast.
If you pick another single-window app, my response is: that is a decision they chose to make. They can also choose to go multi-window or tabbed just like notepad did.
I see little reason why building a UI today cannot be a superset of what was solved in the 90s so I’m curious to know what that solved subset looks like to you
Some platforms publish a list of recommended guidelines which are effectively a standard. For example here's one from Apple about when and how to use charts in an application: https://developer.apple.com/design/human-interface-guideline...
Also they were called GUI standards or UI at the time.
The modern equivalents called UX isn't reflecting the same conglomeration of standards and conventions though. So not talking about the newer stuff.
I'm no expert on it, and it required specialized expertise. It's been abandoned for mobile interfaces and the modern UX stuff, which often optimizes for design over functionality.
If you've never used old software, it's hard to explain. But old Apple or Microsoft GUI standards would cover the basics, but you'd also need to study the applications and how they presented their GUI.
Nowadays man's applications are web apps, build without such frameworks, with less UI research and even where frameworks are used they are often built with a somewhat mobile first approach.
If I visited a site dedicated to hamburgers today, I would not be surprised if the "Log Out" button was presented as an image of a hot dog. It would be a mystery to me what that hot dog did until I clicked on it.
Compare this to 90's UI, where it would pretty much be unheard of to do something like that. It would have been a joke, or a novelty. These days that sort of ambiguous interface isn't presented as a joke or a novelty - it's the real product.
This was broken with Ribbon and the hamburger menus that every application seems to have switched to for no other reason it seems than to copy Chrome.
To be fair Ribbon is somewhat useable again, but the in the first version I have no idea how people were supposed to find the open and save functions :-)
Other problems:
Tooltips are gone. Yes, I can see they are hard to get right on mobile, but why remove them on desktop while at the same time when the help files and the menus were removed?
The result is even power users like me have to hunt through internet forums to figure out how to use simple features.
Back in the nineties I could also insert a hyphen-if-needed (I have no idea what it is called but the idea is that in languages like Norwegian and German were we create new words by smashing other words together it makes sense to put in invisible hyphens that activates whenever the word processor needs to break the word and disappears when the word is on the start or in the middle of a line and doesn't have to be split.)
Back in the nineties I was a kid on a farm. Today I am a 40+ year-old consultant who knows all these things used to be possible but the old shortcuts are gone and I cannot even figure out if it these features exist anymore as documentation is gone, tooltips are gone and what documentation exist is autotranslated to something so ridiculously bad that I can hardly belive it. (In one recent example I found Microsoft had consistently translated the word for "sharing" (sharing a link) with the word for "stock" (the ones you trade).
Hamburger menus I disagree with but sort of understand the logic of - they're basically like making 'fullscreen mode' the default mode, and then the hamburger menu button just sort of temporarily toggles that off. It makes perfect sense on mobile (I don't think that's what you're talking about though), and on the desktop it can make sense in a web browser when you have, essentially, 3 sets of chrome - you have the desktop window, the browser's chrome, and then the website's chrome all before you get to the website's content.
I don’t care if the food is bad because it is too salty, or because it is overcooked, I still won’t eat it.
They later gave this up and almost everything else in their very reasonable guidelines based on actual research when they switched to OS X in a hurry and multithreaded everything. The early Mail application was a disaster, for example. Generally, people started complaining a lot about the spinning beach ball of death.
In contrast, modern UX guidelines are mostly about web design, how to make web pages look fancy. They also recommend instant feedback, but many libraries and application designs don't really support it.
Ya, you know, unless you didn't speak English, needed accessibility, had a nonstandard screen size, had a touch screen, etc
Not my cup of tea, but I got the same problem.
Going to try this: https://faq.i3wm.org/question/2828/open-application-and-fix-...
It's not impossible - tape drives jukeboxes with 5 minute delays aren't a new invention.
* No paging. Entire program in memory at all times. Although it was possible to put a paging library inside an application and let it manage its own paging, which was done for gcc.
* A real time CPU dispatcher. Real time priorities were strictly preemptive. Unblock a higher priority task, it starts now, not when the dispatcher gets around to it.
* The usual test for a hard real time OS is that you have an interrupt routine that senses an external pin. It unblock a user level task when the pin goes high. The user level task turns on an output pin. You hook up a scope and a square wave generator, and watch the latency. If there are outlier values on the scope, something is broken. One implication is that stuff below the OS, such as anything running in system management mode after boot, has to be eliminated. It cannot be allowed to steal cycles.
* While real time work is going on, you can still run compiles and web browsers. They get preempted. I was impressed that that worked. Our real time application had a hardware stall timer, and if the commands coming out were late, a relay tripped and everything shut down.
So, it's absolutely possible to do this. Downsides:
* You're way out of the mainstream.
* You're going through the general case every time. Fast paths are unwanted.
* Everything is always on. Power consumption is constant. No CPU slowdowns, no sleep modes, no battery saving. This is just fine when you're running a rolling mill, and 99+% of the power is driving the heavy equipment. Not so good when you're running a credit card terminal. Battery powered devices cannot run well in this mode.
Is there no possibility for compromise here?
A credit card terminal sits idle 99%+ of the time, does little actual work, and whatever work it actually does, it's in a reaction to an input event (such as user pressing a button, or a command being sent from cash register). CPU power states are, to my understanding, something you can switch between many times in a fraction of a second. Outside of talking to the network, duration of which is by nature unpredictable, everything else falls into two modes - "slow mode" for listening to events (and perhaps doing whatever is needed to keep the network connection alive), and "fast mode" for processing those events. Sounds to me you could put a real-time guarantee on everything except the networking parts, and keep the CPU slow/sleeping for most of the time.
That was the case in ye olden days of single-core CPUs, but is it still true today?
I can easily imagine a computer that has one always-on core for kernel chores and hardware management, and a dozen dormant cores that wake up when userland tasks need those cycles, all running in hard real time. Or even better, a mix of real-time for UI & audio on some cores + non-time-bound jobs on others.
There's no reason that CPUs couldn't be designed to avoid such interference, but don't expect to find that on x86. Not sure about ARM.
In the demoscene days, we used to obsess over 60 Hz guaranteed framerates. Code was constantly being profiled using raster bars. Our effects might have been crappy, but all was buttery smooth.
Then someone decided it would be a good idea to create a common user interface that would work on many different framerates and resolutions, and all was lost. Or people would try for complex 3D graphics and sacrifice smoothness.
Some people tried to convince us that humans have a 150ms built-in visual processing delay, and all would be fine. This is not the case, and we are still stuck with mediocre animations. The content has improved a lot though :)
For those wondering: we'd change the background color from, say, black to blue for the "physics" part, to green for the "drawing" part, to purple for the "audio rendering" part... All in the same frame, while the frame was being drawn.
So you'd see on the border of your screen (usually outside of the drawing area you'd have access to) approximately which percentage of each frame was eaten by "physics", graphics, audio etc. and how close you were to missing a frame.
And it was quite "hectic" typically these color "raster" bars would jump around quite some from one frame to another.
But yeah there was indeed something very special about having a pixel-perfect scrolling for a 2D game running precisely at the refresh rate (50 Hz or 60 Hz back then). And it wasn't just scrolling: games running on fixed hardware often had characters movement set specifically in "pixels per frame". It felt so smooth.
Something has definitely been lost when the shift to 3D games happened: even playing, say, Counter-Strike at 99 fps (was it even doing that? I don't remember) wasn't the same.
People don't understand that their 144 Hz monitor running a 3D game still doesn't convey that smoothness that my vintage arcade cab does when running an old 2D game with pixel-perfect scrolling while never skipping a frame.
I wish modern game developers & graphics pipelines would spend more effort optimizing input latency. I'll stop noticing the graphics 2 minutes into playing the game. But if the input is laggy, I'll feel that the entire time I'm playing. And the feeling of everything being a bit laggy will linger long after I stop playing.
You've already bought the game at this point so most devs don't care.
An insane amount of money and effort that goes into most AAA games. It seems like a stretch to accuse the industry of not caring about the quality of the games they make.
What I meant is that most people are way more sensitive to graphics than anything else. A lot of people will never play a game that is ugly but don't seem to mind laggy inputs.
There are a lot of games that, even without "real" input lag, have tons of animation dampening/inertia to the point that it makes them IMHO unresponsive, just to improve the "cinematic feel" and avoid characters instantly changing direction.
If you prefer responsiveness to having the player character spin be perfectly smooth, well, sucks to be you, because prettier games sell better.
Obviously this doesn't apply to fighting games or competitive shooters or anything like that where responsiveness is the point of the game.
Maybe. I wouldn’t be surprised most people are much more sensitive to input latency than we think. But when there’s input latency, they have no idea that that’s the problem so instead they attribute the slushy feel of the game on something else. Or just don’t enjoy the game as much without knowing why. Game developers talk a lot about “game feel” and this is the kind of thing they’re talking about.
It doubly screws you in the chocobo racing section because the controls assume no delay, meaning you can’t respond to hazards in time and constantly overcorrect movement.
The obvious fix for laggy input is to apply the same processing to local input, but I've not heard of anyone doing this.
(I think this technique is known as netcode.)
Touch events were run through a filter to predict where the touch position would be some number of frames in the future -- presumably to compensate for page-flipping delays. Not sure it made much difference, mostly because smooth animations on old versions of Android were extremely difficult because of inadequate CPUs/GPUs.
The classes used to do predictive tracking don't seem to be used anymore in current best practice/current API sets.
Blur Busters has detailed explanations, e.g.:
https://blurbusters.com/blur-busters-law-amazing-journey-to-...
The actual motion-to-photon delay (measurable with a high-speed camera) in a well optimized competitive FPS is somewhere around 20ish ms or less these days, so it's on par with it in regards to the input lag, and overall it's far smoother because of the higher frame rate, better displays, and way more thought put into tiny nuances competitive players complain about. A proper setup feels extremely smooth, responsive, and immediate.
I.e. more than two frames between input and output. That is quite a lot, and AFAIK a large regression.
Pixel artists would (and still continue to) consider every individual pixel, performing anti-aliasing by hand. Animations would be timed in relative to the (fixed) framerate, leading to optimal smoothness.
Reaching this level of fluidity is only recently possible again, at the cost of vastly higher hardware requirements. Most people haven't even experienced this, because of the suboptimal period we have been in for some time now.
You win some, you lose some.
There are great examples even outside of the gaming world. I wish there were videos around on how fluid were Scala Multimedia presentations on a bare Amiga 500 in the 90s. An 8MHz clocked machine (7.16 MHz here in the EU for PAL synchronization) that in that field would outperform gear costing at least an order of magnitude more. It was hugely successful back in the day in many local TV stations.
Say what you will about JavaScript, but a great thing it's done for modern software is making async programming default and ergonomic. With most APIs, you couldn't block on IO if you tried. Which means web UIs never do
You are mostly right that you couldn't block on IO, necessarily, but you could still hose up the event thread quite heavily.
- Page scrolling/rendering
- If the user triggers an event, like clicking on a button, that JS event handler is independent from the waiting function and can run while it's waiting
- Same thing if some other `await`ed IO resolves; that function can resume while the other one is still waiting
This is why async/await exists. It allows our code to tell the runtime "let me know when this is ready, otherwise do whatever you want in the meantime". Many languages have a similar feature now, but in JavaScript it's especially fundamental to the language and ecosystem, which means code gets written this way by default, which means JS programs don't have the problem mentioned in the OP by default.
I'm assuming it is on the programmer to do some smart debouncing in the event handler?
Done the straightforward way, you will end up displaying whichever one took the longest to return. I've run into this class of bug before, and since then I make sure to design my primitives to guard against it (always assign the latest-requested instead of latest-received result into state)
If it helps you think about it, any await statement can be converted straightforwardly to an old-fashioned promise resolution with a callback:
async function foo1() {
doStuff()
const res = await doIO()
return 'Result: ' + res
}
function foo2() {
doStuff()
return doIO().then(res => {
return 'Result: ' + res
})
}And it's not just OSes that don't particularly care about real-time - modern processors don't either. e.g., in the common scenarios it's possible to load maybe eight 8-byte values per nanosecond (maybe eight 32-byte values if you're doing SIMD), but if they're out of cache, it could take hundreds of nanoseconds per byte. Many branches will be predicted and thus not cost affect latency or throughput at all, but some will be mispredicted & delay execution by a dozen or so nanoseconds. On older processors, a float add could take magnitudes more time if it hit subnormals.
If you managed to figure out the worst-case timings of everything on modern processors, you'd end up at pre-2000s speeds.
And virtual memory isn't even the worst OS-side thing - disk speed is, like, somewhat bounded. But the number of processes running in parallel to your app is not, so if there are a thousand processes that want 100% of all CPU cores, your UI app will necessarily be able to utilize only 0.1% of the CPU. [edit note: the original article has been edited to have a "Real-time Scheduling" section, but didn't originally]
So, to get real-time UIs you'd need to: 1. revert CPUs to pre-2000s speeds; 2. write the apps to target such; 3. disable ability to run multiple applications at the same time. Noone's gonna use that.
I remember the windows lock screen being particularly annoying (an animation had to finish..?) but on Linux Mint I've always been able to start typing away immediately, so much so that I start blindly typing in the password before my monitor has finished turning on. Properly queueing events (keyboard ones at least) should be pretty simple to do properly, but of course many things still get it wrong.
Yes, the screen still locks, it just doesn't eat your keystrokes anymore. If you have Windows Home edition without gpedit.msc there's a registry key you can set instead, you'll have to google it.
1) mlock isn’t meant to be called on all your memory, just something that needs it to operate correctly. The situation the author described where the system comes to a halt as memory contents are paged in and out of memory/disk is (subjectively) worse when the only option is for the OOM killer to begin reaping processes everywhere (which would happen if all apps took this advice).
2) the following excerpt is not how any sane modern OS scheduler works:
> Imagine you have multiple background process running at 100% CPU, then a UI event comes in to the active UI application. The operating system may block for 100ms * N before allowing the UI application to process the event, where N is the number of competing background processes, potentially causing a delayed response to the user that violates the real-time constraint
Modern schedulers calculate priority based off whether the process/thread yielded its remaining execution time in the previous round or was forcibly evicted. Background processes are additionally run at a penalty. The foreground window (on OSes with internal knowledge of such a thing) or terminal owner in the current login session group gets a priority boost. Threads blocked waiting for input events get a massive priority boost.
(But the point stands and it might be a whole lot longer than n * 100ms if drivers or kernel modules are doing stuff.)
This one's interesting since your outcome often depends on what hardware you have. On systems with slow IO, i.e. a slow HDD, it's possible for swapping to make a system entirely unusable for minutes, whereas if swap is disabled the OOM killer is able to kick in and solve the issue in less than a minute. That's the difference between being able to keep most of your work open and none of your work open (because the alternative is being forced to reboot).
Also, the OS will never become the bottleneck for any general purpose software because you'll never have a team of programmers good enough to make it so. All the performance issues be due to mistakes in the application itself.
Maybe some sort of magical framework will figure out how to get the latter without the work in the future.
Unfortunately programs are built out of math and so I can’t usually do this.
Apple iOS actually does do a pretty good job of it, and that is a huge differentiator that I think is a large part of why iPhones feel great to use.
Obviously if you're writing an event-loop web server, then your file IO needs to be non-blocking or you're hosed. On the other hand, if you're reading a small configuration file from disk after your app launches, the most responsive option may well be to just read the bytes on the current thread and continue.
https://web.archive.org/web/20190606075031/https://twitter.c...
https://web.archive.org/web/20211005132519/https://twitter.c...
Oh boy
> File system IO functions belong to a class of functions called blocking functions.
Hasn’t non-blocking IO been a major feature for about a decade now??
On some Windows machines with network-mounted drives, the File-Print-to-pdf dialog takes *minutes* to become responsive, even when all currently open files are on a local drive.
This is the kind of thing the author is talking about. The programmers of that dialog box probably just called a generic "open file dialog" library function, without researching its worst-case performance.
In turn the library writers probably blithely coded something like "check if all mounted drives are accessible", without stopping to consider the worst-case performance.
More like the part of Microsoft that implemented non-blocking I/O for Windows some time ago never bothered to tell the part of Microsoft that writes the generic Windows UI code for things like the open file dialog. Or for Microsoft Office applications, for that matter; I still see Word and Excel block the UI thread when opening a file from a network drive, even though Windows has perfectly good asynchronous file I/O API calls.
no, they just don't care.
The alternative to this is to roll a barebones file picker - there might even be one available in the Windows API.
This doesn't make people use it, and even the ones that try to use it might erroneously expect that open(2) will return in less than a second.
Android (and I believe iOS does too) enforce that.
It's up to the developer to show a meaningful message/animation if an IO operation takes noticeable time.
This is absolutely not true. Android simply detects that there has been no progress on the UI thread "for a few seconds" before force-closing the app [1]. By this time, the interaction has been janky/frozen for WAY too long. If you have seen bad iOS scrolling and lock-ups, you know this as well.
I have worked on mobile software for these apps that have billions of users. When I pointed out how much stuff ran on the UI thread, there was a collective "this is just the way it is" response and life went on.
It's super-depressing.
-----
[1] "Performing long operations in the UI thread, such as network access or database queries, blocks the whole UI. When the thread is blocked, no events can be dispatched, including drawing events.
From the user's perspective, the application appears to hang. Even worse, if the UI thread is blocked for more than a few seconds, the user is presented with the "application not responding" (ANR) dialog."
https://developer.android.com/guide/components/processes-and...
https://developer.android.com/reference/android/os/NetworkOn...
The author is giving the OS a unfair rep here. All three major OS has solutions to this problem. It's more up to application makers to use the tools correctly.
Not sure if this is a problem specifically on Linux or not. I like to think that on Windows, UI-thread/process priority boosts take care of this problem, but I suspect not entirely.
I've been meaning to turn off Intellisense to see if my productivity will increase. What I know for sure: my productivity in VSCode/C++ is hugely tanked compared to Android Studio/java.
I'm not sure we have UI models for dealing with more complicated asynchronous operations on a document model though.
Old versions of Microsoft word used to run pagination and line-breaking on a background thread. The UI would keep a couple of lines around the edit cursor up-to-date, with precisely page layout for anything more than a couple of lines after the edit cursor running asynchronously. (It probably still does, but we probably notice it less these days).
And various elaborate schemes for keeping local and cloud object models synchronized asynchronously.
Oh. And I guess we have compilers and editors that collaborate to take snapshots of code when a compile starts, while allowing editing to continue, and to run Intellisense analysis in the background in the face of continually updating edit buffers. Incredibly complicated stuff, done mostly by brute force of intellect, I think.
Beyond that, I don't think we have a general theory of what to do in user interfaces when there isn't enough CPU to go around.
It's been few years since then so I am light on details.
---
Also, as much as devs hate javascript, it mostly solves this by having 1 thread and offloading all system calls to another via callbacks? So the solution to your issue is: Use electron. (Just kidding, but there is some truth in it)
Not really. It means they chose to degrade performance when too much memory is used, rather than crash.
All options available when "out of memory" are bad. They thought that was the least bad option.
Author is redefining "correct" to mean "has the property I care about, which in this case is performance".
There are many desirable distinct properties of a computer system: correctness, performance, security, etc.
"Correct" usually means something like "gives the right answer". It has nothing to do with performance.
Here's an extremely fast sorting algorithm that's incorrect:
function cobbalSort(arr) { return arr; }The iPhone UI must have either a realtime UI (or something pretty snappy). It responds to finger inputs in a bounded way.
Must be disabled on Samsung flagships, as these are the ones I have experience with, and I wouldn't say touch response is bounded on them, not for any useful definition of "response" and "bounded".
The last mobile device I used that had sane input response was a Sony Ericsson K800i feature phone. Input processing on it was fast, and most importantly, consistent. That means predictable - I could do most stuff on that phone while keeping it in my pocket, because I muscle-memorized the input sequences and how long every action took. Such a thing is impossible on smartphones.
In my experience this is only a problem on Linux. Windows and mac will simply inform you that the system is running out of memory and then you can kill things to set it right. No hung system.
Essentially, composable can recompose with every frame, like for an animation. But, in certain circumstances, this will cause allocations for every frame.
For example, a modifier that is scoped to something like BoxScope. You can't hoist it out of the composable, it has to be in BoxScope. If that scope is animated on every frame, that modifier get re-allocated, every frame. That could be a lot of pressure on the GC.
Edit: Then again, its hard doing anything realtime in GC languages like Java / Kotlin, maybe its possible if doing 0 allocations per event.
UI should be run in a separate processor in real time.
I am tired of clicking to see the screen change after I click and the click registers on the new screen, not what I clicked on.
- I/O like file access or networks, you can punt to an async framework like Rust's Tokio
- CPU number crunching like cracking a password, you can punt to a worker thread (like Tokio's `spawn_blocking`)
The kernel scheduler is smart enough to give time slices to the UI thread even if the password-cracking thread is trying to eat as much CPU as possible, or, God forbid, the network thread is _sleeping_ until a packet arrives. (Networking is a waste of a thread. Most requests and responses are small enough that the CPU could process them on the UI thread, it's just that sleeping the whole thread until the message shows up is goofy)
It's not a lack of processors, it's the fact that good multi-threading and UI totally changes the architecture of a program, and most of us aren't trained for it or incentivized to implement it.
Adding USB ports to a video card would achieve 99% of the hardware requirements. The other 1% is made up of things I have overlooked.
The software would be more complicated, but it would not require rewriting applications. It would mostly be recoding OSes.
Also, the security benefits of linking output the screen to input from the keyboard and mouse in real time would be substantial.
Edit: I was completely wrong, but the article is very interesting.