Computer Latency: 1977-2017
danluu.com
danluu.com
For the next step I want to make it capable of injecting artificial latency, and then do A/B testing to determine (1) the smallest amount of latency I can reliably perceive, and (2) the smallest amount of latency that actually bothers me.
This idea was also inspired by this work from Microsoft Research, where they do a similar experiment with touch screens: https://www.youtube.com/watch?v=vOvQCPLkPt4
If you wonder why we no long we have "twitch" games, this is why. Old school games had a tactile aesthetic lost in the blur of modern lag.
Sublime Text: 17–29 ms
iTerm (zsh4humans): 25–54 ms
Safari address bar: 17–38 ms
TextEdit: 25–46 ms
Method: Record 240-fps slo-mo video. Press keyboard key. Count frames from key depress to first update on screen, inclusive. Repeat 3x for each app.Well, that's how it was done 10 years ago.
So e.g. put a clock in the video
Currently, there is a whole pile of steps to update a UI. The input system processes an event, some decision is made as to when to rerender the application, then another decision is made as to when to composite the screen, and hopefully this all finishes before a frame is scanned out, but not too far before, because that would add latency. It’s heuristics all the way down.
With adaptive sync, there is still a heuristic decision as to whether to process an input event immediately or to wait to aggregate more events into the same frame. But once that is done, an application can update its state, redraw itself, and trigger an immediate compositor update. The compositor will render as quickly as possible, but it doesn’t need to worry about missing scanout — scanout can begin as soon as the compositor finishes.
(There are surely some constraints on the intervals between frames sent to the display, but this seems quite manageable while still scanning out a frame immediately after compositing it nearly 100% of the time.)
This Xorg dude did exactly the tuning you want on wayland https://artemis.sh/2022/09/18/wayland-from-an-x-apologist.ht...
Adaptive sync is beneficial for graphically intensive games where you can't always render fast enough, but IMO this should never be true for a GUI on modern hardware.
That’s a matter of perspective. If your goal is to crank out frames at exactly 60 Hz (or 120 Hz or whatever), then, sure, you can’t send frames early and you want to avoid being late. But this seems like a somewhat dubiously necessary goal in a continuously rendered game and a completely useless goal in a desktop UI. So instead the goal can be to be slightly late for every single frame, and then if you’re less late than intended, fine.
Alternatively, one could start compositing at the target time. If it takes 0.5ms, then the frame is 0.5ms late. If it goes over and takes 1ms, then the frame is 1ms late.
But yeah, for non-fullscreen it helps. See https://github.com/swaywm/sway/pull/5063
We've got servers in 200+ cities around the world, and ask them to ping each other every hour. Currently it takes our servers in Tokyo and London about 226ms to ping each other.
We've got some downloadable datasets here if you want to play with them: https://wonderproxy.com/blog/a-day-in-the-life-of-the-intern...
Some random example:
Azure Application Insights can be deployed to any Azure region, making it feel noticeably snappier than most cloud hosted competitors such as New Relic or logz.io.
ESRI ArcGIS has a cloud version that is "quick and easy" to use compare to the hosted version... and is terribly slow for anyone outside of the US.
Our timesheet app is hosted in the US and is barely useable. Our managers complain that engineers "don't like timesheets". Look... we don't mind timesheets, but having to... wait... seconds... for.... each... click... is just torture, especially at 4:55pm on a Friday afternoon.
Because our product is global, our backends are replicated worldwide too. Otherwise we’d be forcing the pain we go through daily on our users too
With 40 Mm circumference and 300 Mm/s light speed in vacuum you have as a physical limit latency below 70 ms from opposite places in the world.
Could you expand on this? The speed of light is constant, no?
Looking at the map with the blue dots, a cool rainy-day project would be to show the pings and pongs flying back and forth :3
Chicago London New York
Chicago — 105.73ms 21.273ms
London 108.227ms — 72.925ms
New York 21.598ms 73.282ms —The usability is worlds better than what we have now, even comparing a 1990 computer with a 25 MHz m68030 and 16 megs of memory with a four core, eight thread Core i7 with 16 gigs of memory. Interestingly, the 1990 computer can have a datatype added which allows for webp processing, whereas the Mac laptop running the latest Safari available for it can't do webp.
We've lost something, and even when we're aware of it, that doesn't mean we can get it back.
https://news.ycombinator.com/item?id=25290118 (December 3, 2020 — 454 points, 259 comments)
https://news.ycombinator.com/item?id=16001407 (December 24, 2017 — 588 points, 161 comments)
> UIKit introduced 1-2 ms event processing overhead, CPU-bound
I wonder if this is correct, and what's happening there if so - a modern CPU (even a mobile one) can do a lot in 1-2 ms. That's 6 to 12% of the per-frame budget of a game running at 60 fps, which is pretty mind-boggling for just processing an event.
Speaking of games: I had just the other day the realization that we should look into software design around games if we want proper architectures for GUI applications.
What we do today instead are "layers of madness". At least I would call it like this.
Why? Are there any technical reasons?
I think this is a pure framework / system-API question.
They process all normal keystrokes locally, and only send back to the host when Enter and function keys are pressed. This means very low latency for typing and most keystrokes. But much longer latency when you press enter, or page up/down as the mainframe then processes all the on-screen changes and sends back the refreshed screen (yes, you are looking at a page at a time, there is no scrolling).
Of course, these days people use emulators instead of hardware terminals so you get the standard GUI delays and the worst of both worlds.
Every computer systems since then has been a head shaking disappointment, latency-wise.
I read something about this being intended for use with high end wired gaming mice, where the end to end latency between mouse and cursor movement is theoretically lower if the signal doesn't go through the USB bus on the motherboard, but rather through whatever legacy PS/2 interface is talking to the equivalent-of-northbridge chipset.
https://static.tweaktown.com/content/1/0/10071_10_asus-rog-m...
Latency is lower because it's interrupt-based and much simpler than the polled USB stack. IMHO if you're going to always have a keyboard and mouse connected to the computer, it makes perfect sense to keep them on the dedicated simpler interface instead of the general USB; especially when the dedicated interface will be more reliable. The industry may be partly moving away from the "legacy-free USB everything" trend that started in the 2000s, finally.
AFAIK all SuperIOs support a pair of PS/2 ports, so from a BoM perspective it's not an extra cost to the manufacturer, but they still market it as a premium feature.
Another comparative anecdote I have is between Windows XP and OS X on the same hardware, wherein the latter was less responsive. After seeing what GUI apps on a Mac actually involve, I'm not too surprised: https://news.ycombinator.com/item?id=11638367
I do wonder about MS-DOS on different machines, technically it would be the BIOS and VBIOS doing much of the heavy lifting so vendor variations might have a real impact here. Same as a CSM on UEFI; DOS wouldn't know the difference, but traversing a firmware that as as complex (or more complex) from DOS would cause a whole lot of extra latency.
10K iterations of $null=$null took 38 milliseconds on my laptop.
10K parses of the letter "a" took 110 milliseconds.
10K parses of a 100K character comment took 7405 milliseconds.
10K parses of a complex nested expression took just over six minutes.
You're probably imagining a lexer written in C that tokenizes a context free language and does nothing else. In PowerShell, you can't run the tokenizer directly, you have to use the parser, which also builds an AST. The language itself is a blend of two different paradigms, so a token can have a totally different meaning depending on whether it's part of an expression or a command, meaning more state to track during the tokenizer pass.
On top of that, while it was being developed, language performance wasn't a priority until around version 3 or 4, and the main perf advancement then was to compile from AST to dynamic code for code blocks that get run a minimum number of times. The parser itself was never subject to any deep perf testing, IIRC.
Plus it does a bunch of other stuff when you press a key, not just the parsing. All of the host code that listens for the keyboard event and ultimately puts the character on the screen, for example, is probably half a dozen layers of managed and unmanaged abstractions around the Win32 console API.
But yeah, I agree with the other comment that powershell is likely adding less than 1ms.
Instant response. A full reboot was a control break away. Instant access to the interpreter. Easy assembly access.
I thought, it executed.
I remember our school moving from the networked bbc’s to the PC’s and it was a huge downgrade for us as kids. Computer class became operating a word processor or learning win 3.11 rather than the exciting and sometimes adversarial (remote messaging other terminals, spoofing etc) system that made us want to learn, to just more drudgery.
Having an ordinary key on the keyboard that would effectively kill -9 the current program and clear the screen was a crazy design decision, especially for a machine where saving data meant using a cassette tape!
Your program would still be in memory with an \>OLD command.
As long as it was a basic prog, the machine code loaded *RUN was lost and had to be reloaded from tape, yes.
A pain for games but I don’t really recall accidentally pressing the break key much it was out of the way up right.
I could talk about this all day!
But any data was lost and I saw break get pressed accidentally fairly often at school and amongst friends.
Example: https://youtu.be/ls-PxqRQ35Q?t=178
In general, predicting user input to reduce latency is a great idea and we should do more of it, as long as you have a good system for rolling back mispredictions. Branch prediction is such a fundamental thing for CPUs that it's surprising to me that it doesn't exist at every level of computing. The JavaScript REPL's (V8's REPL) "eager evaluation" where it shows you the result of side-effect free expressions before you execute them is the kind of thing I'm thinking about https://developer.chrome.com/blog/new-in-devtools-68/#eagere...
While I am not a fan or proponent of AR / VR. One thing that will definitely be an issue is latency. Hopefully there will be enough incentive for companies to look into it.
> Computer results were taken using the “default” terminal for the system (e.g., powershell on windows, lxterminal on lubuntu), which could easily cause 20 ms to 30 ms difference between a fast terminal and a slow terminal.
If that was the only source of noise in the measurements then ok, maybe, but compounded with other stuff? For example, I was thinking: the more time passes, the further we drift from the command-line being the primary interface through which we interact with our computer. So naturally older computers would take more care in optimizing their terminal emulator to work well, as it's the face of the computer, right? Somebody's anecdote about PowerShell performance in this thread makes me feel more comfortable assuming that maybe modern vendors don't care so much about terminal latency.
Using the "default browser" as the metric for mobile devices worries me even more...
I like Dan Luu and I SupportThismessage™ but I feel funny trying to take anything away from this post...
Wonder about hat as we just talked about the importance of sub-second response time in 1990s (full screen 3270 after hitting enter; even if no ims or db2 how can it be done …). The terminal keyboard response is fine (on 3270). Network (sna) …
1977 still have mainframe and workstation.
https://colin-scott.github.io/personal_website/research/inte...
I didn't dig into the text blob to ferret that out.
Did anybody?
Because this doesn't pass the sniff test for data I want to trust
He does mention that he includes the keyboard latency; other latency test results he found exclude that step.
I find it fascinating that you think no one would bother to read details about a nerdy subject on Hacker News. Why else are we here?
---
Does this actually even matter today when every click or key-press triggers dozens of fat network request going around the globe on top of a maximally inefficient protocol?
Or to summarize what we see here: We've build layers of madness. Now we have just to deal with the fallout…
The result is in no way surprising given we haven't refactored our systems for over 50 years and just put new things on top.
So that is about 0.175ms.
I know people might think this over the top, but AFAIK this is far cheaper than say to work around latency within the software system. Next we need to work on Display Latency and Display Connection Latency.
[1] https://www.razer.com/gaming-keyboards/razer-huntsman-v2
For completely imperceptible computing, display refresh must be dealt within 50 ms, roughly. Input must be sampled, all relevant computations for the current rendered display must be computed, or in easily accessed storage, and all updates to the display propagated to the framebuffer. Sensitive humans can notice as little as roughly 50 ms of lag or jitter.
This means for a program dealing with highly interactive graphics that are linked to an input device, the core event loop must execute in less than 50 ms or so. Even with current blazing-fast machines, this is a tough challenge for anything complex. If this deadline is not met, potentially perceptible lag in rendering will occur. In a 3D rendered scene, the graphics may perceptibly hang or tear for a couple frames. This is perceptible by a human, though with sustained suspension of disbelief, we can mostly ignore it, much as we can ignore the various visual artifacts in 24 fps cinema.
It just adds up to the overall latency.
Also real latency of web-pages is measured in seconds these days. People are happy when they're able to serve a request in under 0.2 sec.
For example, I had a support call with Azure asking them why the latency between Azure App Service and Azure SQL was as high as 13ms, and they asked me if my target user base was "high frequency traders" or somesuch.
They just could not believe that I was expecting sub-1ms latencies as a normal thing for a database response.
I think I'm just learning this the hard way, given the down-votes of the initial comment. :-)
Maybe people really don't see the issue with adding layers after layers of stuff, and that we've reached, no, surpassed even, some tragicomic point? Computers are thousands of times faster, yet the end-user experience becomes more sluggish with every year passing. We have an issue, I would say. And it's actually not even funny any more.
Managed to fix the problem and restore performance by suggesting we disable the MySQL query cache which was slowing down the system with more cores.
And this was nothing to do with banking or high frequency anything. Just code that did lots of small queries to fill a screen full of data to support persons.
This approach for a fix sounds very unintuitive.
Did you manage to understand why this was like that? The explanation is likely interesting!
What I could think of: Context switches across cores constantly invalidated data in the CPU caches. In such a case CPU pinning would help maybe. (Just speculating! I'm not an expert on such things. But I know that the CPU cache, and memory I/O in general, is the single most important topic when dealing with usual performance issues on modern CPUs. Today's CPUs are only fast when they have their data available in their local cache(s). Fetching form RAM is by now like fetching from spinning rust a decade ago; it will kill your performance no mater how fast your CPU is).
Tangentially related: https://news.ycombinator.com/item?id=14888360
It was meant to accelerate PHP code in the 90s with single core CPUs. It was never designed to be scalable and in-development (back then) versions had already disabled it, which gave me the idea that it was something to try.
Indeed unexpected. But interesting to know.