The 64 Milliseconds Manifesto
apitman.com
apitman.com
A search for "hacker news" took 186 ms, with 111 ms spent waiting.
64.ms itself takes 3 ms for DNS lookup, 39 ms for the initial connection, and 29 ms for the SSL - total of 71 ms before any content download happens.
I'm on 802.11 AC WiFi seated about 5 feet from my router.
Apparently, users are quite willing to wait longer than 64 milliseconds.
I'm no web developer, so I don't claim to have expertise, but it really does seem like an unreasonable expectation. As soon as you have any sort of complex functionality the idea falls apart:
"Wait for the credit card payment to process"
"Download a high resolution image" (that counts as a page of content right?)
"Start watching this 4K video"
"Download a CSV of your 100,000 row data set"
To be honest, I have no way of knowing or testing, because ~180 ms is already faster than I can see or react to.
64.ms itself takes 3 ms for DNS lookup, 39 ms for the initial connection, and 29 ms for the SSL - total of 71 ms before any content download happens.
Then my browser waits for 250ms (slower than Google).
Whether the app/site you requested can be interacted with depends on whether the app/site devs care about interaction latency.
Talking about network response times is irrelevant to the point here. Clients can, and do, make changes on-screen before the network response.
And really, even offline interactive apps have no chance of reacting this fast.
Try opening TextEdit or Mail - they don't even draw the window frame in 64 ms.
Performance is important but normal users only care about it when it's getting in the way of functionality. If we chose our preferences by performance alone, Windows Phone and webOS phones would still be here and Android would be gone, Microsoft Office would be the least popular office suite, and Steam would be the least popular PC game store.
What do you mean? Lots of games run at 60Hz, and lots of games are considered unplayable if they render frames slower than 64ms. You might consider your terminal broken if it took more than 64ms to respond by displaying your input. https://danluu.com/term-latency/
> Try opening TextEdit or Mail - they don't even draw the window frame in 64 ms.
The OS shows you the app is loading. This seems like you're intentionally trying to prove some point and not listen, rather than being open to understanding. Arguing that an app doesn't respond quickly when the app isn't running seems obtuse.
> If we chose our preferences by performance alone,
Nobody said that, so this is another straw man. The post is suggesting that UI & frontend devs use generally accepted interaction principles, and nothing more.
As an avid video game player, 180ms is roughly the visual reaction time of good players. "Sound" to be maybe 150ms or faster however.
However, trained players consistently can "time" and percieve fast events. The throw-break window in Guilty Gear is 3-frames, or 50 milliseconds. Players still hit throw-breaks.
There's also the 1-frame jump, which improves jumping speeds by 1-frame (17 milliseconds), but requiring 1-frame precision.
--------
Street Fighter players have a variety of combos requiring "1-frame links", a window of only 17 milliseconds where the combo would work.
Humans may not "react" to 17 millisecond events, but they absolutely perceive them and can work with them. I've run calculations on musicians (ex: Flight of the Bumblebee), and "proper timing" of songs is also to a precision of less than 100ms.
EDIT: Case in point: a 16th note at 144 bpm (Flight of the Bumblebee speeds) is active for ~104 milliseconds, and probably needs to be "precisely" played at a quarter of that time period. (if you are 25 milliseconds off, a good musician probably will notice)
Humans are way more perceptive of sound than vision. So maybe its not "fair" to study music or musician perception.
My web browser is an interactive software application and it doesn't even start up and put an about:blank page on my screen in 64 ms. Repeat this statement for literally any GUI application, the containing window don't even appear in 64 ms.
And anyway, how big is a "screen?" What is the measurement of a screen of content? Do I have to fill up a 320x480 screen or a 3840x2160 screen? How many characters is a "screen? Pictures? Video frames?"
If you make some kind of transform in Illustrator or Photoshop, why would it be reasonable to give up the quality of the end result just to hit that arbitrary 64 ms goal? Can Illustrator even load/render a simple design like a company logo in 64 ms? A lot of interactive software combines batch processes with live interaction, too, so why would it be reasonable to be unwilling to wait more than 64ms for calculation, rendering, or transformation procedures? Of what use would the unfinished content be if it were shown to me within 64ms?
Essentially no application follows this guideline and never will.
This website is more of an exercise in snark and elitism than anything else.
This is a straw-man. You didn't ask your web browser to open, you asked your OS. Your OS shows you the response to your request to load an app or web browser, and it shows that response immediately. The launch icon blinks, your cursor changes, and a blank frame appears all within the span of a few frames.
> And anyway, how big is a "screen?"
You get to decide. I don't think the post was meant to be taken as so extremely literal as that.
> Essentially no application follows this guideline and never will
Browsers and games already do follow this guideline. I think you're mis-reading what's here, and what's here is already generally accepted and very much used broadly. The audience for this post may be devs unfamiliar with interaction best practices, new web developers who don't think to show changes on screen before a network response, or don't want to do extra work when it seems like a .then() is completely sufficient.
> This website is more of an exercise in snark and elitism than anything else.
Why do you feel it's snarky or elitist? Is it possible you're mis-interpreting this post, or just in a bad mood? As someone who's worked in games and interactive apps, I interpret the post to mean that for user interactions, please keep performance-as-a-feature in mind. There have been many threads on HN on this over the years, and that's just a drop in the bucket compared to all the talk of how to keep interaction from being frustrating to the user.
Well, it should! Why does it take more than 100 million cycles to put up a window? Because it's not a priority, not because it's particularly hard.
> Illustrator or Photoshop
It can render in the background, but it should remain responsive the entire time and anything that doesn't need batch processing should be instant.
> Essentially no application follows this guideline and never will.
Most applications could follow it with only a medium amount of discipline.
Full screen of content within 64 ms.
Sure, I definitely expect a checkbox to do something in < 100ms. Show that its state changed and at the very least show me some kind of loading icon so I know something is happening.
But a full screen of content in 64 ms? Unreasonable, and really a completely undefined unit of measure (what is/how big is a "screen" and how much content can fit within? How do you measure audio or video content? What happens when I'm working with vector graphics that can't possibly render in 64 ms?)
There were plenty of applications in the 80s that could not render a full screen in that timeframe. That did not stop those applications from being useful.
(And I honestly think a DOS directory listing can take longer than 64ms to fill the screen)
I assume the intention isn't speed alone, but part of the user experience. When you interact with an application, you should be able to tell immediately that it's recognized your input and something is happening.
The "Test your click speed" button is fun to play with. Times 50-100 ms for mobile and 150-300 ms for desktop are reasonable long for performing some pre-loading.
[1] https://instant.page/ [2] https://github.com/GoogleChromeLabs/quicklink
- 16ms: one frame
- 32ms: two frames
- 64ms: four frames
So it's really 1, 2, 4. Maybe we could have 1, 2, 3, but I'm guessing the "48ms Manifesto" wouldn't sound too great: they probably preferred to have a nice power of two.
And if I may, I think the manifesto is basically right. Something I have to wait several frames for before it reacts feels sluggish. When I type a character, I want it now, not 5 frames later. When I move the mouse, I want it in sync, not lagging behind my hand. Touch screens are even more exacting. They only feel perfect when latency falls under one millisecond. But that last one isn't achievable in practice on current stock hardware…
> What is advised when these rules, especially the "full screen of content" cannot be met because of things like external dependencies out of the site's control?
This is missing the point, and many people here seem to be attempting to pick the same nit. You never have to wait for external dependencies to show changes on screen. This is how browsers and well designed high traffic websites like YouTube already operate. When you take action, you are shown UI changes on-screen that indicate the action is in progress. The manifesto is saying to acknowledge the user action so the user knows the application received the request. It is not saying everything in the world needs to be done in 64ms. Applications can always acknowledge the request immediately, and yet some still don't.
In the context of a UI element or screen layout, "content" is the appearance of the UI elements on screen. The post, and my comment, isn't referring to web content or video content, it's referring to UI content.
Twisting my words to comment on YouTube's quality is funny, but irrelevant to interaction latency.
Well, most actually could be. And a snappy "acknowledgement" only impresses me the first time. The second time I stare at the loading indicators in disbelief, while waiting for the not-so-snappy action to execute.
Loading indicators are simply one single example of how to respond to user actions on-screen. The idea is more broad than this, the idea is for clients to respond immediately to any inputs. This is true for opening menus, clicking buttons, typing, pushing on the D-pad, waving the Wiimote, pressing on your MIDI controller, waving your arms for the camera... it applies to any and all input devices for all applications.
This idea is not in the least bit controversial, I'm surprised by the amount of push-back here in this thread. Devs of games and interactive apps have always done this. If the machine doesn't acknowledge user input immediately, the user doesn't know if the input was captured, and very quickly decides the machine is slow or buggy.
In contrast to "not-so-snappy actions"? Snappy actions.
As I tried to say before: Instantaneous feedback is great (even necessary) for first time users. But once I get used to an app, waiting for the termination of a loading indicator is just as annoying as waiting for {page|script}-{load|parse|eval}.
Edit: > This idea is not in the least bit controversial.
I agree that content generation in one-digit milliseconds is not controversial.
Typical human reaction time is 200ms or so. This effectively means that for 200-264ms you're not priming to repeat the action or wondering whether it "took", because there's an immediate response and the action itself does something useful quickly.
Waiting for termination of a loading indicator on a web page is irrelevant to this post, it's a tangential topic completely.
"content" in the post is not referring to web page content, it's referring to UI content. We're not talking about content generation in the web page sense.
You get upset when buttons have a 'clicked' state, or something?
It's hard for me to picture how acknowledgment of input isn't a good thing.
If you add intermediate network gear, it's hard enough to get across half the globe in under 180ms (I'know, I live in Argentina and that's my ping time to almost anywhere non local).
You could fake it somehow in some scenarios or with an awful lot of money.
It's not that high if you live in a rural area (And I'm not even talking about a third world country, I live in France). Most of the time, the main culprit is Bufferbloat[1].
I'm currently playing World of Warcraft with 250ms of ping, and even for solo level up it's annoying. League of Legends with 350ms isn't fun at all…
A game dev can't fix it, but it should not be taken as a ground truth of the world. It needs to be fixed. Combine that with a smaller amount of prediction, and things come out pretty well.
This is a hint that you're writing for the wrong platform. Write real apps.
Your second sentence explains why the first is wrong. Yes, you can’t have the entire world using a single server but it’s never been easier to get servers around the world and the new edge compute services are further improving that.
Only if your «world» means «developed countries only»…
Akamai's coverage is better still; they've got just about the entire world covered: https://www.akamai.com/us/en/resources/visualizing-akamai/me...
Cloudfront, Google, Fastly, etc. all have pretty good coverage, though you're right they're spottier when it comes to S. America, Africa, and a few other places. Their POPs are, from what I understand, have higher capacity but are less dense. Good for cost savings, though probably worse for latency.
There's also the consideration that most web sites have most of their users in the developed world (though that equation is shifting quickly for a some).
Out of curiosity, what is the significance of "<<" and ">>"? I'm assuming they're not bitwise shift left and right. I've seen them more often as of late, and they seem to come mostly from non-Americans. I ask because they're quite difficult to run an internet search on without knowing what they are called.
They're quotation marks in many European languages and I believe (though I'm not sure) that the correlation isn't non-Americanness, it's non-English-keyboardness. People from English-speaking countries will still use the quotation marks you're used to.
This is a really good example. Did you look at Africa? For the record, 1.2 billion people live there nowadays ;)
> Out of curiosity, what is the significance of "<<" and ">>"
Oh, you mean '«' and '»'? They are French quotation marks[1], I try to use the English ones '“' and '”' when I'm writing English, but sometimes I forget. Same for the space before the question and exclamation mark (there is a space in French, but not in English).
Yeah, I looked at it. There are datacenters within 2000km of everywhere, and that's including the Sahara. About 1500km outside of the Sahara.
1500km as the crow flies, so let's say 2400km of fiber. At the speed of light in fiber, that's 11.5 milliseconds. Double it, add time for ten routers, that's 27 milliseconds to go from any ISP line to a Cloudflare data center.
Sounds good to me. If there are infrastructure problems that's a separate issue from whether Cloudflare has enough POPs.
Beyond the obvious point that the growth of commercial cloud options make it much easier to get servers around the world without having to research providers and negotiate contracts in each country, you might in particular want to spend a couple of minutes looking at current CDN coverage, especially if your experience is either limited or old. Coverage is a lot better globally than it used to be — gone is the time when “Africa” either meant servers in Europe/Israel or, at best, one POP in South Africa. Next, think about services like CloudFlare Workers, Lambda@Edge, etc. and ask what percentage of people on the planet are likely to be close to one of the POPs where you can run code:
https://www.cloudflare.com/network/ — 194 cities in 90 countries
https://aws.amazon.com/cloudfront/features/ — 191 locations in 33 countries
This assumes that we need remote servers to do our computing for us. :~) P2P software doesn't require remote servers, plus it means your apps still work offline.
The right answer for web and networked applications is to update the screen with UI that acknowledges the user action and shows the user that the result is pending. Ideally, progress is visible, but that’s tangential to the point of the manifesto.
A client can, in fact, almost always respond to actions within these time constraints. The point is to do something, rather than wait for the network response.
More useful:
People can often respond to things within a few frames (1/60 second = 16ms). For some interactions, it's useful to respond within one frame. For others, more latency is acceptable. For yet others, sub-frame latency is required (e.g. VR/AR). Light can move ~3,000 miles in that time, which is a hard limit on latency.
* Websites making controversial claims should state clearly the claim in the first paragraph
* Controversial claims should be backed up by at least a summary of a well reasoned argument within the first 5 paragraphs.
* Controversial claims should be backed up by citing sources.
Sure, with all those laggy applications we have right now, we tend to become desensitised. That's not an excuse, though.
Also note that in some cases, we're not even close to acceptable latency. Finger tracking on touch screens for instance, require a one millisecond response time to feel perfect. Otherwise the objects we drag will lag behind our fingers.
The 64ms goal by itself is clearer and more useful than this combo of three goals.
I thought it was a presentation or something and I needed to use arrows to navigate it..turns out.
Well not really sure what the point of the thing is tbh.
Curl will do one round trip to your resolv.conf configured resolver
Then you have one round trip for tcp.
Then another for TLS hello (presuming TLS 1.3 or false-start)
Then one last round trip for HTTP.
Assuming cached DNS, you need a ping to the server of about 21ms to get content via https in 64ms. And that's assuming you can get enough content for a screens worth in the initial congestion window (10ish packets, so let's say 15kb). Good luck!
Oh, and my DSL connection adds at least 20ms to the path, and that's not unusual.
My intuition says that's steep. Noticeable change 1 frame later might be doable with a web app, if the app is structured well and focuses on fast response times. However, I don't think actionable information two frames later, or a full screen of content four frames later is feasible for a network-based application, especially given the fact that just crossing local layer 3 network segments can add 10s of ms to your latency, depending on the router's load.
Really, these numbers strike me as only achievable by an application where all data and processing is done locally on the computer that controls the display. Even then, these target numbers still feel steep.
If the user does something interactive to warrant further layout changes, then you're allowed to start changing things again.
The idea is to make it cost something to change the layout. Then people will have to actually try not to do it excessively.
https://developers.google.com/web/fundamentals/performance/r...
In many cases the Input Lag ( Especially Touch Screen ) in itself is more than 16ms. That is excluding the processing from CPU to GPU rendering and your Display Lag.
If you really want to minimize latency on a modern system, and can choose whatever commodity hardware you want, you can get the input+rendering+scanout delay below 10ms. At that point it's a matter of how often your screen accepts new frames. And with a high-end, frame wait plus pixel response time can also be under 10ms.
20ms for local end-to-end latency is a completely achievable goal.
https://www.blurbusters.com/gsync/gsync101-input-lag-tests-a...
Their best combo here achieved an end-to-end latency of 12-15ms, on a 200/240Hz screen with counterstrike running at 2000+ fps.
At a more mainstream 144Hz to the screen, running the game at 'only' 288Hz, end-to-end latency was 13-19ms.
Even at 60Hz they could easily keep the latency below 30ms.
As for touch screen technology, the Apple Pro Pen I believe is only 20ms and is state of the art. While I do believe that developer ux has contributed to lazy resource management, it's also clear that sub-60hz latency has regularly required exotic hardware.
Touch is hard to keep ultra fast. But with physical buttons it's mostly apathy.
A couple of years ago I was thinking about this either and someone suggested a study defined 0.1, 1 and 10 seconds as limits: https://stackoverflow.com/questions/17138585/what-is-the-max...
16 ms → 62 Hz
32 ms → 31 Hz
64 ms → 16 Hz.
In particular:
> 0.1 second is about the limit for having the user feel that the system is reacting instantaneously, meaning that no special feedback is necessary except to display the result.
[1] https://www.nngroup.com/articles/response-times-3-important-...
By focusing on UI responsiveness you vastly improve discoverability and interoperability because it naturally decouples the problem into a model/view/controller design. Best of all you can't mis-design: the responsiveness is your constraint, and if the design doesn't give you responsiveness you KNOW it is a poor design.
This is the ultimate quality litmus test.