Performance Matters (2019)
hillelwayne.com
hillelwayne.com
But, software that ignores, delays or discards user inputs? Absolutely F*$#ing Unacceptable.
Unfortunately, most user-interface software (including the ePCR software described in this article) commits this "Unpardonable Sin" of UI software. So, the EMTs, being insulted by the software developer, just don't use it.
If you write UI software, you have one job. To utterly cherish user input, painstakingly preserve it, and lovingly and promptly provide user output. That's it. Almost anything else is optional, and can be forgiven.
That said if this UI failure was easier for the developers to code, such as no code and download from NPM, then it is most correct even if both the business and end user are both catastrophically harmed immediately.
Do a bug report I suggest.
Software should never break because flesh-and-bone human user is "too fast" for it.
My guess is, whoever came up with that threshold has never played any console game.
I've once worked on a system that was objectively slow. Some actions would take seconds to complete. It's not like people would refuse the use the system, it was the only way to do their jobs. The public had no choice, either.
Initially, I didn't think it was such a big deal. Yes it was a bit slow, but nothing _terrible_ and why did it matter, since there were so many other procedures that took way longer than the software step. It was a government facility, lots of red tape. Surely optimizing some of the paper processes would be more beneficial.
But we optimized it anyway. Not because the UI was slow(which we didn't even measure directly), but because it was slow for a reason - backend processes were taking time, and resources. If we could optimize, we thought, we would save resources (and also increase the runway until we had to expand again, buy hardware and the like).
First round of optimization was completed. Lots of low hanging fruits(many database-related). Testing indicated that the backend would be faster by almost an order of magnitude. Deployed.
THE SYSTEM BECAME SLOWER.
How was that possible? Maybe we had missed something. We found more bottlenecks, optimized those. Deployed again.
Everything is slow again. What's happening?!
We went to the 'field' - as in, talked to the users. Well, this is what we discovered: we saw the backend working harder - because the users were more productive! Instead of staying after the facilities were closed to the public and then catching up on whatever paperwork they were unable to enter in the system, they could process everything almost as fast as people came in. Which means that they could go home early.
What we saw as "the system being slow again" was due to bad metrics - we were watching backend load, not UI response times. Because, admittedly, that didn't bother us directly. Not until we talked to the poor souls that had to use the system. Lots of async processes got triggered and queues filled up, but it didn't really matter to the users, as long as the UI said things were submitted and they could move on.
After that, we optimized the heck out of all the UI interactions we could find.
1. It highlights that user impact determines success more than resource metrics.
2. It shows that assumptions infect even when you think you’ve accounted for assumptions.
Honorable mention: don’t just prove the abstraction and call it done. It’s got to prove itself in real world scenarios before you can trust its claims!
However, I will always avoid - if I can - using a product where the UI itself is slow, sluggish or bad. This includes many modern websites and also quite a few apps.
Responsiveness is the key. If I click on something, then something needs to happen. People can deal with a loading bar, but everyone gets terribly confused and / or annoyed by a sluggish drop-down menus, buttons that seem to fire only after a couple of seconds, or, my favorite, a scrolling action that is either sluggish or leads to load processes that interrupt itself.
This was when I first realised that UI delays have a non linear effect on productivity. The productivity loss from 10x latency increase is more than 10x. I don’t know how to explain it other than by example, but if you have multiple 1000ms delays you’re more likely to be distracted - to check your email or get a cup of tea or answer a slack message - than if you’re just ploughing through the work. That was my experience anyway.
A few years later I saw the problem in a different context (ETL), so I drew up an informal scale from small to long delays using this kind of human centred thought process, and I realised that, for example, a 1 hour delay to process ETL data is effectively identical to a 6-12 hour delay, because the user will set the process off and go do something else productive for the rest of the work day, which had a huge impact on project delivery (of course).
Once I had seen this in action, I couldn’t unsee it, and I spent the next few years getting rid of a pile of bottlenecks to make the UI performance and feedback loops as fast as possible, for example by refactoring complex operations. My proudest moment was reducing a very complex, 8 hour process down to a 5 minute (perceived) process by doing most of the work in advance, so the data was almost entirely ready before the user needed it. Our users thought the system had broken because it finished so quickly!
I’m sure there are formal studies about this phenomenon, but it was so obvious once I’d experienced it, and our focus on UI latency made a huge difference to both the quantity and quality of the work we and our customers could get work done in a given time period.
I think the "The Economic Value of Rapid Response Time" from 1982 is rather illuminating:
https://jlelliotton.blogspot.com/p/the-economic-value-of-rap...
The problem here isn't performance. The problem here is that the company which is building this software are so remote from end-user that they don't hear feedback.
If they knew that performance is the problem and multidollar contract with some huge network of hospitals were at risk, I can bet you, this performance issues would have been fixed really fast.
More to the point: unfortunately the software needs in the world, and the ways the world are underserved by software, are competing for resources and organizing talent with organizations that fundamentally don’t serve those needs. And both categories are competing in a limited pool, because software’s capabilities have outpaced the available talent to take advantage of those capabilities.
Good for my bank account, I guess. But bad for people and the world we (I) inhabit. And bad for me too, probably even more than I know.
Anyway this is far afield of parent comment’s point, but I felt it was a good place to add a little depth as an engineer who gives a damn but rarely sees the opportunity to apply damns.
I remember back when FogBugz was a thing and Joel claimed that the correct way to provide customer service was track bugs, have one owner, and only allow the original reporter of the bug to mark it as closed. I'm sure that's not completely feasible but It's surprising to me how I know of ZERO software developers that follow anything even remotely close to this practice except for a few open source projects.
Want to report a bug on Windows? Photoshop? Almost any game ever? Good luck finding out if a dev ever saw your report and that it didn't just get dropped by some underpaid customer service center rep.
The current industry standard is directing such reports to official "support forums", where users try to help each other and nobody with any relevant expertise is present. After all, why would the crew of a modern and enlightened software project stoop so low as to talk with actual users, where extensive telemetry provides all the information they need?
s/, but only slightly.
in the case of EMS, it sounds like they could recognize it, but probably didn't report it back to the company. Why bother? They have used their product; they see it as "good enough".
But very often, end users get frustrated by things like input latency, and can't express what is making them frustrated in specific terms. So they tell their IT department that "it's slow", and IT goes back and starts hammering on the __ team (networking, server farm, whatever).
It's remarkably powerful if you can help you users develop vocabulary to recognize and report what they experience. (it's also flipping hard)
> The ambulance I shadowed had an ePCR. Nobody used it. I talked to the EMTs about this, and they said nobody they knew used it either. Lack of training? No, we all got trained. Crippling bugs? No, it worked fine. Paper was good enough? No, the ePCR was much better than paper PCRs in almost every way. It just had one problem: it was too slow.
That is a crippling bug. The UI is a soft real-time system, [0] and it's doing such a poor job of meeting its deadlines that the user considers the system to be unusable.
If you're writing an autopilot system, it's not enough for the system to eventually make the right decision, it must arrive at the right decision before the deadline. Failure to do so would by definition qualify as a bug.
> Most of us aren’t writing critical software. But this isn’t critical software, either: nobody will suddenly die if it breaks.
Won't they? If the software corrupts the patient data, someone could die, right? Elsewhere the article essentially says as much:
> An error might waste valuable time as nurses chase invisible problems or ignore obvious ones. Worst case, it leads to the wrong treatment. In emergency situations these mistakes can be fatal.
[0] https://en.wikipedia.org/wiki/Real-time_computing#Soft
edit I see brundolf's comment already makes some of these points
Consider the new trend in login forms these days whereby you are forced to enter you username or email THEN press a button THEN enter your password THEN press another button. What used to be simple is now no longer simple. Why is it like this? To accommodate the x% of people who seem to get confused in some manner when presented with "too many choices". I forget the reasoning now, and disagree 1000% but don't want to sidetrack my argument any further.
Anyhow... compare the original form to an imagined GUI since we're not presented with the software to make a proper comparison. If you needed to fill out the paper form quickly you can tick..tick...tick..tick..write a bit.tick...tick..read..write.etc.
Now with the software, does it use a mouse? Probably not. It's most likely a touch screen. Possibly a tablet. So now every choice required more interaction. More choices. Opening a dropdown? So you mean to say the developers have decided to hide important information until you request it? The paper version has everything you ever need to know at a glance. One side of a sheet, no need to even flip it over.
This is all conjecture, but I experience this type of thing all the time when a software "solution" comes to fix a real-world "problem". It can be done well. Most UK Govt websites are incredibly well-done IMHO. But usually they're not.
FWIW, this is done to accommodate SSO (single sign-on), which matters for any software that's going to be used in corporate or governmental environments. You have to submit your login first, because it's used to determine what authentication method and provider to select.
That said, I hate this flow too, and there must be a better way. It also doesn't excuse products that do not support SSO that still implement such split flow anyway.
Determine the Auth provider from the username
Feed both username & password to the provider
And be done with it?
I’m not clear why the user needs to be made aware of the SSO setup..
Story time: I got hit by performance these very days. The app I'm developing for one of my clients has to process images from an USB camera. Under my development everything is dandy. Works like a charm, images gets processed and when the user hits the on-screen button that image gets stored in database as part of the entire process. Neath, yeah?
Well, it turns out my client is using a cheap from last decade tablet (I'm a purist so this decade will end on 31st Dec 2020) that due to processing 30 frames each second from USB camera, has little time for actual GUI responsiveness. And the app feels sluggish, with 1 second delay between my tap and the combobox firing up. Turns out, it didn't need to actually process all those 30 frames each second. On per second will suffice, hence I've implemented to only process on frame every second. More than enough for customer's needs and now the app is also flying on that old hardware.
My 2 cents.
That's merely a workaround, not a fix. Next day someone will use a 8k camera on the same slow tablet. Another day someone will run your app in parallel with with some other process consuming all CPU cores.
A fix would be making so that however slow the computer is, processing frames from the camera doesn't affect GUI latency, at least not by much.
You probably gonna need multithreading for that. And if that 2010 tablet only has a single CPU core without hyperthreading, you might need to adjust the priority of that camera's thread. But it's all doable.
You have seat-belts on your car? Is that a fix or a workaround? Because a fix would be to actually have a car that doesn't crash at all. But would not be economically viable.
You have plastic insulator around your electricity wires to prevent you getting electric shock. Is that a fix or a workaround? Because a fix would be to actually have continuous current at max 12V as power lines. But that's not economically viable.
You have kids going to school and strangers are educating your kids a good portion of their life, molding them sometime against your values. Is that a fix or a workaround? Because a fix would be to have them home-schooled under your eye. But that's not economically viable.
I can do this all day.
Your examples are wanting.
Wires are also insulated to stop them from touching each other, not just to prevent electrocution. A 12V wire touching another can definitely result in sparks and fire. It can even amputate your finger if you bridge with your wedding ring.
In virtually all situations wire needs insulation - it's a feature not a workaround.
Like I said, it all comes down to economics.
Look up https://en.wikipedia.org/wiki/Knob-and-tube_wiring that is in some cases operational even after a century.
Lmao. I like this perspective. I guess all if life is basically a workaround to not dying.
For experienced developer using the right tools/architecture it's not even that expensive to develop.
Did they ask the people who would actually use the software if they would use the software? All the way from design to implementation to the team that purchased the software. This is what happens when no one checks with the person that will actually use the software everyday.
You end up with checked boxes and wasted resources.
Here is the full quote: "We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil. Yet we should not pass up our opportunities in that critical 3%."
What most people do is ALSO pass up optimizing that critical 3%, so software gets slower and slower. What he (probably) really meant is: "don't micro-optimize what a compiler can do better, optimize your algorithms and data-structures"
Some common excuses for bad code:
- Shuffling data around needlessly? Not a problem. It's fast enough.
- The algorithm is bad. Not a problem. It's fast enough.
- This is obviously bad. Not a problem. It already takes minutes, so a few more won't matter.
- The code is slow. Not a problem. It's only used internally.
[1] http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Considering the no. of fields on that form, with that kind of lag, I'll be ditching it too. I mean most of it is check-boxes; just replicate the damn thing on html. Though without knowing more about the setup, I would guess that hardware with terrible input is also involved.
I went back and tested the various changes I'd made to that point and thankfully found that most of them were needed to sustain that performance improvement but it was a good reminder to break out the profiler first, not last.
I strongly disagree. Anyone on the project would have (hopefully) known what it was being used for and how critical timing is for that task. Heck, the entire reason the project exists is to expedite the task. That translates directly to minimum-latency being a project constraint, so this to me just sounds like a project that failed to meet its constraints.
If you're an application developer, you're generally going to have a clear idea of what the priorities are for your project and what does and doesn't matter (in terms of performance and otherwise).
Maybe there's a case here for systems/library programmers erring on the side of performance because they genuinely don't know who will be using their code for what. But even then- if a library isn't fast enough (put differently: if the library's priorities diverge from the application's priorities), the application developer should know that and not use that library for their application.
This trend of universalist performance-puritanism is exhausting. Performance is one of many priorities that a project may have. Know your target user, set your priorities accordingly, and develop your code in line with those priorities. No single priority gets to be so important across an entire field that everything else takes a back-seat.
There was only one problem with Java. Even on new machines, you would get noticeable random pauses doing simple things. This never happened in C++ apps even with full MFC bloat. The GC just wasn't that good yet. Programmers would also include FAR more code and libraries because it sped development up. Even today, C++ applications only include libraries if they absolutely have to because of how painful it is.
All of C++'s disadvantages become advantages when you look at them from the lens of performance. Manual memory management means that even if you are slow, you are consistently slow. You don't get the peaks and valleys that really irritate humans who are incredibly sensitive to rhythm. Memory leaks are honestly not even a problem unless you have long running servers. Even then, you could honestly just restart. The near-complete lack of standard libraries and complete lack of a package system highly reduced the amount of fluff you had. If you wanted it in your software, you had to do it yourself. This led to very simple, non-pretty, static interface that just tried to look like Word without the Toolbar of Death. The best optimization has always been to Do Less Stuff.
It's only going to get worse too. The same companies who think C++ is too complex for their programmers are going to laugh at Rust. People in Java or .NET shops aren't going to move over. "Native" is increasingly becoming a JavaScript space which has all the terrible performance of JavaScript usually combined with the slowness of internet connections. For good and ill, package management is now standard.
Microsoft Office used to run fast on a 486 with megabytes of ram. Think about that. Are you a more complex or intensive program than Word? I bet that ePCR has far higher specs. The program running on it isn't inherently more complex. It's literally a form filler, but it isn't used because of how ridiculously bloated even the most basic software has become. In the 90's, you could have made that machine by installing Windows on a machine and programming a VB program in a few weeks that hooked up to a printer. It would have been more responsive than a state-of-the-art ePCR that sits unused today.
Not your main point, but having to kill a program or an OS and restart whatever I was working on because some process can't be bothered to give back memory it doesn't need (and that I now do need) is a major pain.
Firefox is a counterexample to this statement.
Rust appeared because experienced Firefox developers simply couldn't manage C memory manually, C memory with garbage collection, C++ memory manually, or C++ memory with garbage collection.
Also last-minute loading and verification of classes. Interpretation overhead and JIT compilation wouldn't help either, especially on a single-core machine. Many aspects of Java make more sense for servers than for desktop UIs. (This might finally be changing, with recent progress on ahead-of-time compilation of Java without requiring an obscure proprietary JVM.)
There is a 0.0% chance that this happens (in my opinion as a paramedic). If I had to guess, 0.1% is probably the rough order of magnitude of the frequency with which doctors look at the PCR at all. We give a verbal report which covers the important details. I would be hard-pressed to come up with an error that could possibly waste an hour of someone's time.
"Did that quarter-second lag kill anyone? Was there someone who wouldn’t have died if the ePCR was just a little bit faster, fast enough to be usable? ... It could have saved the person the EMTs couldn’t get to because they lose an hour a week from extra PCR overhead."
Similarly, no. There is no situation in EMS where 250ms is the difference between life and death. It's not like the tones drop for a call and we say 'gee, I wish I could go help that person, but I still have this chart to write...'.
Performance absolutely matters (ironically, I'm a software developer for an EMR at my day job), but a little lag in an ePCR isn't going to kill anyone.
That wasn't how I read it. I may be incorrect, but my interpretation was that the 250ms didn't make the software too slow to be effective but rather too slow to be user friendly.
The idea isn't that it directly killed people by being slow, but rather that it was not adopted because people didn't like using it, and on the paper alternative, mistakes were made that would have been impossible to make on the computer.
The only place where PCR errors matter is on the witness stand... The PCR is certainly important, but any critical information is conveyed in a number of redundant ways, and I would be very surprised if there's been a single incidence of a PCR error resulting in a patient's death.
PCRs are not treated with a great deal of trust, nor should they be (I say this as someone who has written thousands of them). We're dealing with incomplete and conflicting information, patients that actively lie (or, more charitably, forget) about their medical history, and in the case of critically sick patients (the ones that are at risk of dying in the first place) I am more focused on the acute management of their condition. No one expects a PCR to be accurate in every detail. That doesn't mean it's useless, but it also means the system expects there to be errors and has redundancies in place to account for that.
The author made such a good praise of performance that I was half ready to learn ASM to code all my apps ... then he admitted he was using this Electron app because it was more featurefull than the native one ...
- animations under 10 ms - reacting to user interactions under 50ms