To expand on your point, one of my favourite examples of how slow a lot of software has become: A few years back I tested load time for a default Ubuntu emacs install
in a console against "booting" the Linux-hosted version of AROS (an AmigaOS reimplementation) with a customized startup-sequence (AmigaOS boot script) that started FrexxEd (an AmigaOS editor co-written by the author of Curl).
AROS goes through the full AmigaOS-style boot, registering devices, and everything.
It handily beat my Emacs startup.
Now, you can make Emacs start quicker, and FrexxEd is not by default as capable (but it's scriptable with a C-like scripting language), but there's no wonder people feel software has gotten slower, because so often the defaults assume we're prepared to wait.
E.g. one of my pet peeves with typical Emacs installations: Try misconfiguring your DNS and watch it hang until the DNS lookups fail.... It's not an inherent flaw in Emacs; you can certainly prevent it from happening, but so many systems have Emacs installations where you face a long wait. Normally of course this is not a big issue, but those setups still have a DNS lookup in the critical path for startup that adds yet one more little delay.
All of these things add up very quickly. In the cases where these things are an issue in older software, they tend to either need to be explicitly enabled, or there is concurrency.
I submitted some patches to AROS years ago, to implement scrollback buffers and cut and paste in the terminal, and one of the things it really brought back is how cautious AmigaOS was in making everything painstakingly concurrent all over the place. At the cost of throughput but cutting apparent responsiveness.
E.g. when you cut and paste from a console window on AmigaOS, data about the copied region gets passed to a daemon that will write it to a clips: device. It gets passed to a separate daemon because the clips: device that the clibboards are stored to, like everything else in AmigaOS can be "assigned" to another location. By default it is stored in T: (temporary). By default T: points to a RAM disk. However, clips: or T: could very well have been reassigned to MyClipboardFloppy:. In which case copying would prompt you to insert the floppy labeled MyClipboardFloppy: after which writing the copied section would take way too long. (That actual write to the floppy would happen in yet another task)
So copying as a high level system service is handled in a separate task (thread).
Everywhere throughout the system everything that could ever potentially be slow, and on a 7.16MHz 68k machine with floppies and limited RAM that was a lot of things, would be done in separate tasks so that at least the user could just get on with other things in the meantime. "Other things" then as a consequence often meant "just keep using the current application" because the concurrency would mean a lot of these things would lazily happen in the background.
For e.g. a typical shell session, just pressing a key will involve half a dozen tasks or so, e.g. gradually "cooking" the input from a raw keyboard event, to an event for a specific window, to a event to a specific console device attached to a window, to an event to a high level "console handler" that handles complex events such as e.g. auto-complete, to the shell itself. It is inefficient in terms of throughput, but because the system will preempt high priority tasks (including user input related ones), and react with low latency and offload less latency critical tasks to other tasks, the system feels fast.
A key to this was that on a system that slow, a lot of the things that could be slow, would regularly be slow for the developers and would get fixed so that you'd opt in to the slow behaviour. E.g. I doubt most people using Emacs are ever aware if their installation is slowed down by DNS lookups on startup for example, because most of the time it's fast enough that you won't really notice that one extra little papercut.