It can. How would you like them to move into the future? It's not as if there's been no recent evolution in terminal emulators.
Notice that evolving into the future and using old ideas are not incompatible. Modern cryptography still uses Euclid's algorithm from more than two millennia ago.
For starters: stop using in-band signaling.
In-band signaling — or rather, in-band markup — is a big part of what makes a "terminal" (i.e. a TTY/PTY device) semantically a "terminal": a hybrid character-grid / event-log that works both as a sink for streamed-in text, and as a "plotter" for making line-printer art. A device that both streams out as, effectively, an input-event log (where this log can be captured for replay, as with script(1) or most logging systems); but which also maintains a notion of being a ring-buffer "containing" a rectangularly-bounded volume of text (and control-events), such that clients just connecting onto it can begin streaming just from the beginning of that buffer, to end up with one complete "image" of the latest PTY state, minus any scrollback, without needing previous history.
All that is kind of predicated on control-characters being embedded in the text and "following" the text around, such that taking a slice of the text (as the PTY's ring-buffer does every time a line is expunged) will preserve the corresponding slice of events.
What would reading back the contents of a PTY device look like in a world with out-of-band TTY signalling? What would flow over a serial port? Would it even be a character-stream device, or are you imagining a TTY/PTY as operating in something more akin to a structured datagram event-stream mode, such that you'd use sendmsg(2) and recv(2) on it rather than write(2) and read(2)?
I mean, it's not an impossible dream; but that really is a "start over with a whole separate ecosystem that no existing software works with until made compatible" kind of change. Effectively it'd be a separate thing from TTYs, that just happens to have similar functionality. But it wouldn't support any existing software, or any existing hardware, except by virtualization (i.e. running a PTY emulator process inside your modern OoB-signalling terminal emulator.) Kind of like what Windows has been going through to replace its own command-line.
---
Personally, I'd prefer to keep in-band signalling (in the "in-band markup" sense above, not the "you have to recognize conventional escape-code sequences heuristically to even know they're not regular text" sense.)
But I'd rather just make the in-band signalling structured — i.e. to make TTYs into a data-stream containing a variable-length self-synchronizing bit-encoding with clear prefix separation for control- and data- packets.
Y'know, like UTF-8 is for text.
...or, well, speaking of Unicode: we could just use Unicode for this, reserving another block† of control characters to go with the 30-odd ones that sit at the beginning of the BMP. Then "is this is a control codepoint, and if so, what does it mean" could just be answered by consulting a Unicode table. (In such a setup, CSI command parameterization would be accomplished with zero-width joiners, variant selectors, and other things. Just picture control-characters as invisible emoji — specifically like the flag emoji that are formed by spelling out country-codes in a sort of "flags meta-alphabet"; or like that family emoji [https://emojipedia.org/family/] with the combinatoric variants.)
† Why not use the Private Use Area? Because this would be an explicitly inter-compatible signalling standard, not a proprietary usage. It's not text, but it is a standard signal within a text document. Just like emoji — or like the existing control codepoints in Unicode.
I have some frustrations with terminals, which I have always interpreted as being caused by their adherence to some old standards. After reading this comment, I think I can see how it's more the terminal paradigm itself that is responsible for some of these things.
Now, in the future I will be able to look at these frustrations in a new light, and hopefully understand a little better why it makes sense to continue using the terminal paradigm despite them.
This is something I've struggled to understand for a long time, and your comment, obvious though it may seem to some, is one of the first helpful answers I've seen.
Keep in mind that I do not deeply understand the stack of standards and implementations that comprise a terminal. That said, my assumption is that this complaint is an unfortunate, unavoidable byproduct of that stack.
While introducing newcomers to terminal usage, there are simple actions from the gui world that do not work at all, or even seem to break the terminal display. It's extremely confusing for them, they get over it eventually, and come to accept that the terminal just can't do certain things. This is what I mean by being unable to move into the future.
As a point of comparison, consider how much of a mess it is to deal with spaces in file names, in bash. To a newcomer, this seems like a crazy hassle, some antiquated nonsense. It's hard for them to understand how experienced Linux users can stand to deal with it. It's annoying, but it's a direct consequence of the choice to use the space character as the delimiter between tokens in bash. This choice is actually super convenient most of the time, because that is the best use of the space character/key in a terse shell language.
What is the analogous rationale behind the shortcomings of the terminal paradigm? What would we need to give up, in order to make the terminal interface a little bit more modern?
For example, you could get a terminal that allows selecting text with keyboard but what happens when a user inevitably wants to or accidentally selects text from the non-input part of the terminal? Should the input part of the terminal and the display part of the terminal be treated differently? In Firefox as I type this I get a nice big input box where I can do multi-line paragraphs but if I start clicking and dragging to select the text in the input box it'll never let me select text from your comment, likewise in reverse, but on a terminal this would be undesirable behaviour for the common pattern of selecting text to copy output for documenting or troubleshooting. Some terminals have plugins or scripts to allow such selection of text but not for input purposes, if you decide to only allow text selection on the input field as a means of improving text input do you just get two such methods of being able to select text? The sane approach might be to allow selecting all text but you'll end up with paper cuts where a user doesn't care about what is going on before the $ but might end up in situations where text is selected far before the $ because of a reverse search gone wrong, etc.
As an aside, terminals/bash do have some form of text-editor like functionality, readline is the usual library involved (man bash, /^readline) and has reasonable support for moving cursor around words, move cursor to character search, deleting/yanking/pasting words, etc. It's not the best text editing interface but for dealing with a single command line it's usually sufficient. There's even a vi-like editing mode (set -o vi) built into bash if you feel like you need a modal editor for a single line but it seems even less intuitive and harder to grok.
https://readline.kablamo.org/emacs.html