The problems start when you try to manipulate serialised data, which is not safe to do this in the general case. You should instead construct a proper representation of what you desire, and then serialise that, depending on the serialiser to take care of all of this sort of stuff. This approach has always been fairly popular in compiled languages and languages that like types, but dynamic languages have historically significantly preferred to manipulate strings, I suspect because they don’t have good ergonomics on the other approach, and it’s probably slower in interpreted languages—you’ll note that React felt the need to extend JavaScript to make its approach acceptable to people.
Most JavaScript stuff that supports server-side rendering now is working in this way, crafting a DOM tree and then serialising that. Svelte is a notable exception in that it takes a declarative DOM tree and essentially serialises what it can at compile time, thereby still retaining the required safety guarantees.
There are definitely downsides to strict adherence to the model of crafting a data structure and then serialising it; most significantly, you can’t start streaming a response until you’re done. The solution for this is to use an append-only data structure (or possibly one that allows you to “commit” the document up to a given point, while still allowing mutations in anything that occurs later in the document); thus serialisation can begin before you finish writing the document.
You know the old favourite about parsing HTML with regular expressions? <https://stackoverflow.com/questions/1732348/regex-match-open...> (If not, enjoy!) This is the thing people need to understand and realise in the general case: serialised data should be treated as opaque, and only interacted with after real parsing and before real serialisation.
HTTP headers aren’t strings; "Date: Tue, 15 Nov 1994 08:12:31 GMT" is a serialised HTTP header, representing the actual header that’s more like {Date, 1994-11-15T08:12:31Z}. And that latter is the form you should interact with it in.
HTML isn’t strings; "<p>Hello, world!</p>" is the serialised form of a paragraph element containing a text node with data “Hello, world!”. And that’s the form you should interact with it in.
Yes, I am presenting a strongly-opinionated position that lacks any shade of pragmatism. Yes, my website is generated with templates that manipulate serialised HTML. Eventually I’ll replace it with something more sound.
One last note: at the start I said valid HTML, because it’s not enough to just serialise an arbitrary HTML DOM tree, as you can easily craft invalid HTML DOM trees, like nesting hyperlinks. In most regards, the XML syntax of HTML (still a thing) is actually a safer target to serialise to because then you don’t even need to validate your tree to be confident it won’t get mangled by the serialise/parse round-trip.