That people produce HTML with string templates is telling us something
utcc.utoronto.ca
utcc.utoronto.ca
excellent, and in hilarious contrast with the responses in this thread...
Digital desire paths? https://en.wikipedia.org/wiki/Desire_path
I was out hiking a couple of years ago on a very steep trail with lots of signs telling people to stay on the trail because they were trying to regrow the forest in the surrounding area to prevent land slides... and what did people do? They cut through it anyway. No wonder they had to shutdown entire portions of the trail.
Sometimes things have to be done a certain way for other reasons. The most convenient technical solution is not always the right solution either.
But my comment wasn't about trail management, I'm recounting an anecdote from 25 years ago. The point was to check your assumptions against reality, and adjust accordingly.
Is that an assumption, or what the ranger said? No need to make up requirements that don't exist.
The playwright is not wrong because the play could not be understood by a man who doesn’t understand the language the play is written in. There is an expected background that allows the play to be more than a blast of noises designed to only interact with your basic senses (which would also assume the audience has those senses in the first place), and the audience can be wrong for not meeting that expectation.
Of course, the intended audience may have no relationship to the audience he will get — in this case the playwright is unreasonable, though the play may still be correct.
I would say both the audience and the playwright operate in a symbiosis; they serve one another, and they both have responsibilities in the matter.
The point of having a market of plays, games, and restaurants is that we can match producers and consumers with each other. People are going to watch movies they don’t like, eat food they hate, and watch plays that they think are boring. That doesn’t mean that we have to assign responsibility (or blame) to anybody for it! Not everybody has to like your play.
For example you talk about "The playright" and then refer to "the group he serves" which causes one to imagine the archetypal playwright as a man. Again, the person who doesn't understand the language is "a man who doesn’t understand."
I know this style of writing was the norm in the past but I found it quite jarring to see it today and honestly I had to read it again to catch your point. On the re-read I caught that only the "random person off the street" was a person and not a man.
Anyway, hope it's ok to call it out, don't mean to come across as unfriendly.
You may enjoy reading R. G. Collingwood’s The Principles of Art. Here is a relevant excerpt:
“ Next, with regard to the arts of performance, where one man designs a work of art and another, or a group of others, executes it. Ruskin (who was not always wrong) insisted long ago that in the special case of architecture the best work demanded a genuine collaboration between designer and executants: not a relation in which the workmen simply carried out orders, but one in which they had a share in the work of designing. Ruskin did not succeed in his project of reviving English architecture, because he only saw his own idea dimly and could not think out its implications, which was better done afterwards by William Morris; but the idea he partly grasped is one application of the idea I shall try to state.
In these arts (I am especially thinking of us and drama) we must get rid, to put it briefly, of the stage-direction as developed by Mr. Bernard Shaw. When we see a play swathed and larded with these excrescences, we must rub our eyes and ask: ‘What is this? Is the author, by his own confession, so bad a writer that he cannot make his intention clear to his producer and cast without composing a commentary on his play that makes it look like an edition for use in schools? Or is it that producers and actors, when this queer old stuff was written, were such idiots that they could not put a play on unless they were told with this intolerable deal of verbiage exactly how to do it? The author’s evident anxiety to show what a sharp fellow he was makes the first alternative perhaps the more probable; but really there is no need for us to choose. Whether it was the author or the company that was chiefly to blame, we can see that such stuff (clever though the dialogue is, in its way) must have been written at a time when dramatic art in England was at its lowest ebb.’
I am only using Mr. Shaw as an example of a general tendency. The same tendency is to be seen at work in most plays of the later nineteenth century; and it is just as conspicuous in music. Compare any musical score of the late nineteenth century with any of the eighteenth (not, of course, a nineteenth-century edition), and see how it is sprinkled with expression-marks, as if the composer assumed either that he had expressed himself too obscurely for any executant to make sense of the music, or that the executants for whom he writes were half-witted. I do not say that every stage-direction in the book of a play, or every expression-mark in a musical score, is a mark of incompetence either in the author or in the performer. I dare say a certain number of them are necessary. But I do say that the attempt to make a text fool-proof by multiplying them indicates a distrust of his performers on the part of the author which must somehow be got rid of if these arts are to flourish again as they have flourished in the past. This cannot be done at a blow. It can only be done at all if we fix our eyes on the kind of result we want to achieve, and work deliberately towards it.
We must face the fact that every performer is of necessity a co-author, and develop its implications. We must have authors who are willing to admit their performers into their counsels: authors who will re-write in the theatre or concert-room as rehearsals proceed, keeping their text fluid while the producer and the actors, or conductor and orchestra, help to shape it for performance; authors who understand the business of performance so well that the text they finally produce is intelligible without stage-directions or expression-marks. We must have performers (including producers and conductors, but including also the humblest members of cast and orchestra) who take an intelligent and instructed interest in the problems of authorship, and are consequently deserving of their author’s confidence and entitled to have their say as partners in the collaboration. These two results can probably be best obtained by establishing a more or less permanent connexion between certain authors and certain groups of performers. In the theatre, a few partnerships of this kind are already in existence, and promise a future for the drama that must yield better work on both sides than was possible in the bad old days (not yet, unfortunately, at an end) when a play was hawked from manager to manager until at last, perhaps with a bribe of cash, it was accepted for performance. But the drama or music which these partnerships will produce must in certain ways be a new kind of art; and we must also, therefore, have audiences trained to accept and demand it; audiences which do not ask for the slick shop-finish of a ready-made article fed to them through a theatrical or orchestral machine, but are able to appreciate and enjoy the more vivid and sensitive quality of a performance in which the company or the orchestra are performing what they themselves have helped to compose. Such a performance will never be so amusing as the standard West-end play or the ordinary symphony concert to an after-dinner audience of the overfed rich. The audience to which it appeals must be one in search not of amusement, but of art.
This brings me to the third point at which reform is necessary: the relation between the artist, or rather the collaborative unit of artist and performers, and the audience. To deal first with the arts of performance, what is here required is that the audience should feel itself (and not only feel itself, but actually and effectively become) a partner in the work of artistic creation. In England at the present time this is recognized as a principle by Mr. Rupert Doone and his colleagues of the Group Theatre. But it is not enough merely to recognize it as a principle; and how to carry out the principle in detail is a difficult question. Mr. Doone assures his audience that they are participants and not mere spectators, and asks them to behave accordingly; but the audience are apt to be a little puzzled as to what they are expected to do. What is needed is to create small and more or less stable audiences, not like those which attend a repertory theatre or a series of subscription concerts (for it is one thing to dine frequently at a certain restaurant, and quite another to be welcomed in the kitchen), but more like that of a theatrical or musical club, where the audience are in the habit of attending not only performances but rehearsals, make friends with authors and performers, know about the aims and projects of the group to which they all alike belong, and feel themselves responsible, each in his degree, for its successes and failures. Obviously this can be done only if all parties entirely get rid of the idea that the art in question is a kind of amusement, and see it as a serious job, art proper.
With the arts of publication (notably painting and non-dramatic writing) the principle is the same, but the situation is more difficult. The promiscuous dissemination of books and paintings by the press and public exhibition creates a shapeless and anonymous audience whose collaborative function it is impossible to exploit. Out of this formless dust of humanity a painter or writer can, indeed, crystallize an audience of his own; but only when he has already made his mark. Consequently, it is no help to him just when he most needs its help, while his artistic powers are still immature. The specialist writer on learned subjects is in a happier position; he has from the first an audience of fellow specialists, whom he addresses, and from whom an echo reaches him; and only one who has written in this way for a narrow, specialized public can realize how that echo helps him with his work and gives him the confidence that comes from knowing what his public expects and thinks of him. But the non-specialist writer and the painter of pictures are to-day in a position where their public is as good as useless to them. The evils are obvious; such men are driven into a choice between commercialism and barren eccentricity. There are critics and reviewers, literary and artistic journals, which ought to be at work mitigating these evils and establishing contact between a writer or painter and the kind of audience he needs. But in practice they seldom seem to understand that this is, or should be, their function, and either they do nothing at all or they do more harm than good. The fact is becoming notorious; publishers are ceasing to be interested in the reviews their books get, and beginning to decide that they make no difference to the sales.
Unless this situation can be altered, there is a real likelihood that painting and non-dramatic literature, as forms of art, may cease to exist, their heritage being absorbed partly into various kinds of entertainment, advertisement, instruction, or propaganda, partly into other forms of art like drama and architecture, where the artist is in direct contact with his audience. Indeed, this has begun to happen already. The novel, once an important literary form, has all but disappeared, except as an amusement for the seine-literate. The easel-picture is still being painted, but only for exhibition purposes. It is not being sold. Those who can remember the interiors of the eighteen-nineties, with their densely picture-hung walls, realize that the painters of to-day are working to supply a market that no longer exists. They are not likely to go on doing it for long.‘
- Storing credentials in plaintext
- Not validating input
- SQL injection
- etc
All more convenient than doing it the right way.
Using string templating makes the DX better without compromising UX, since users just see the rendered output. Implementing bad/nonexistent web security also makes the DX easier since there's simply fewer features to implement, but this obviously has negative consequences on UX when folks have their accounts/credentials easily stolen.
If security is a goal, there is a difference between doing anything and actually having a secure system. There is also such a thing as secure enough.
Hacker news thread talks about how to do HTML. Guy writes article refuting the thread.
But it happens backwards
TENET!
https://ontology2.com/the-book/html5-the-official-document-l...
One big problem is that systems like this are orders of magnitude slower than text-based template systems. Another one is a problem with namespaces. If you mash together two arbitrary documents they could have identical id attributes (forbidden) or identical classes (talk about wires getting crossed.). You ought to be able to transcoded an arbitrary HTML document into another one but you’d need to rewrite the CSS to eliminate conflicts in some cases, which I think is possible but is rarely done.
Many of the responses in this thread are highlighting that the analysis here is useful but oversimplified: if people are "doing it wrong" it's an opportunity to reflect on our approach to UX/DX, but accepting populism for its own sake is throwing out the baby with the bathwater.
It's also been pointed out by numerous commenters that the central qualifier - that everyone uses string templates - isn't even true. The most popular front-end systems in modern stacks are structured html. It's in common use today.
Lisp macros make adding HTML syntax easy. You won’t find anyone using string templates in that language because a handful of macros means you can just program like it’s just lisp.
Strings such primarily because they don’t establish regularity. If you don’t understand everything fully and follow their patterns exactly, it’s easy to accidentally lose your pseudo-macro hygiene and output garbage.
JSX was a revelation simply because it was a “macro” (DSL) that ECMA had already designed an entire spec around (E4X) and thoroughly baked into the language. Like with Lisp, you could just use your normal coding patterns to interface.
A custom HTML macro baked into the syntax makes sense in a language where almost everyone using it is going to need HTML. It would make far less sense to dedicate all that syntax space in a more general purpose language.
And in JS, even with all that design time spent on E4X, you are still back to doing string interpolation the second you step away from that specific syntax (or you’re forced to express everything as HTML even if it’s not a good fit).
The world would be a better place if JS had been scheme and people had been forced to learn a lisp.
They were merged with another company. Reddit guys knew both Lisp and Python. Other company coders only knew Python.
Since then, they've had massive issues scaling Python in general and their ORM dependence in particular.
It migrated to python very early in the game (via Aaron Swartz (rip)), and it's my understanding that it was that move that allowed them to scale it.
...how does this change conditional probability? If of those people who use Lisp for web development, nobody uses strings, it's unrelated to how many people use Lisp for web development.
I kid, I kid. Lisp is great.
None of the modern web would be around though because we’d still be waiting for a sufficiently advanced compiler.
Chez scheme is probably about as fast as JS JITs and with only a fraction of the time spent creating it. If you restrict continuations, you can get even better performance. On the flip side, new JS features like BigInt would have existed from the start (along with generators, rest/spread, typed arrays, let/const, etc). Features like threads that don't exist probably would exist.
On the better side, all the terrible things people complain about like hoisting, `with`, type coercion, weird prototypal inheritance, bad Java-based dates, etc simply wouldn't have happened because Scheme already specced out most of the relevant things.
HTML would have likely disappeared over time because innerHTML and string parsing would be radically less efficient than just using the macros.
We wouldn't have 10 different versions of JS because most of the new stuff either would have been baked into the first version or could be easily accomplished with macros. Major versions would be little things like adding optional type hints or
CSS wouldn't exist because you'd create sets of styles with lisp lists then pass them in. It would be a better version of CSS in JS, but done 25 years ago.
JSON wouldn't have been discovered because lists do all the things better. Likewise, there wouldn't be a need for the "lost decade" of XML development because those same scheme macros would do that job and transformer macros are far easier and better to write than XSLT.
I would be fairly surprised to hear someone complain about with-statements. My impression is that most folks don't even know it exists, and I'd be very shocked to see it actually being used in the wild.
Huh? This doesn’t make any sense. I don’t think people have done a lot of Scheme JITs, but Scheme has some pretty damn impressive compilers—Chez[1] first and foremost. Certainly ones with better codegen than pre-V8 JavaScript ones. Scheme (the standard fragment) is less dynamic than JavaScript, not more (which has been used as an argument against that fragment by more OG-Lisp-inclined people).
(The one potenial problem I can name is tail calls—IME LuaJIT is much, much worse at compiling Lua’s Scheme-like proper tail calls than it is at the same code expressed as a loop. But then the price for LuaJIT’s small size is that it’s a bit picky at which code it’s willing to compile well. Production JS engines probably handle that better, if at a cost of a couple of orders of magnitude more code.)
The biggest language difference I can think of is the guarantees about numeric types. JS can easily compile to native float or integer operations, when it's hard to do that in scheme.
What other scheme features do you have in mind that make it harder to compile? ( Maybe ignoring call/cc)
HTML templating is even popular in Lisps. See djula and selmer.
For Clojure (re selmer), hiccup is a way more popular way of doing HTML (probably even the de-facto standard), and it's not doing string templates.
More generally, syntactic sugar matters a lot. E.g. Python's list comprehensions are merely syntactic sugar for filter + map, but they make functional code so much more readable.
user> (require '[rum.core :as rum]
'[hiccup.core :as hiccup])
nil
user> (rum/render-static-markup [:p "<script>alert('you have been pwned')</script>"])
"<p><script>alert('you have been pwned')</script></p>"
user> (hiccup/html [:p "<script>alert('you have been pwned')</script>"])
"<p><script>alert('you have been pwned')</script></p>"
(Note that Rum is also a React wrapper, but you don't have to use that part of it; you can simply use it for static rendering of HTML.)That is to say Clojure/lisp programmers are likely more familiar working with deeply nested lists of data.
You could write a very similar library in Python if you wanted using lists, I'm just doubtful anyone would use it because it's not "Pythonic".
More readable than filter/map/reduce _in Python_, because Guido dislikes them and wanted them out of the language (and succeeded in moving reduce to functools). Python's excuse for lambdas makes this so much worse.
Compared to functional languages, though? I'll take filter/map/fold/reduce from Haskell or Clojure versus list comprehensions any day.
Imagine if JSX produced a standard data structure that could then be
- Rendered by React or another framework
- Serialized to a string
- Deeply inspected/compared
etc. And framework integration only happened when you make the actual framework call, not globally at build-time
Your custom jsx function can return data in reusable format, that is passed to different functions for difference use case.
ts-liveview is using this approach to use jsx for both server-side rendering and compact over-the-wire updates
Compared to what? Languages with actual filter & map generics seem far more legible to me than list comprehensions...
I was doing imperative programming for years, and more functional now.
You've got for loops, list comprehensions, newer dict comprehensions, map/filter, generators, etc. in Python.
Python programmers tend to prefer for loops and comprehensions. List comprehensions, especially, dict comprehensions being much newer.
Living mostly in the pre-Java 8 Java world, I used for loops as much as the next guy, like the rest of the Java world. But I really liked list comprehensions. And then Java got streams and map and filter, etc., and wow, I loved it. Even most Java programmers started embracing the new "functional" paradigms.
Now that I've been both (and its more powerful cousins in Clojure), I'd say map/filter are "more readable". It depends on the implementation too: try merging two dictionaries in Python, ugh [1] [2]. Awkward any way you slice it. Watch out what version of Python 3.x you have! Watch the order of arguments, and some variations modify in-place! You're also never sure if your result will be a list or a dict.
Try map'ing a dict in Python: don't, it's not worth it [3]. These things are actually pretty easy, even in Java, let alone Clojure.
For the record, Clojure has map/filter and list comprehensions (called for comprehensions, and they're even more powerful than Python because they can terminate early [0]).
So, yes, in Python, map/filter can often be a pain. They were afterthoughts, and less readable than the alternatives. For languages with powerful map/filter and friends, these functional equivalents are much more flexible and readable. And it appears that most people that are familiar with both prefer the functional variety.
[0]: https://clojuredocs.org/clojure.core/for
[1]: https://stackoverflow.com/questions/38987/how-do-i-merge-two...
[2]: https://gist.github.com/SZanlongo/bc4baa90d3795db7c6ed7e8d41...
[3]: https://stackoverflow.com/questions/23862406/filter-items-in...
> No one has structured HTML creation that's as easy as string templates.
Well, maybe it's about the "easy" part, maybe the author considers JSX string templates. I assume the former, because the latter would be absurd.
The advantages of so-called frontend "frameworks" such as React et al is 50% this point. Also a reason for Lit's subjectively slow adoption.
Can't say I agree. Maybe it's because I worked with simple functional programming primitives like filter and map before I ever really worked with Python, but I find Python list comprehensions weirder and harder to read than things like Clojure's thread macros, Ruby blocks, or even just chaining functional method calls together using the normal dot syntax in Java or Scala.
They do look better than the equivalents shown in the official Python documentation¹:
squares = list(map(lambda x: x**2, range(10)))
but that is exceptionally hideous and not what I'd consider a normal way to write functional code for dealing with collections or streams.Seems like a Python workaround for a Python problem to me. ¯ \ _ ( ツ ) _ / ¯
--
for example, imagine if your example could be written as:
range(10).map(lambda x: x*2).list()
but as a list comprehension it's slightly shorter, but autocomplete isn't as good:
[x*2 for x in range(10)]
Also some historical context, Python's list comprehension are based on Haskell's
I do also find the comprehensions more readable when one is actually using all of their components, i.e., when the function in the map behind the sugar would not be the identity function. But occasionally, when all the author really wanted to do is filter, you'll see stuff like
[x for x in ... if ...]
and that always strikes me as a bit weird and unfortunate.> Also some historical context, Python's list comprehension are based on Haskell's
I didn't know that! I do kinda like Python's preference for English keywords over arrows here, even though the Haskell syntax is more concise or whatever.
[x^2 | x <- [1..10]]
which looks about as legible as the Python version.But what people actually need are grammars.
The exact same reason why parsing HTML with a regex unleashes Zalgo is why generating HTML with string templates is bad. Because both treat HTML as a string, not a grammatically restricted language.
I don’t actually disagree with you for the most part, but I feel that an important caveat has gone unacknowledged.
Grammar formalisms have the same weakness compared to dealing with raw strings as sound static type systems do compared with dynamic typing: there are small, mostly isolated islands of feasibility in a sea of intractable (often undecidable) generality, and if your problem doesn’t fit inside those borders things start to get nasty (cf how even GCC’s handwritten rec-descent parser didn’t get its lexer hack interactions correct in all cases[1]).
I still agree that we spend criminally little time on syntax. Starting with the simplest cases: with how much time is spent in school on “order of operations” you’d think we could take a moment to draw[2] a damn syntax tree! But nooo. There are in fact working mathematicians who don’t know what that is. (On the other hand, there are mathematicians who can explain that, in a sense, the core of Gödel’s incompleteness is not being able to reason about arithmetic—it’s being able to reason about CONS[3], which arithmetic happens to be able to very awkwardly do.)
[1] https://gcc.gnu.org/bugzilla/show_bug.cgi?id=67784
[2] https://mlochbaum.github.io/BQN/tutorial/expression.html
Generating JSON data using string interpolation or templating is clearly wildly insane, right? You don’t do it.
Maybe for some config file generation scenarios you might just run a template JSON file through a token substitution or env var interpolation or something. But you’d feel bad about it, because it’s so easy to NOT do it that way. And even then you’re not interpolating in JSON fragments like ‘“age”: 25’ - you’d have the decency to only interpolate in values like ‘25’.
In the node ecosystem it’s so easy to switch from a .json file to a .js file, too, if you want to build the json dynamically.
For some reason people feel more willing to attempt it with YAML. And then regret it when they realize how significant indenting has screwed them.
And then with HTML people just give up and go ‘yup, it’s all text, even the angle brackets’
I'm sorry to report that I've seen a lot of JSON generated by string concatenation and templating, in different projects.
Often using 'printf' or 'echo' in various languages. Sometimes using whatever's used for HTML string templating if the JSON is embedded in HTML or served as a resource similar to HTML.
Yes, its horrible and breaks if fed variable values that have characters like quotation marks in. People do it anyway.
Even in languages that have perfectly good data structures and JSON libraries.
I've seen a fair amount of parsing values out of JSON using regexes too, assuming specific formatting of the supplied JSON.
This file generates a feed of events (rehearsals for my high school play) to be rendered by the FullCalendar JS plugin. FullCalendar required a particular data schema that didn't match the format of my MySQL table, which meant I couldn't just json_encode() the MySQL results. I guess I just didn't conceptualize that I could create a new object that matched the FullCalendar format, and then call json_encode(). So, I generated JSON with strings.
Honestly it's a toss-up whether the JSON generation is the worst thing about this file. It looks like I also made a separate database query for every single row to get the username, because I apparently didn't know how to do joins. Could probably spend an hour listing some of the other little nuggets of awful in there. But hey, it got the job done! :)
But I'd never try to implement my own parser or output deeply nested JSON.
It knows when to do escaping and how. It also can detect, though dynamically, when fragments have been combined into an illegal sequence which would be rejected by the full grammar. It can not however guarantee that the result will parse only that it can not detect that it would fail.
The ugly quasiquoting seems unfortunate (I’ve a half-serious suspicion the reason Template Haskell never got popular is that it looks so bad), and the GLR sledgehammer precludes ever having a lightweight implementation, but otherwise it seems like a interesting entry in the extensible languages story.
5 - 2(3 - 1) - 5 = ?
Also, respectfully, it doesn’t matter. Not having learned maths in English, I don’t know the mnemonic, I don’t care to know it, and I find even the concept of it completely asinine. (For eighteenth-century mathematicians, addition and subtraction bound tighter than multiplication and division, and they could calculate perfectly fine.) You can look up the precedence table if you need to—as long as you need to understood the idea of precedence (and not order of operations, for goodness’ sake). You won’t then be able to calculate fluently, but fluency is a different problem with a tedious and time-consuming solution, and given the time crunch I’d rather talk about some actual Maths as She Is Spoke instead.
6 / 2 (1 + 2) = ?
I approach the problem the same as I would 6 / 2(x + y). When the multiplication is missing 2(x + y) is a single term. The implicit multiplication is part of the parenthesis and reduces the problem to 6 / 6. People who argue that you have to strictly use PEDMAS left-to-right will divide 6 / 2 first and get 9.
Neither way is wrong as long as you can explain the process but everyone wants to argue and have there be a single answer.
My older relatives are the ones I see repost these inane order-of-operation tests, getting the answer consistently wrong.
Moreover, important parts of HTML processing would be significantly more brittle and complicated and less powerful with objects: "escape some completely arbitrary text to valid PCDATA or a CDATA section, whatever is shorter" is strictly more general, robust and principled than "render a Street Address to a fragment that isn't supposed to contain markup".
HTML is a grammatically restricted subset of text.
I can take arbitrary text and embed it in HTML by escaping characters within it. That produces a grammatical fragment of HTML that represents the arbitrary text, but it is not the text.
As an example sentence, take the following:
"The French equivalent for the English "Good Evening!" is "Bonsoir!", whereas Italians might say "Buonasera!" to one another for similar effect."
There are four languages in that sentence, two of which are English. You may need three editors to deal with them, or you can flatten the sentence and simply edit everything assuming you knew all three.
This is not even true of natural language, which has a vocal representation that is at least as important as the written representation.
Though I agree that representing HTML as objects is a poor substitute.
The problem with HTML is that its syntax trees are relatively unpleasant to use.
For anyone who hasn't seen it yet, top answer from https://stackoverflow.com/questions/1732348/regex-match-open...
Besides not "being proper" or whatever your argument boils down to, people (arguably) are doing useful things by just manipulating strings.
I'd argue most of the web is probably built with just strings and duct tape holding all the pieces together.
It would be better if programming languages discouraged people away from those mistakes and prodded them towards the pit of success by making string concatenation harder and providing better tools for constructing grammatically sound structures.
Python's MarkupSafe (used in jinja) and go's html/template are good examples.
But if we were talking about outputting CSV data you wouldn’t be able to say
In lisp the templating systems are S-expression based instead of string interpolation based, which at least models the tabular structure of CSV documents
Because hierarchies don’t model tables especially well.
So the suitability of lisp for outputting HTML feels slightly coincidental.
I mean it basically worked with JSON too.
This is something that literally every single framework, front-end and back-end has had to deal with in some way since the 90s. From chucking ugly <?php tags to more elegant solutions like curly braces.
Having one standard in the spec would standardize something that is currently done a million different ways.
https://developer.mozilla.org/en-US/docs/Web/HTML/Element/te...
This gets complex fast, I cannot imagine having something like this as a slow-moving spec.
Also, XSLT is horrible.
That said XML is like violence, if it’s not solving you problems you need to use more.
XQuery can do everything XSLT can but with different syntax. XQuery 3 can also handle JSON. It's a clean way to generate well-formed XHTML.
That’s not to mention years of having to cook up hacks to deal with inconsistent browser implementations that violated the document structure you’d be trying to create.
In the early 2000s I was working on web projects. We built such data structures, we did XML/XSLT, we were very careful to make sure everything was well-formed… and we still ended up using string templates somewhere. The tools just didn’t always exist to do everything the way we wanted, so we had to work with what we had. It hurt every time we resorted to it, because we knew we’d have to clean it up someday. But sometimes you just have to do what gets you home in time for dinner.
I believe people use string interpolation to construct HTML, or just about any other language (eg. SQL, JSON, and even human languages) -- is that they see it as a nail and the hammer of string manipulation is almost universally acquired in the very first lessons of any programming language, only second to arithmetics.
People construct HTML with string because the language and environments they use doesn't have mainstream and suitable -- in terms of accessibility and efficiency-- constructs for building HTML.
This problem doesn't exists with React, Elm, and friends that has first class constructs for building HTML.
E.g. using Common Lisp as an example, this:
(html output-stream
((a href "https://example.com") "foo"))
could translate to the code: (write-string "<a href=\"https://example.com\">foo</a>" output-stream)
which just dumps a string literal to the stream.One issue in Common Lisp html generators is that if you want to use the backquote syntax for it, you're steered toward an implementation that just lets the backquote do its job of constructing the list, and then walk the list. The reason being that backquote expands in an implementation-defined way. If backquote expands to a macro syntax like Scheme quasiquote, you can suppress its evaluation and then walk it yourself to give it your own meaning:
(let ((url "https://example.com"))
(html output-stream
`((a href ,url) "foo")))
Here, html could intercept the quasiquote syntax, walk it itself and spit out code like: (let ((url "https://example.com"))
(html output-stream
(write-string "<a href=\"" output-stream)
(html-write-attr url output-stream)
(write-string "\">foo</a>" output-stream)))
Historically, an example of a Common Lisp HTML formatter which walks a constructed nested list object is Tim Bradshaw's htout library. An example of an efficient generator of write-string calls is CL-WHO.CL-WHO doesn't use backquoting for interpolation. The code template uses keywords for indicating expressions that are HTML tags. Other expressions are implicitly Lisp to be executed. Inside evaluated lisp, the htm macro switches back to HTML templating. E.g.:
(with-html-output (*http-stream*)
(:h4 "Look at the character entities generated by this example")
(loop for i from 0
for string in '("Fête" "Sørensen" "naïve" "Hühner" "Straße")
do (htm
(:p :style (conc "background-color:" (case (mod i 3)
((0) "red")
((1) "orange")
((2) "blue")))
(htm (esc string))))))
which, according to the documentation, generates code similar to: (let ((*http-stream* *http-stream*))
(progn
nil
(write-string
"<h4>Look at the character entities generated by this example</h4>"
*http-stream*)
(loop for i from 0 for string in '("Fête" "Sørensen" "naïve" "Hühner" "Straße")
do (progn
(write-string "<p style='" *http-stream*)
(princ (conc "background-color:"
(case (mod i 3)
((0) "red")
((1) "orange")
((2) "blue")))
*http-stream*)
(write-string "'>" *http-stream*)
(progn (write-string (escape-string string) *http-stream*))
(write-string "</p>" *http-stream*)))))a{ nested:tree with:{lots-of:data and:more}}
https://clojure.org/guides/learn/hashed_colls#_maps
As an outsider this strikes me as a purely stylistic argument, like people who argue bubble sort is the worst sort and should never be used. If it's sufficient to the task, who cares?
If you told, for example, suckless that their page builder code[0] should not use string wrangling they'd certainly laugh and ignore the advice.
[0] https://git.suckless.org/sites/file/build-page.c.html#l23
To someone reading (or authoring) a document, the end state/goal is having that string of text visible and formatted well enough. They don't care about DOM objects.
The main reason we use HTML is not because people enjoy DOM-traversal or parsing or abstract syntax trees. It's because a little markup in your strings can make them format nicely and make it easy to link to and embed other stuff like images, video, and audio.
String templating/interpolation is the goal.
Hand-written HTML is static and thus doesn't use user-provided strings, while HTML generated by other means automatically escapes any strings given to it.
From a security perspective, when it comes to HTML output you're concerned about injection. A HTML opening/closing tag are code - they're read by the browser and have technical meaning for the renderer. The text in between those tags is content: you generally want to render that as is. So these two parts of the string have different purposes, they're contextually different.
Injection happens when a malicious actors gets data into your content that the browser will think (& interpret) as code. The best way to avoid this on the server side is to be aware of whether you're outputting content or code at any given point in your html document.
That's possible with string templates: its called output escaping and you just wrap and variable printing in a function that escapes special characters to avoid them being interpreted as HTML code - this is simple as there's only 5: gt, lt, ampersand, single- and double-quotes.
Every modern string templating system does this by default and its ok. But it's only a start, and is pretty limited it how far it can go.
There's actually three types of injection for HTML: injection actual HTML is just one. There's also JS or CSS injection (CSP is starting to allow mitigation of some of this but no-one uses it, mainly as it's discouraged by Google & others) and attribute injection.
Output escaping only deals with the first type of HTML injection. It can be extended to try and deal with the other two, but it's very error prone if it doesn't have knowledge of whether it's printing a variable inside a script tag, inside an HTML attribute, or just in a normal tag content. It's very difficult for string templates to get that context: structured generation gets all that context for free.
Once you have all that extra context for security, it's also useful for a bunch of other stuff (debugging/analytics/dynamic server side formatting/etc.)
https://learn.microsoft.com/en-us/dotnet/standard/linq/funct...
Related: I’ll admit that I used to use XSLT for my static site generators [1] and now I’ve switched to string interpolation (with conservative escaping as the default) [2]. Edit: In my case, the goal was to simplify and reduce dependencies.
tagFunction`string text ${expression} string text`
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...https://pitsidianak.is/blog/feed.xml
NoScript on firefox is blocking it for me, but there's no javascript involved, only XSLT, HTML and CSS.
Also, non-programmers dig XSLT specifically because it's not "code".
I manage a team of reporting analysts who have to maintain data feeds from our clients to brokers. They are from business and accounting backgrounds and use XSLT for it all. They don't need to worry about build steps or any of that junk, we just store the transforms in a database and they use IntelliJ to incrementally build out the files to the brokers specs, and we can display the transforms in a UI to other staff so they can know whats going on whenever there are questions.
Hundreds of reports, couldn't manage it all without XSLT.
With vanilla JS you have to translate to JSON, then build elements in the front end.
The solution to building elements in JS is to use a framework, but compiling Javascript just never sat right with me. It's like ok, now I have an incomprehensible bundle of minified JS and I probably need to run a JS backend if I don't want to go insane.
Alternatively, you can skip all of that and just template.
I'm excited for htmx but haven't used it yet. Simplicity wins. LAMP stack was great for its simplicity. Go monoliths are a modern solution.
https://github.com/Knio/dominate is a Python lib that implements this principal.
I wrote a little wrapper class to help with it:
All the IDE tools just work: refactoring, auto-completion, usage search, debugging.
Also I get really nice exceptions and can express stuff about templates in their types.
When combined with Classes = Templates, static functions = macros and using implementation inheritance (the one permissible case haha) it covers all the usual corners of nice templating engines.
Sure the syntax and indentation is a little wonky, but for me it pays off in so many more convenient ways that it is easily worth the tradeoff.
Is this the best implementation of the concept? Probably not, but it is good.
There is no issue, until you forget to use escaping (or use the wrong one) for one variable, and someone uses that hole to inject arbitrary HTML and/or JS into your page. As long as all your escaping of interpolated variables is perfect, producing HTML with string templates is fine.
String formatting on the other hand, yeah, no good way like that in a language not designed for it.
Not sure which you and GP meant by "string templates".
That is what made JSX so neat.
addFragment : (String, IntermediateHtmlAST) -> IntermediateHtmlAST
renderHtml : IntermediateHtmlAST -> String
There is a sanitation pass that occurs either in the final conversion of the intermediate data structure to an HTML string (renderHtml), or immediately on the function call (addFragment).This is similar to how database query libraries let you build up a SQL query via an intermediate data structure and then convert that to a prepared SQL statement (most common) or do data sanitization on the input fragment (less ideal).
If it's parametrizable why bother with completely dynamic generation? That's just more code to maintain.
The main drawback I see is how large a door it can let open on security side. Not that I can’t be done with reasonable security check, but it’s far easier to inadvertently shoot oneself in the foot.
I'm a large fan of Embedded DSLs (Domain-Specific Languages) for these types of problems, as they allow using normal HTML syntax directly in the rest of your code. Combined with macros (for compile-time parsing/analysis) and interpolation (for safely templating values), DSLs can be safer, more performant, and overall more maintainable than other approaches. For languages like HTML and SQL that are standard and well-defined, this is undoubtedly a better approach in my mind.
I've been developing my own approach to Embedded DSLs as my thesis research with my own programming language Rhovas [0], which addresses concerns with syntax restrictions (e.g. string quotes), semantic analysis (AST/IR transformations), and most importantly tooling. Happy to answer any questions about work in this area.
[0]: https://blog.willbanders.dev/articles/introducing-syntax-mac...
I think it just happens to be that the string representation is one of the most familiar and accessible formats to many people.
I don't subscribe to the idea that HTML somehow fundamentally is a string. However, even if we see it as a data structure, to many languages, data structures are not at hand, while objects with parochial APIs are. So you end up with a flurry of different libraries with their own way of doing things, instead of "this is just data, I'll use my language's generic data manipulation tools to deal with this".
JSON is conceptually simpler and there's still a lot of quirks between libraries for that already (e.g. null/undefined, integer/decimal representation, and large numbers). XML has more going on to start, and then you get all the different libraries inventing their own abstractions as you said and it picks up a lot of pitfalls. FWIW; I've messed around a lot with configuration languages and it's definitely hard to get right so I understand how this differences accumulate.
I'm also surprised that TypeScript finally managed to sway people. I would bet a large number of people here won't even remember the whole ES4/ES5 debacle. There was a time when types (and much more) were being added to ECMA Script proper through the standards process before Yahoo and others killed it.
Racket (and most other Lisps) have built-in syntax transformation features that will let you do it as your own macros, or even inlining with different parsers, without some distinct preprocessor kludge.
Here's a macro example: https://www.neilvandyke.org/racket/html-template/
Here's doing it with data: https://www.neilvandyke.org/racket/html-writing/
You could also make a Racket reader so that you could have inline HTML in its customary angle-bracket syntax.
For example, make linear types Html, HtmlOpen, HtmlBody, and HtmlClosed, with the addition rules (which could be functions or what have you)
HtmlOpen + HtmlBody = HtmlOpen
HtmlBody + HtmlBody = HtmlBody
HtmlBody + HtmlClosed = HtmlClosed
HtmlOpen + HtmlClosed = Html
and only Html has whatever it takes to be valid output.
A separate point is that string templating is simply a local maximum.
f : HtmlBody -> HtmlBody
f x = someOpenFragment + x + someClosedFragmentAs for linear types, I mean that somewhere the type system won't allow progress if you have an open type that wasn't converted to some final state. This is strictly necessary.
Also a reminder that I should read articles before commenting.
Unfortunately, it seems that byte strings as the ultimate ABI won out in this case.
But it shouldn't be surprising that people turn to string-handling for handling something where most of the inputs are strings and all of the output is.
Recently I did an experiment in python using the XIST library for generating HTML vs a Jinja2 template. Things I found:
1. The XIST approach was 10x slower, probably related to XML serialization vs string building with Jinja2
2. Although an HTML template looks a bit like a set of nested functions (like the lisp people say it is) you actually end up wanting to pass context deep down the call tree. For example. You end up adding parameters to all the parent functions or you pass in a "context" dict which makes it hard to see what is actually being consumed from the context and by what.
def make_header(user) -> html.Div: <-- user passed
return html.Div(
make_logo(),
make_user_section(user) # <-- so this can use it
)
This might not seem that different from normal programming but html-as-functions can mean you end up with a very deeply nested call stack with each function a thin wrapper around the next (Russian Doll programming) passing parameters down and down and down.3. As demonstrated by the example above, instead of a set of "what you see is what you get" html strings you end up with a lot of component/function calls which actually make it hard to correlate what is rendered and sent to the browser vs where it came from.
I still like the idea of saying button(icon="plane", text="Book ticket") vs 3 lines of html, and the idea of getting type checking on my html generation functions but this experiment put me off it.
Not sure what to make of the second point, but I'd probably try to log it in debug mode somehow, hopefully avoiding performance degradation when not used.
- pass it down manual - use Context, essentially thread local variables but React
both have tradeoffs and it really depends on what you're building your codebase.
A better way is to use html as what it really is, a dom document and with a simple glue language fill it with data, in this way you keep the html code unchanged for better maintainability and reuse.
For example a very basic way is to use a list of css selectors with the data that you want to fill eg:
div#id > span.class a = "lorem ipsum"
#product > .product-title = $product.title
Something like https://github.com/givanz/vtpl
Here's their example page for arrays, you can see how the template is completely valid HTML without having to make the engine render it: https://www.tinybutstrong.com/examples.php?e=dataarray&m=tem... (then on the upper-left you can click "Result" to see how the "tr" elements, as specified in "block=tr" in the template, are duplicated for the passed-in array)
It also has some weird gotchas (have to use className when you mean class).
document.body.className = 'class1 class2'
JSX chose to align names to the DOM spec [0]. Same for htmlFor and friends.[0] https://dom.spec.whatwg.org/#ref-for-dom-element-classname%E...
The className thing is transitory. `class` was reserved in all non-string contexts previously, but in modern JS, you can add it bare in places like object literal keys. Preact allows you to use `class`, but react is more conservative.
Doing this makes testing trivial, and you also don't need to worry about closing tags as much. As far as I can tell, the compilation is fast enough to where the overhead of translating to HTML doesn't seem to take any amount of time.
Plus I just really like Clojure :)
Whenever I had to write JSX I have a lot of weird internal "switching" going on in my head.
JSX: <div><div><div></div></div></div>
Compiles to: createElement('div', null, createElement('div', null, createElement('div', null)))
That said, the problem always came down to the tooling aspect. If I'm just building my own toy web pages, this works great. See https://taeric.github.io/Sudoku.html for a rough look. Want to integrate with content authors and take in the pages that they are making? Yeah, this is terrible for that.
I wholeheartedly disagree. Occam's razor:
Doing it the right way takes knowledge that string substitution is wrong. Many people don't have that.
Also, it's hard(er) than the simple string solution. It takes knowledge of dom APIs, or frameworks, or other things that a fresh grad wouldn't intuitively create in 5 seconds to insert a value.
Seems fair.
Get real, the whole point is some strings in HTML are HTML, and some are malicious code.
You're being adversarial to equate the two.
Similarly, structured programming makes assumptions that are violated by common idioms such as unrestricted GOTO statements jumping all across your tangled spaghetti code.
Maybe the header-footer idiom is not a good idiom.
When producing PDF document back when I was in charge of SkPDF (used by billions!), we had to generate the PDF files that way, since the serialized format is very precise.
How about "the people are right in wanting systems that automatically prevent extremely dangerous injection vulnerabilities", and "lazy developers are wrong", but "there's a tiny minority of amateur developers that use the works of lazy developers because they don't yet appreciate the dangers of string templating, but soon will"?
"One of my fundamental rules of system design is when people keep doing it wrong, the people are right and your system or idea is wrong."
I am not sure under what circumstances this is true, but for security it is definitely false. People keep forgetting to sanitize inputs, for example. Does that render input sanitization a broken idea?
If everyone forgets to sanitize input, that means our current processes are fucking broken.
I like to imagine if this excuse happened in aviation how it would sound.
“Everyone forgets to put down landing gear before landing, so does that mean landing gear is a bad idea?”
No, but it does mean the process for landing is catastrophically broken, and needs to be overhauled. If that happened twice, we’d implement strict rules around landing like checklists with standardized procedures to have a second person verifying that the process was followed and each task was accomplished, and defining a “sterile phase of flight” where talking about anything other than the landing process was disallowed.
So, input sanitization isn’t a broken idea, but the fact that it keeps going wrong clearly demonstrates that out process and culture around security are catastrophically broken.
I absolutely agree that psychological acceptability is key for building secure systems, and that the human factor must not be neglected. I also agree that process and culture around security are deeply broken. I just do not agree with the idea that if all are doing it, there must be some deep truth beneath. Yes, people habits must not be ignored, but to claim that an idea that goes against those habits is wrong, as the author claims, is inaccurate.
Well ... yes? If people are forgetting to do something that is required, then that something is explicitly needed.
Better to make that "something" implicitly added in the process no matter what the user does[1]. Or make the process break if the user "forgets"[2].
All difficult things to be sure, but easier than expecting the user to remember which of the 100 different NON-DEFAULT sanitation packages to install, configure and use, for output to HTML, SQL, JSON, Logs, and more.
If the language makes it easy to do, and there is very little blowback in terms of security[3], then users are gonna do it.
[1] Auto-sanitise strings, obviously. I dunno how you'd actually do this though.
[2] When running as a web-service, have Django/whatever-framework configure Python/whatever-language to emit warnings whenever string interpolation is used without any escape function.
[3] For HTML, the exploitation of string interpolation are few and far between; they're so rare as to be lost in the noise. Hence, users don't use it. For SQL, injection was a real problem, with the risk of getting pwned on string interpolation being close to 100% on a good day (and actually 100% on a bad day), and so users actually used the mitigations there were.
Maybe it's not possible to do any better (no magic language that can tell what a given string of bytes will eventually be used for) but that doesn't change the fact that the reason inputs don't get sanitized is because the system requires an unreliable component to be reliable.
It's also a bad example because it's not really a choice (no magic language).
The developers aren't all idiots, they are following your incomplete documentation and using the buggy and incomplete example code you provided. Or they are using the buggy example code they found online because you didn't provide any. If your API isn't useful without first taking a college course on the problem domain then that is a problem with your API and probably documentation. The whole point of an API is to encapsulate the complexity and provide the user a tool they can use without having to first learn everything about it.
Possibly. You could prevent careless use of unsanitized inputs with a type system. Fail-safe designs should be more common.
The input sanitization system could be improved by requiring sanitized inputs, reminding people that they need to sanitize inputs, automatically sanitizing inputs, etc. Clearly the current system is leading to a lot of insecurity.
An alternative that has worked well for us - especially if you have a dedicated web designer (not just a graphical designer, but someone who can get it all the way to clean HTML/CSS) - is to have the designer build the end result in terms of visuals and then a developer go back and "wire it up". (svelte is a fantastic framework for this model, BTW)
Sometimes we'll instead have a developer do a rough implementation and then have the designer go back and make it look good, but in general the first approach works best because a good designer is often better at taking into consideration all the intricacies of HTML organization, and also be able to own changes to it over the long term.
But, perhaps worse, people produce most everything with string templates and string construction of some sort. Just look at how the LLM craze is panning out. We aren't leaving it, we are doubling down on it.
And on top of that you have no end of "web-purists" that claim that this is "standards-compliant browser-supported way of doing things unlike that non-standard abomination of React".
String templates are literally not much more more than this [1]:
tagFunction`Hello ${firstName} ${lastName}!`
becomes tagFunction(['Hello ', ' ', '!'], firstName, lastName)
// note, first parameter isn't strictly an array
That's it.Yet people use them as a (rathe poor) substitute for proper DSLs and macros.
77 years since Von Neumann, and we're still stuck holding our instructions in the data
That said, even if that existed, there'd still probably be some dangers you'd have to be careful to avoid.
iframe sandbox with srcdoc? Not so elegant, but it works. [Alternatively, iframe to another document and use CSP header there to ban everything]
- "No one has structured HTML creation that's as easy as string templates." -> I think this is usually true, but not always. Lisp allows use of S-expressions to define HTML (via SXML), and I genuinely find that more pleasant and powerful than writing HTML, or HTML with string templates. JSX is also arguably easier than adopting a string templating system since you can just drop into it instantly and write it as a function; it's language-integrated.
I also actually classify languages such as Haml as structural containment, which might be counterintuitive, since they're largely just seen as a syntax sugar on writing HTML. But because these languages are completely different to HTML, they're necessarily defined and parsed as languages in their own right, and the capacity for interpolating variables, etc. is integrated into the definition of that language as part of its syntax. (Consider how you can't provide mis-structured HTML with Haml, for example.) So this is a bit of a surprising result where something which is just seen as sugar on top of HTML actually takes you from string templating to structural containment. Thus using languages like Haml is a simple and easy way to move to structural separation in ecosystems where Haml or similar technologies are available. Since the whole point of Haml is to make writing HTML easier/lower effort, I also think this is another counterexample to "no one has structured HTML creation that's as easy as string templates."
- "One of my fundamental rules of system design is when people keep doing it wrong, the people are right and your system or idea is wrong." -> Yeah. If web developers keep adopting the wrong solution, it does suggest a lack of better leadership in this area in terms of the web development tools and languages they're provided with. I agree with this constructive view of leading via positive example; it's not reasonable to blame people for not using the wheel if they live in a world where nobody has ever seen one.
- It's mentioned that Go came up with html/template rather than come up with a better solution. It seems like the problem here is that in order for a language to be flexible enough to allow articulating HTML in a structural way that is actually ergonomic, it basically has to either a) be as flexible as Lisp, or b) be a DSL or language extension specially designed for the task (JSX, Haml).
While Chris takes a different view here I find his POV interesting, so I've added a link to his page from mine.
My hunch is that it’s for the same reason folks love other plaintext formats: they’re readable and easily manipulated.
HTML has the added upside that it’s pretty straightforward for humans to write, which whether one likes it or not is why some folks would rather not learn some else’s DSL for generating it.
Btw, I suspect there’s a similar division in the write your own SQL vs ORM/linq world.
It's probably me, but I'm not grokking this ending sentence. Is this praising go's "html/template" or is it meant as satire? I believe it isn't but genuinely can't say for sure. From a cursory look, html/template looks like doing the right thing, but come on, you have to tell it explicitly and procedurally which kind of escaping(s) you want for every individual var expansion. Compared to SGML which automatically detects where expansion of an entity reference occurs (within attributes, within CDATA or RCDATA, within normal parsed character data, and so on) while also considering the entity type and type-checking the final expanded result against the expected content model (not just for preventing <script> injection) this is a joke. Maybe SGML templating is what the author considers 'strict structural containment'? I mean, SGML isn't just any rando templating system, it's what HTML is based on, has been with us since 1986, and should be the point of reference.
This sounds similar to the concept of narrow waists: https://news.ycombinator.com/item?id=30483914
ALL complexity has a cost. String manipulation is such a core operation for so many software engineers that for many of them, the complexity of using an imperfect tool for the problem is outweighed by the complexity of bringing a new tool into the toolkit. (Or at least the perceived costs, at the time the decision is made.)
The trick is that JSX looks like a string template. Compared to building a tree of nodes (or using the DOM API), XML transforms, or other "more proper" ways to generate HTML, it's more intuitive.
Of course, there are issues with string interpolation, like XSS.
There would likely need to be some new kind of syntax delimiter for it to be a native feature. E4X-style would conflict with JSX-type syntax.
You would also need to look at client-side and server-side code separately.
On the client, it seems to me people are using string templates much less than before, and instead using JSX.
With external templating libraries, where you write your HTML in a separate page and add slots for the content, usually there is a lack of transparency in how that data gets filled, by who and when, especially when seeing the code for the first time.
2. browser engines try very hard to accept as much malformed slop as possible and still give usable results ( https://en.wikipedia.org/wiki/Robustness_principle ), this was probably a mistake
Something like 30% of the time this revealed a real bug, like double escaping or no escaping.
People gave up and went back to HTML, which is non-strict and will happily parse gibberish.
https://github.com/julvo/htmlgo
Which definitely looks interesting.
What are you gonna do, recompile and ship new and different executables to all {1, 10, 100, 1000, …} of your customers or end users for each point release?
WTF? That’s bananas. No thank you.
Because string templates have one less layer of abstraction.
I'd rather learn SQL and HTML than every language and framework's personal take on how to do this "right".
There are, as any number of other comments on this topic are pointing out, a huge number of "correct" ways to do it that exist today. None of them are as simple as
html = "<p class=\"" + params["post_type"] + "\">"
and they can't be, because the only thing that beats a one-token string append is a zero-token string operation... which to the extent programming languages have them, is also always an append.(Before you say your way is simpler than that... you need to include the setup code. That has no setup code; it is built into the language. If you so much as type "import " followed by anything you've already exceeded that in complexity, let alone if you add documentation you have to read for so much as 15 seconds or anything like that. It's not just about character count at the payload location.)
To fix this problem, first "room" has to be made to make the right thing easier than the wrong thing, and that involves the counter-intuitive requirement of not building the broken operator right into the language. Simply jamming two strings together must become more complicated than it is.
It also suggests that there ought to be something about the process of concatenation that raises alerts to a human reader, something like the minimum concatenation operation being
format_string("%RAW;%RAW;", str1, str2)
or perhaps "%UNSAFE;" or something.But I don't expect to see that happen anytime soon. Should a new language even try it, it would simply become the number one complaint people have about using the language ("omigosh this language is so broken you can't even concatenate strings" "wow thanks for that now I know never to try it out"), and a newborn language can't afford that.
We need to come up with something, though. The historical #1 security issue has been memory safety issues. It may still be, such things are hard to measure, but if they are, they're hanging on by a thread. The issue of string handling is rapidly taking over the #1 security slot, and it's probably going to be a similarly multi-decade adventure for the programming community to figure out what to do about it. I do know that just as we eventually realized git gud was not a solution to memory safety, we're going to realize git gud is not a solution to this problem either. I am pretty gud, have a deep understanding of the issue, and it's still walking through a minefield for me whenever I'm dealing with this issue.
I have no idea how to fix this thought, even in theory, without "making room" of some sort and making string concatenation fundamentally a harder operation. (And honestly most programmers would simply bash together an "string_append" function anyhow and call it "sa" or something and just bypass the attempt at security.)
This... isn't true. JSX is structured HTML creation. Lisp too obviously but that's not as popular in modern web dev. JSX is surely the most popular HTML creation method in modern web dev?
> I don't have any particular answers to why string templating has been enduringly popular so far
I was disappointed to read to the end and see this, as the article title wording - "is telling us something" - seemed to hint that the central topic would be trying to pose answers to what that "something" is.
> (although I can come up with theories, including that string templating is naturally reusable to other contexts, such as plain text).
This is something I think JSX is really close on, but not quite there - it might be nice to see a fully HTML/XML-compatible syntax that works exactly like JSX does. I know JSX started out with that idea, and the incompatibilities introduced were well considered and made sense, but I'd still be curious to try.
Overall, while I think the author makes a good point (fighting user tendencies head on is mostly futile) the benefits of structural generation are clear, while strings have no real benefits beyond UX/DX - acknowledging that the fight is futile is ok but I think it's still worthwhile to continue the fight in a more compromising/self-aware/middleground type approach. JSX fits this really well.
Would be nice to see the concept spread further outside of the React ecosystem, maybe even to other languages in some form.