But what people actually need are grammars.
The exact same reason why parsing HTML with a regex unleashes Zalgo is why generating HTML with string templates is bad. Because both treat HTML as a string, not a grammatically restricted language.
But what people actually need are grammars.
The exact same reason why parsing HTML with a regex unleashes Zalgo is why generating HTML with string templates is bad. Because both treat HTML as a string, not a grammatically restricted language.
I don’t actually disagree with you for the most part, but I feel that an important caveat has gone unacknowledged.
Grammar formalisms have the same weakness compared to dealing with raw strings as sound static type systems do compared with dynamic typing: there are small, mostly isolated islands of feasibility in a sea of intractable (often undecidable) generality, and if your problem doesn’t fit inside those borders things start to get nasty (cf how even GCC’s handwritten rec-descent parser didn’t get its lexer hack interactions correct in all cases[1]).
I still agree that we spend criminally little time on syntax. Starting with the simplest cases: with how much time is spent in school on “order of operations” you’d think we could take a moment to draw[2] a damn syntax tree! But nooo. There are in fact working mathematicians who don’t know what that is. (On the other hand, there are mathematicians who can explain that, in a sense, the core of Gödel’s incompleteness is not being able to reason about arithmetic—it’s being able to reason about CONS[3], which arithmetic happens to be able to very awkwardly do.)
[1] https://gcc.gnu.org/bugzilla/show_bug.cgi?id=67784
[2] https://mlochbaum.github.io/BQN/tutorial/expression.html
Generating JSON data using string interpolation or templating is clearly wildly insane, right? You don’t do it.
Maybe for some config file generation scenarios you might just run a template JSON file through a token substitution or env var interpolation or something. But you’d feel bad about it, because it’s so easy to NOT do it that way. And even then you’re not interpolating in JSON fragments like ‘“age”: 25’ - you’d have the decency to only interpolate in values like ‘25’.
In the node ecosystem it’s so easy to switch from a .json file to a .js file, too, if you want to build the json dynamically.
For some reason people feel more willing to attempt it with YAML. And then regret it when they realize how significant indenting has screwed them.
And then with HTML people just give up and go ‘yup, it’s all text, even the angle brackets’
I'm sorry to report that I've seen a lot of JSON generated by string concatenation and templating, in different projects.
Often using 'printf' or 'echo' in various languages. Sometimes using whatever's used for HTML string templating if the JSON is embedded in HTML or served as a resource similar to HTML.
Yes, its horrible and breaks if fed variable values that have characters like quotation marks in. People do it anyway.
Even in languages that have perfectly good data structures and JSON libraries.
I've seen a fair amount of parsing values out of JSON using regexes too, assuming specific formatting of the supplied JSON.
It's less common, because JSON is simpler, so the tradeoff point for using a grammar is lower, but it still makes sense in things like shell scripts, and other cases where the equivalent of `print(obj)` (or `eval(totally_not_rce)`, but let's pretend that's not available anyway) doesn't happen to produce (or consume) valid JSON by coincidence.
# Using:
grep -oP '(?<="bar": ")[^"]+' foo.json
# and
printf("{\"count\": %i, \"type\": \"%s\"}\n",nfoo,tfoo);
is a general-purpose solution that can be adapted to pretty much any text-based format just by looking at examples, without having to cross-reference with a external specification (that the thing you're feeding input to or pulling output from may not even correctly implement anyway), so obviously people do that!See also various discussions under the heading "Worse is Better". Whether it's the right thing or the wrong thing, it very clearly is a thing.
I think everyone agrees that those approaches with JSON are bad.
This file generates a feed of events (rehearsals for my high school play) to be rendered by the FullCalendar JS plugin. FullCalendar required a particular data schema that didn't match the format of my MySQL table, which meant I couldn't just json_encode() the MySQL results. I guess I just didn't conceptualize that I could create a new object that matched the FullCalendar format, and then call json_encode(). So, I generated JSON with strings.
Honestly it's a toss-up whether the JSON generation is the worst thing about this file. It looks like I also made a separate database query for every single row to get the username, because I apparently didn't know how to do joins. Could probably spend an hour listing some of the other little nuggets of awful in there. But hey, it got the job done! :)
All is fair in love, war, and programming.
Taking the opportunity to ask for more sane alternatives. What do the crowd here use for manifest templating?
But I'd never try to implement my own parser or output deeply nested JSON.
Takes all the tree and hierarchy management away, makes it so ordering doesn’t matter.
If I’m generating JSON from batch scripts it’s my preferred tool (easier than fighting jq for many tasks)
It knows when to do escaping and how. It also can detect, though dynamically, when fragments have been combined into an illegal sequence which would be rejected by the full grammar. It can not however guarantee that the result will parse only that it can not detect that it would fail.
The ugly quasiquoting seems unfortunate (I’ve a half-serious suspicion the reason Template Haskell never got popular is that it looks so bad), and the GLR sledgehammer precludes ever having a lightweight implementation, but otherwise it seems like a interesting entry in the extensible languages story.
5 - 2(3 - 1) - 5 = ?
Also, respectfully, it doesn’t matter. Not having learned maths in English, I don’t know the mnemonic, I don’t care to know it, and I find even the concept of it completely asinine. (For eighteenth-century mathematicians, addition and subtraction bound tighter than multiplication and division, and they could calculate perfectly fine.) You can look up the precedence table if you need to—as long as you need to understood the idea of precedence (and not order of operations, for goodness’ sake). You won’t then be able to calculate fluently, but fluency is a different problem with a tedious and time-consuming solution, and given the time crunch I’d rather talk about some actual Maths as She Is Spoke instead.
6 / 2 (1 + 2) = ?
I approach the problem the same as I would 6 / 2(x + y). When the multiplication is missing 2(x + y) is a single term. The implicit multiplication is part of the parenthesis and reduces the problem to 6 / 6. People who argue that you have to strictly use PEDMAS left-to-right will divide 6 / 2 first and get 9.
Neither way is wrong as long as you can explain the process but everyone wants to argue and have there be a single answer.
A test question like "6 / 2 (1 + 2) = ?" is not asking for the mathematical meaning of those symbols it is asking for "Guess what I as thinking when I wrote this".
(Unless it is a programming class and you are learning how the compiler reads your code)
My older relatives are the ones I see repost these inane order-of-operation tests, getting the answer consistently wrong.
÷ x
is equivalent to 1
* —
x
The thing schools don’t do a great job of doing is explaining when transition from doing ‘arethmetic’ to doing ‘algebra’. Many of the symbols you use in arithmetic continue to be used in algebraic notation, but what they mean changes subtly. Arithmetic is a procedural activity - performing a series of operations to get to an answer. Algebra is a declarative activity - making truthful statements about the world.For example in arithmetic
x + y
means ‘add y to x’. But in algebra it means ‘the sum of x and y’. In arithmetic ‘=‘ means ‘gives the result:’; in algebra it means ‘is equal to’.The failure of teaching to explain that you’re moving on to use those symbols to do something fundamentally different is, I think, one of the things that leaves some kids behind and dooms them to always annoy their relatives in Facebook comment threads about operator precedence.
Moreover, important parts of HTML processing would be significantly more brittle and complicated and less powerful with objects: "escape some completely arbitrary text to valid PCDATA or a CDATA section, whatever is shorter" is strictly more general, robust and principled than "render a Street Address to a fragment that isn't supposed to contain markup".
HTML is a grammatically restricted subset of text.
I can take arbitrary text and embed it in HTML by escaping characters within it. That produces a grammatical fragment of HTML that represents the arbitrary text, but it is not the text.
As an example sentence, take the following:
"The French equivalent for the English "Good Evening!" is "Bonsoir!", whereas Italians might say "Buonasera!" to one another for similar effect."
There are four languages in that sentence, two of which are English. You may need three editors to deal with them, or you can flatten the sentence and simply edit everything assuming you knew all three.
This is not even true of natural language, which has a vocal representation that is at least as important as the written representation.
Though I agree that representing HTML as objects is a poor substitute.
The problem with HTML is that its syntax trees are relatively unpleasant to use.
For anyone who hasn't seen it yet, top answer from https://stackoverflow.com/questions/1732348/regex-match-open...
Besides not "being proper" or whatever your argument boils down to, people (arguably) are doing useful things by just manipulating strings.
I'd argue most of the web is probably built with just strings and duct tape holding all the pieces together.
It would be better if programming languages discouraged people away from those mistakes and prodded them towards the pit of success by making string concatenation harder and providing better tools for constructing grammatically sound structures.
Python's MarkupSafe (used in jinja) and go's html/template are good examples.
But if we were talking about outputting CSV data you wouldn’t be able to say
In lisp the templating systems are S-expression based instead of string interpolation based, which at least models the tabular structure of CSV documents
Because hierarchies don’t model tables especially well.
So the suitability of lisp for outputting HTML feels slightly coincidental.
I mean it basically worked with JSON too.
The probably slightly newer aspect is producing an intermediate representation that is then serialised to HTML, though I think that’s still going to be back in the ’90s. But the oldest examples I know of (while I was yet a small child) used functions and methods to produce serialised HTML strings directly, which was more efficient (at least in the languages in question) and also allowed you to mingle with string templating.
Perl’s CGI.pm let this example be written no later than 1997 (no idea when it was actually written, can’t be bothered searching harder for older than CGI.pm 2.32): https://github.com/Perl/perl5/blob/54310121b442974721115f936...
For stuff that worked on the frontend, it’s still way older, though it tended more to XML-based stuff like XSLT (… which still works in browsers now, e.g. https://chrismorgan.info/blog/tags/meta/feed.xml is an Atom feed but the <?xml-stylesheet?> processing instruction is basically a pointer to the file for the browser to use to convert it to HTML which it then renders). But there were definitely things in this vein even on the frontend in active use more than five years before Mithril, though I can’t be specific as my memory is fuzzy as I wasn’t paying much attention to it all back then.
Literally cringing as I read the readme. We decided over a decade ago that writing HTML with code is rediculous but somehow it comes up again and again.
A designer shouldn't need to code JavaScript to edit your design.
Web applications aren't just HTML though, that's why code might be a more appropriate format.
You can argue that designers need better tools to edit structured markup in other formats, but that doesn't entail that HTML should be the default format. For instance, something like repl.it for mithril or similar that immediately renders the output so you can see the results would be useful.
This is something that literally every single framework, front-end and back-end has had to deal with in some way since the 90s. From chucking ugly <?php tags to more elegant solutions like curly braces.
Having one standard in the spec would standardize something that is currently done a million different ways.
https://developer.mozilla.org/en-US/docs/Web/HTML/Element/te...
This gets complex fast, I cannot imagine having something like this as a slow-moving spec.
Also, XSLT is horrible.
That said XML is like violence, if it’s not solving you problems you need to use more.
The idea of producing visual presentation by running XSL-T and then XSL-FO was the kind of thing that makes you think writing TeX macros might be easier.
Hilariously over engineered for the problem users actually wanted to solve (making data driven web pages look pretty).
It's a very natural proclivity of a language designer designing a template language for language X to want to find a way to articulate that template language in that same language X also. I think most engineers can't help but love ideas like that. It's probably the same reason everyone who creates a programming language wants to make a self-hosted compiler for it.
In this case though, XML being a rather verbose and arguably limited semantic markup language for textual documents, it's an extraordinarily unergonomic choice for templating itself (in notable stark contrast to SXML combined with Lisp code, which is an example of homoiconicity between the structure being templated and the templating language works very well).
And Lisp S-expressions are a fine way to model mark up and transformations on markup.
XQuery can do everything XSLT can but with different syntax. XQuery 3 can also handle JSON. It's a clean way to generate well-formed XHTML.