BTW, when you say "inference of omitted tags" did you mean "inference of omitted close tags"? Because if so, this is not a feature. It's a patch to cover up a design flaw in SGML, namely, that close tags are required to match and so it is possible to make the mistake of omitting or mismatching them. This is one of the reasons S-expressions are superior to SGML: S-expressions are DRY. SGML isn't.
Early in the web history, browsers tried to out-do each other in guessing what broken HTML means and render it. That led to a nasty situation of everyone having to emulate everyone else's bugs and hacks.
To give a concrete example for tannhaeuser's point, consider this document:
<!DOCTYPE html>
<title>…</title>
<p>…
This is a completely correct, valid HTML document. The first thing to notice is that it's not a tree made up of elements. That first line is not a tag, and isn't part of the DOM tree.Then we get to the DOM tree. It's got several implied start and end tags, so the tree itself looks like this:
html
head
title
…
body
p
…
Again, this is a completely correct, valid document and this is the correct way to parse it. While you may be able to represent the parsed DOM as an S-expression, you can't represent all valid HTML documents as S-expressions, at least not in the convenient way people assume – and it's not "broken HTML" to blame.Of course you can. Here is how to express your example as an s-expression:
((:!doctype html) (:title "This is the title") (:p "..."))
Here it is being rendered by CL-WHO:
? (princ (html ((:!doctype "html")) (:title "This is the title") (:p "...")))
<!doctype "html">
<title>This is the title
</title>
<p>...
</p>Basically, it's stuck in-between two states doing neither correctly. It doesn't represent the actual HTML document, and it doesn't represent the parsed document structure. It's an alternative model of the HTML document that serialises to something that would be parsed in an equivalent way. I'm sure that's useful in a whole bunch of different situations, but it's not as simple as "S-expressions can do everything HTML can, in a convenient way".
S-expressions are great, and very useful. But they aren't the right tool for every situation. HTML is an odd markup language that only appears simple superficially, with all kids of irregular corner cases creeping in when you dig into the details. S-expressions would be a great fit if HTML were as simple as it appears on the surface, but it's not.
Markup is meant as a text format for content authors that can be parsed into a hierarchical structure, rather than as general-purpose data representation syntax, even though XML is being frequently (ab-)used for this purpose.
The original use case for markup is that you can take a piece of plain text and then mark it up with tags, unlike s-expr and/or JSON which arise out of the syntax of a programming language and need eg. verbatim text to be written as string constants/with quotation characters.
Yes, that was the original use case, but in actual practice HTML has not been used that way for a long time. Nowadays HTML is de facto used as a programming language for the visual representation layer of a browser. No one actually uses HTML to mark up documents by hand any more, for two reasons: first, no one writes plain text documents to use as source material for markup. They write Word documents, or TeX documents, but plain text source is all but unheard of nowadays. And second, HTML syntax is too clumsy and places too many demands on the user. So when ordinary people want to produce HTML they use WYSIWYG editors. When geeks want to produce HTML (and remember I'm talking about documents here) they use markdown. The only time anyone writes HTML nowadays is when they want to make a browser do something fancy.
S-expressions are a data structure, different from the DOM, but S-expression syntax is a syntax. Normally S-expression syntax is parsed to produce S-expressions, but can also be parsed to produce other things. S-expression syntax can be parsed to produce a DOM. The easiest way to do this is to parse S-expression syntax ino S-expressions, render those S-expressions into HTML code, and then use an off-the-shelf HTML parser to parse the HTML. But you could also write a parser that parsed S-expression syntax directly into a DOM if you wanted to. You could also write a transformation program that compiled S-expressions directly into a DOM without going through the intermediate HTML.
The answer to your question of how to add an attribute to an implied element is that it is not possible to do that in HTML. It is only possible to add an attribute to an implicit element of the DOM produced by parsing an HTML document that omits that element (because at that point the element is no longer implicit). The exact same thing is possible using S-expressions. For example, here's how you write tables in my library:
(:table (header header ...) (data data ...) (data data ...))
This string of characters is parsed by the Lisp reader to produce an S-expression that has a one-to-one correspondence with the string you see above. But then there is an extra processing step that transforms that into a different S-expression whose printed representation is:
(:table (:tr (:th header) (:th header) ...) (:tr (:td data) (:td data) ...) ...)
At that point you can manipulate that S-expression in the same way that you manipulate the DOM (because they are both just data structures). Once you're done, you convert the S-expression to a DOM. At the moment that is done by rendering to HTML, but as I noted above that is just an implementational convenience to take advantage of the fact that HTML->DOM parsers are available off the shelf. You don't have to do it that way (and indeed the world would be a better place if it were not done that way).
All of this is trivial when dealing with S-expressions precisely because of the strict 1-to-1 correspondence between data structure and visual representation that does not exist in SGML-derived languages. That is why writing code for SGML-derived languages using S-expression syntax is so advantageous. (Actually, this is true for any language, not just SGML-derived languages. It's just a little more obvious for SGML-derived languages because SGML syntax already kinda sorta looks like a data structure representation so it's a little easier to grasp what is going on.)
I believe this is where the confusion is coming from. When you parse HTML syntax, you get a data structure; this is the same as when you read sexpr syntax, you also get a data structure. Both these data structures are different from the DOM tree.
Try this example:
<pre>
<span>one
</span>
<br>
<span>two</span>
<br />
</pre>
Can CL-WHO generate HTML that matches that? (i.e. feed both into a tool like BeautifulSoup and produce the same data structure?)Outside of CL-WHO and Hiccup-type libraries, you can of course use S-exprs to represent the same data structure. Here's a hypothetical S-expr syntax that might produce the same data structure:
((pre)
"\n " (span) "one\n " (/span)
"\n " (br)
"\n " (span) "two" (/span)
"\n " (br/) "\n"
(/pre))
Which is what I believe JimDabell meant by:> you can't represent all valid HTML documents as S-expressions, at least not in the convenient way people assume
In the case of S-expressions that is true. In the case of HTML it may or may not be true. It depends on how the HTML parser is implemented. There is a "natural" mapping of HTML onto a parse tree that is different from the DOM, but that is not part of the standard (AFAIK).
> Can CL-WHO generate HTML that matches that?
Yes, though native Common Lisp does not provide c-like string escapes so putting in newlines is a little awkward. You could, of course, bring in a string interpolation library, but here's how you can do it without that:
? (defun nl () (who (fmt "~%"))) ; NL = NewLine
NL
? (defun nli () (who (fmt "~% "))) ; NLI = NewLine + Indent
NLI
? (princ (html (:pre (nli) (:span "one" (nli)) (nli) (:br (nli) (:span "two") (nl)))))
<pre>
<span>one
</span>
<br>
<span>two</span>
</br></pre>
Or you could do this: (html (:pre "
<span>one
</span>
<br>
<span>two</span>
<br />
"))
which looks like cheating but is actually closer to the spirit of the original.The PRE tag is really weird because it actually changes the way things inside it are parsed. You can actually implement that in Lisp too via reader macros. CL-WHO doesn't support that out of the box, but it's not hard.
I can't imagine anyone actually wanting to do that, though. The PRE tag is for presenting pre-formatted text without changing its appearance, so embedding other tags inside it is kinda perverse. [EDIT: I was wrong about this. See below.]
pre provides the simplified line breaking and usually a monospaced font. However, tags are available to do whatever else.
A major example is that the Vim editor uses pre for formatting syntax colored code to HTML (when you do that with :TOhtml).
The output is a pre block containing various span elements which are styled with CSS.
BTW where in the HTML spec does it say that the interior of pre is parsed differently?
If we are parsing HTML (to Lisp objects or whatever), we should preserve the exact whitespace. The reverse generation should regurgitate the original whitespace.
If we take the license to eliminate newlines, then we ruin pre. The fix is simply not to do that.
I was wrong about that. I had a vague memory of putting HTML inside a PRE tag once and having it come out as if it were escaped, but apparently I hallucinated that.
> A major example is that the Vim editor uses pre for formatting syntax colored code to HTML (when you do that with :TOhtml).
OK, I stand corrected on that too.
> If we are parsing HTML (to Lisp objects or whatever), we should preserve the exact whitespace. The reverse generation should regurgitate the original whitespace. > If we take the license to eliminate newlines, then we ruin pre. The fix is simply not to do that.
Right.
Actually, I just realized that I mis-read the example. I saw <br /> and thought it was </br>. (Maybe the OP edited it?) In any case, the example now reads:
<pre>
<span>one
</span>
<br>
<span>two</span>
<br />
</pre>
And you can render that in sexpr syntax as: (:pre "
" (:span "one
") "
" (:br) "
" (:span "two") "
" (:br) "
")
This is a particularly bad example to demonstrate here because the whitespace in the code plays badly with the whitespace in the HN markup. But I tried running this code and it does work. Here is the output copied-and-pasted verbatim from my listener: <pre>
<span>one
</span>
<br />
<span>two</span>
<br />
</pre>
Note that both BR tags are rendered as <br />.This is why you would need separate tags to emit them properly with an S-expr syntax (tag), (tag/), and (tag)(/tag) in my example.
<tag> ==> (:tag)
<tag/> ==> (:tag nil)
<tag></tag> ==> (:tag "")
Using (:tag/) is a bad idea because that would screw up attributes.CL-WHO doesn't support this, but that would be easy to change if it ever actually mattered to anyone.
In HTML, <script></script> is valid. <script /> is invalid. <br /> is valid. <br></br> is invalid. So they are represented differently.
> Using (:tag/) is a bad idea because that would screw up attributes.
For my example?
((:tag/ :attr "value")) => <tag attr="value" />
((:tag :attr "value") "..." (:/tag)) => <tag attr="value">...</tag>
> You actually can distinguish between those if you really want to. It's just a matter of picking a convention.That sounds like it could work. So a leading `nil' would be treated as a special case (not a child node):
(:pre "
" (:span "one
") "
" (:br) "
" (:span "two") "
" (:br nil) "
")OK, then the best way to handle that is to let the HTML-renderer know that different tags need to be rendered differently if they're empty. Are there any cases where you would ever want to distinguish between the various kinds of empty tags?
((:tag/ :attr "value")) => <tag attr="value" />
((:tag :attr "value") "..." (:/tag)) => <tag attr="value">...</tag>
No, that's not what you want. Let's start with this general form:((:tag attr value ...) content ...) => <tag attr=value ...> content ... </tag>
Let's assume we have no attributes so I don't have to keep typing those. Then we have:
((:tag) content ...) => <tag> content ... </tag>
In this case (no attributes) we can unambiguously remove the parens around (:tag) and get:
(:tag content ...) => <tag> content ... </tag>
Now if we have no content we get:
(:tag) => <tag></tag>
All this is still completely regular, no special cases. But now if we write (:br) we get <br></br> which is not what we want. So we need to tell the renderer that some empty tags get rendered one way, and other empty tags get rendered another way. CL-WHO does this.
Notice that we have not actually typed any / characters. This is important. The role played by / in HTML is played by the close-paren in sexpr syntax. If we re-introduce the / into our new syntax we will have a hopeless mess.
> So a leading `nil' would be treated as a special case
That is exactly right. If (and this is a big if) we want to be able to write something equivalent to both <tag/> and <tag></tag> in the same document we have to be able to distinguish between those two things in the markup somehow. I just looked this up and the distinction that HTML makes between <tag /> and <tag></tag> is that the former content is EMPTY while the latter content is "" (i.e. the empty string). So really the Right Thing would be:
(:tag) => <tag />
(:tag "") => <tag></tag>
That will work, but now we have to remember to add an empty string in some situations, e.g.:
((:script :src "...") "")
Personally I would find this annoying, so I would choose to go with the lookup table.
#!/usr/bin/lisp
(defun foo () a b c)
has an "inferred PROGN" around a b c and the first line isn't part of the tree.What you're not getting here is that the above broken HTML has a canonical HTML representation. That canonical HTML can go to S-exp.
If we are doing HTML-in-Sexp, we can throw out some of the non-canonical aspects, keeping the ones we like. We can certainly infer element wrapping if someone using our HTML-in-Sexp finds that useful.
Is that really necessary?
> the above broken HTML
As I very clearly stated, it's not broken. It's completely correct, valid HTML. Stick it in a validator if you don't believe me. Yes, I know a lot of people assume otherwise. That's because HTML is only superficially simple but has unexpected irregularities once you dig deeper. This just reinforces my point that HTML isn't the nice neat package that fits well with S-expressions you think it is. The canonical HTML representation of that "broken" HTML is simply the HTML I provided, unaltered, which is not conveniently representable as an S-expression. Please, before trying to reinvent HTML-as-S-expressions, take the time to learn what is and isn't correct HTML. You seem to be assuming the language is simpler than it is and any irregularities are because the sample HTML provided is "broken". This isn't the case.
<a><c></c></a>
End-tag inference is semi-syntactic. You can tell that the following might have a missing end tag: <a><c></a>
But the previous example definitely does not. <!ELEMENT html O O (head,body)>
tells SGML that the html element should contain a head element, followed by a body element. "O O" (capital letter O for omission) are tag omission indicators (in this case meaning that both the start- and the end-element tags for html can be omitted).This is covered in depth in the linked paper/slides (in fact, covering JimDabell's example exactly).
<!document ...>
<a><c></c></a>
<!document ...>
<!element a O O (b,c)>
<a><c></c></a>
They just have different semantics. If parsed as HTML, the first produces a DOM with two nodes and the second produces a DOM with three nodes.By way of contrast, this is syntactically incorrect:
<table><tr><td>data</img>
So is this: <img src=x<y.html>
There is no DOM that corresponds to those two examples.One of the (many) problems with SGML is that it muddies the distinction between syntax and semantics. That is one of the (many) reasons that using S-expression syntax to write SGML-like languages is advantageous.
Are you talking to me? I've just pointed out how SGML works and didn't say anything about syntax/semantics.
> > If you have a good strategy for validating your template files, I'd love to hear it!
> Use S-expression syntax instead of SGML syntax [emphasis added]
Note the use of the word syntax. I'm talking about syntax. All of your responses have been at best irrelevant or at worst wrong because you either don't understand what syntax is, or you chose to ignore it. It's damned annoying, particularly when you start making demonstrably false claims like, "it's nowhere close to the power of SGML as a text format" (https://news.ycombinator.com/item?id=13569991).
(I hope for the sake of SGML and HTML that you're the one confusing character level syntax with tree manipulation.)
[1]: http://drdobbs.com/a-triumph-of-simplicity-james-clark-on-m/...