Why I like XSLT
guy-murphy.github.io
guy-murphy.github.io
You know that a technology has got serious problems when the best thing you can say about it is "well, it is at least preinstalled everywhere"!
>Benefit #2: Separation of concerns
If you want to put business logic in your XSLT, you can. It is a turing complete language (unlike, say, jinja2).
True, the language is so terrible that you will probably do everything possible to avoid putting business logic in your XSLT and that would help with your separation of concerns. I do not consider this a benefit, however.
>Benefit #3: Transformations
This is not a benefit of the language. This simply what it was originally intended to do.
>Benefit #4: Focus on API: The focus on XML as the result of a controller naturally focuses development on the application API
As opposed to other languages where you don't focus development on the application API?
>Benefit #5: Sand-boxed
Definitely a benefit but far from a unique one. It is also doesn't have particularly good sandbox implementations: https://www.mozilla.org/en-US/security/known-vulnerabilities...
>Benefit #5: Ease of testing
Compared to what?
> You know that a technology has got serious problems when the best thing you can say about it is "well, it is at least preinstalled everywhere"!
... plus it's not even true. For any collection of XSL stylesheets, representing any kind of real world complexity, the chances you can port them as-is from A to B is highly dependent on which XSLT features were used (1.0, 2.0 and now 3.0 ), what the XSL processor supports plus "host" framework idiosyncrasies ( Java, .NET, PHP ( http://php.net/manual/en/xsltprocessor.registerphpfunctions.... ! , etc. ).
> This is not a benefit of the language. This simply what it was originally intended to do.
If a technology intends to be good at something, and then is, that's not a benefit of the technology? Huh?
In any case, transformation, i.e. push vs pull templates, is indeed an interesting distinction and feature of XSLT. Using XSLT for push-style problems (like templating a book, a manual, etc.) is where it is actually useful.
The point about the syntax is a good one, but I don't like the idea that people should just equivalently like any syntax that expresses the same basic ideas.
I can't say why exactly but something just feels fundamentally wrong with writing code in XML. I felt the same way when I had to write MSBuild scripts, which are XML. I guess it could be trying to do imperative things in a declarative language, but that doesn't seem to be a problem for Lispers.
Made my day.
Syntax is verbose, but felt that the power was worth it. Never found a declarative, functional way to do the same kind of transforms on JSON data, but I'm sure it must exist by now.
But then came JSON, Firefox and other browsers that made XSLT/XML look a less supported and ageing tech.
I tried to write a js equivalent in 2008. And after a couple of iterations it finally ended up with this: http://gitHub.com/pure/pure
It is declarative. With CSS selectors instead of XPATH. Do not mix at all with HTML. And fast rendering.
Now for my own e-commerce ETL system I actually had to do a transformation engine encompassing JSON, XML, SQL and NoSQL in Python and rather decided to stick with the KISS principle, allowing declarative stuff for basic simple operations such as multi-key joins and offloading advanced stuff directly to customizable Python scripts.
From the ground up, I admit I like XSLT. But: maintaining large sites that inherit templates (call-templates and apply-templates) and style through <xsl:include/> and <xsl:import/> over lots of files tends to become a template mess. After a while only the experienced devs are able to maintain the templates.
Then, it's a functional language, where lots of things are overly complicated to do; you can do just everything, but that doesn't mean you should.
I work together with a bunch of developers on a website that renders its HTML output with XSLT. What you address is architecture.
What we basically did is we have one XSL file for each "element". Elements are: button, image, headline, etc. Each of those elements inputs "our own" XML syntax with wisely chosen defaults (some HTML attributes like class,id,style get just passed through) and outputs Twitter Bootstrap HTML. Writing responsive websites in our framework is extremely easy. So each element has a XSL file where the magic happens, an XML file with documentation and tests and additional files like scass/css/js files with the same name. Accessing them is extremely easy when you use something like Sublime Text (what we all do).
All those different templates get written into just one XSL file automatically with the help of node.js' gulp.
The website gets delivered in Apache Cocoon which has a pipeline concept similar to gulp. Once a page got rendered, it is cached until the next update (except for pages with dynamic content of course).
The whole frameworks documentation is basically a website written in this framework with automated tests.
We use XSLT 1.0 + EXSLT and have written around 30 (EXSLT) functions to make sure everything works like we expect.
Everybody in our team loves it.
If you start out "right", you probably will be able to enjoy XSL.
I would love to see someone propose a modern alternative to XSLT, that works for the same use cases, but avoids XSLTs mistakes.
Just having a standard format which is XML-based which is able to take in XML and spit out: different XML, plain text (e.g. CSV, HTML, bespoke UTF-8 formats, etc), and so on is a wonderful idea and XSL does in fact accomplish it.
But, yes, writing XSL is extremely painful. Much of it needlessly so. Although XSL 2 did fix some of my whines (DateTime library?!). I'm still legitimately annoyed (irrationally so) that indexes start at 1 and not 0 in XSL.
The way template matching works is clunky. I think the first answer in this Stackoverflow question illistrates my point well:
https://stackoverflow.com/questions/3127108/xsl-xsltemplate-...
Use an XPath library for your favorite language, and write in your favorite language. XPath is the vast majority of the good in XSL, and there's no reason to put up with the steaming pile of the rest of it just for that bit anymore, and hasn't been for a while.
If XSL hasn't been replaced, well, it's because it's been replaced by too many great languages for doing the same work rather than too few.
I will stop using XSLT as soon as there is a better way to... transform XML. It is a pure functional language you can use to make web pages, yet it is really automata non grata, if you will, at least to the extent that many people want to kill XML itself.
I wish Microsoft had implemented XSLT 2.0. It seems to me that Michael Kay has almost singlehandedly kept the whole thing alive, but running Saxon in .NET over IKVM (or in Javascript over GWT) is a bit too high a tax for me.
What do you mean by that? Can't you just parse XML into a Clojure map, transform it in Clojure, and then emit the result as XML again? Using this: https://github.com/clojure/data.xml
Or did you mean something else?
I use XSLT as part of a reporting engine that generates HTML from fairly complex XML documents on the back end, and whenever I touch it I'm amazed by a) how powerful it is and b) how painful it is to use.
It isn't clear that it's possible to get similar power without similar pain, though, and I think it's beholden on anyone who claims it is to step us and give us an example. A lot of really smart people were involved in DSSSSL and XSLT, and they couldn't solve the problem. Even stripped of the awkward, ugly, syntax, the kind of thing you want to specify in XSLT is complicated and messy.
"Emit X when you encounter P except if the third element of type Z (if it exists) of the grandparent of P is a Q with attribute TUESDAY having value FIZBIN emit Q's content transformed according to rule Y" is never going to be pretty, and while you can argue that only badly designed document formats require things like "the third element" to be specified, the world we actually live in is full of such things. And because we will always re-purpose document formats to ends their designers never envisaged, it will always be the case.
[#] I can't actually remember how many S's DS*L has in it, although I once did know most of the words to the DSSSL Song.
And that is the logic that does not belong in the view. You've failed at properly separating the M,V, and C.
The database query is complex. The presentation of data should not be. Mixing these together is a design failure to a lot of people. TFA even mentions this, if you subsequently need to generate a .pdf you're going to embed a redundant copy of the logic in that display code too. If you get the logic down at the lower level where it can be shared, then you can have as many views as need be without duplicating logic that doesn't belong there anyway.
What one does is transform different input sources (using XSLT) into a common XML data format, applying stupid rules like the above. Then one applies another transform from the common data format (again using XSLT) to something like HTML, or XSL-FO, or CSV.
"Separation of presentation and query" is really orthogonal to both the use of XSLT and the inevitability of needing to apply complex rules when processing real-world data. XSLT and XPath just happen to be better at expressing those transformations than most other languages, and they certainly have the largest installed base.
XSLT v1.0 is severely limited and many of the idiosyncratic behaviour ppl associate with XSLT in general is because they hit some very hard to parse/understand/grok things in XSLT v1.0
I remember in 2001-2002 sending the following link, when anyone said 'I am having a problem with XSLT transform'
Are there now any free, portable XSLT 2.0-compatible processors (ideally, that match xsltproc in ease of use)? Surely no-one is comforted by "this version fixes the problems of the previous one, but you can't use it"!
EDIT: OK, Saxon (I won't link to it because I've always felt a little skeeved out by their "let's make people pay for schema awareness" business model, however legitimate it is)—it's even available through MacPorts. I don't know if I just never knew about that, or if the XSLT 2.0 support wasn't always free.
https://github.com/cgrand/enlive
See Brian Marick's 5-part tutorial for a good overview:
https://github.com/cgrand/enlive/wiki/Table-and-Layout-Tutor...
which could easily be ported to js
I especially like the power of combining XSLT with Batik and FOP. You can take backend data and present it as an interactive table, a visualization, or a PDF report really easily.
I will say that the syntax is awfully obtuse, though. I find that support for various powerful-but-obtuse 1990s/2000s solution stacks splits along an IDE-user line: those who use IDEs are generally comfortable working with XML and XSLT in Java while text-editor devotees tend to hate that class of solution.
Alternatively, something Perl's Twig let you manipulate the DOM without too much of a headache, and is preferable to XSLT in many cases.
* first bugbear - XSLT is not an implementation of lisp (or insert lang here), we shouldn't care if it implements said language well or not
* XSLT is extremely well (saxon) implemented in some processors and poorly in others ... unsure what that has to do with the language itself.
* XSLT has more implementations then a lot of other popular languages (Ruby, Perl, python come to mind), that doesn't make it better (or worst) then those languages ... its just a fact.
* there are no security vulnerabilities in the design of XSLT that I am aware about (well, there are, and most certainly are in implementations, but at that level I can call out issues with every language and its processor(s)) ... stating it is insecure by design is rubbish
* anyone using XSLT v1.0 should know that there has been, for years, an XSLT v2.0 production ready version in the form of saxon-ce (w/ appropriate open source license) ... I can't explain why folks insist on using XLST v1.0, some are doing so cause they can't migrate to XSLT v2.0, many of them are unaware of new options.
I can understand why folks find XSLT hard, or hurts their brain or makes them want to cringe at the mention of an angle bracket ... teaching basic programming idioms is hard for most people, learning fp principles is even more of a bender (esp if you come from OO world) ... XSLT template matching is a powerful mechanism which allows one to perform transformations with a succinct set of matching rules ... but yes it is verbose and like many languages, XSLT has its ugly points to the language.
If you are allergic to angle brackets, then any language that uses them is going to be anathema to them, but every person I've known who have learned XSLT became a better programmer for it. Thats my own personally biased single data point observed 'fact'.
My biggest gripe with XSLT is debugging. It's just not very much fun. Insofar as it forces you to reason very carefully about your code, that's probably good practice if you haven't already had plenty of practice -- e.g. developing websites before we had decent browser consoles, or writing INITs and TSRs in the 80s.
Lately I've been working a lot in React, and I've found that just understanding the new syntax can be learned in a weekend -- instead of transforming an XML doc, I load JSON into my state, and instead of writing <xsl:template>s I write React classes. The sort of skills that one refines over years, namely how to think of your rendered page as being composed of various smaller templates, benefit me everywhere. XSLT was a really intensive way to force myself to learn to think that way. Even my mixed PHP/HTML that I occasionally still have to write is considerably better.
to your point about XSLT v1, I can speak to why it's so rarely used at least by PHP devs: v1 has been included as a default extension in PHP builds for many years, but v2 has always required a separate installation of saxon. I just checked on my reasonably up-to-date ubuntu default install, and it's XSLT 1.1.28. That's still far better than Go, where the only approach I've seen is to just install command-line xsltproc and execute that. v2.0 isn't going to catch on while v1.0 is what's considered the de-facto standard, and v1 being the de-facto standard is a big part of why so many have a poor opinion of xslt.
The big pain points were all around the library support. The XSLT lib was generally fine, but the XML processing PHP provides was rubbish. The main complaint I heard was that even PHP5.5 wouldn't support outputting as html5, and you need to manually just echo out a '<!DOCTYPE html>' to work around that. I couldn't just import HTML directly from elsewhere as rendered by my CMS into my document tree either - it needed to have string literals of tags concatenated to it just to work. The sheer number of piddly little flags that are barely documented but you need to get exactly right in order for it to work at all was a big nuisance too. I also needed to turn off all parser warnings so it wouldn't barf on reading html5 docs, and when a big advantage of the language is that you can be strict about your markup, it seems quite silly to have to turn off warnings.
Looking back on it, none of my problems were really with XSLT as a language, just with how it's implemented in various languages. I feel like a lot of people got turned off when they learned about it in school, and consequently nobody was interested in doing a good job of implementing its supporting libraries in any language.
Email me if you'd like to chat about it!
Getting rid of the noise is my biggest concern. Ideally, a developer should never have to know about or touch the HTML that's being generated. It's a lofty goal, for sure.
Being using them since 1999 too. Still use them today although more work is being done with JSON/JS.
How is writing XSLT not coding?
This is why templating languages are better at helping you maintain separation of concerns - they physically won't let you put business logic in the view layer because they aren't turing complete. XSLT will.
This might further discourage putting business logic in the view layer, but, OTOH, it just as easily result in developers starting to go down the logic-in-view road, running into things that don't work there, and then ending up with some business logic in the view and some business logic in more appropriate layers.
And there's really nothing inherent in view transformations vs. business logic that necessarily makes the latter sufficiently served by a weaker computational model than a Turing complete one, so not having a Turing complete language for the view can result in view transformations in other layers or constrain the flexibility of presentation if that isn't done.
There is. View transformations are essentially configuration. As such, the ideal language for describing them is declarative.
To keep declarative languages clear, you need to make them not be turing complete. If they are turing complete they become unpredictable and difficult to understand.
Business logic requires turing completeness and usually requires mutable state.
Also, I've managed to setup a system to generate my resume via an XML/XSLT transform. I haven't published how I do that (or the public version [which is out of date]).
It's a tool and it can be very useful if you understand what you're doing. Additionally, there is Unix support for XSLT processing.
https://en.wikipedia.org/wiki/Comparison_of_web_template_eng...
For my CS honours thesis, I created XSLT debugging program - my research paper is here
http://ww2.cs.mu.oz.au/~jbailey/papers/wise.ps
I miss it :) but now that I'm doing lots of web development & javascript, I do wonder that it is somewhat "slow" or "heavy" than other approaches.
Nonetheless, it was an awesome experience diving deep into the technology.
Those of you that know and love XSLT will understand when I say I wish I could be cast into an alternative universe where XSLT continued to be mainstream.
I'll continue to leverage this power-tool (and get paid handsomely by by clients for wielding such power) for the rest of my career.
However, I can say with certainty that I'm the only one within my team that shares this view.
Using it for transformation of XML to HTML, semantic validation (does my XML contain A and B when C='Individual'?) and mapping from one XML flavour to something JAXB friendly. Implementing much of this stuff in Java or other would be a nightmare and/or unproductive process with a LOT more code required.
Anyhow, I see most of developers pain with XSLT is because they never learned the basics of the language and thrash about ... perhaps thats the biggest fault of the language, in that its difficult to tell its story with an analogy to where people start from. Anyhow, learning any fp language tends to improve a developer as a programmer
People who hate XSLT are not usually fans of XML either.
>But even if you don't like the syntax, so what - there aren't any alternatives to XSLT.
There are hundreds. Simply use any normal programming language (javascript, python, perl, ruby, etc.) with a templating language (handlebars, jinja2, etc.).
The results will be much cleaner, easier to debug, and will remain cleaner as your transformations get ever more complex (code has a habit of doing that).
These hundreds of languages have no pattern matching but all have mutable state; XSLT is declarative and this removes a ton of complexity. If you had a bunch of related tables would you process them as Python dicts and arrays or rather load them in an RDBMS and write SQL instead? Especially as the transformations get more complex?
The 'exciting' thing about microformats was that it got you slightly higher positioning in SERPs. I don't know if that's true any longer. That was hardly a benefit of the language design.
I've never heard of components.
>These hundreds of languages have no pattern matching but all have mutable state
There are hundreds of bad languages. I wouldn't defend any of them. Most widely used languages have some form of decent pattern matching, though.
Mutable state can cause headaches, but if you want to use a functional language to help solve that - these days you're spoilt for choice. Most modern languages can be written in a functional style also, if that's your thing.
>XSLT is declarative and this removes a ton of complexity.
The OP made this claim by comparing it favorably to a javascript program that parsed HTML using regular expressions. Which makes me wonder (and shudder about) just what you and the rest of the XSLT defenders are actually comparing it to. And also why you seemingly haven't seen or used any good parsers or templating languages.
>If you had a bunch of related tables would you process them as Python dicts and arrays or rather load them in an RDBMS and write SQL instead
That would depend entirely upon what I was doing and why. I have done both in different contexts because they had different trade offs.
>Especially as the transformations get more complex?
SQL is capable of handling a great deal of complexity (more than xslt certainly), and its declarative nature certainly nets you a lot of things for free (less code, faster queries, no deserialization/serialization overhead), but it still has its limits. Yes, for really complex transformations I would probably write them in python and not SQL.
https://www.youtube.com/watch?v=2toYLLcoY14
It's basically custom HTML tags like:
<tabs>
<tab>Alpha</tab>
<tab>Bravo</tab>
<tab>Charlie</tab>
</tab>
and some JavaScript library ("a new type of library for the Web") that processes these custom tags to display tabs on a web page.I'm talking about pattern matching for templating rules; e.g. I want to handle a "value" token differently depending on whether the parent is a "field" or a "value-list" With XSLT I do this with a pattern: "field/value" vs "value-list/value". It's very clear and very efficient behind the scenes (about as speedy as a regex). I'm not sure I can do it in Python as easily.
As for parsers: parsers are a good thing, certainly, but look at what they do: they take an unstructured sequence of symbols and try to recover the hidden structure of it using a set of grammar rules. The end result is a tree that still contains the original sequence but dissects it into a hierarchy of discovered non-terminals.
Doesn't this look familiar? It's an XML document. XML documents are basically preparsed sequences or serialized parse trees. Maybe they have been parsed already by a specialized parser; or maybe the structure was added there manually when we marked up a text in a natural language; or maybe the sequence was generated by a machine in the first place and we just retained the original structure. What XML does is that it saves the result so there's no need to parse the same sequence again.
(Well, there is; we still need to deserialize the XML document and to interpret certain sequences as strings and another as numbers or dates and this is also parsing; but the serialization notation is much simpler than the grammars that were involved to parse or build the original tree.)
This is why parsing HTML is a waste of time regardless of whether we do this with regular expressions, recursive descent, Earley, LALR, Tomita, PEG, or whatever parsing technology we like :) It has been parsed already; we only need to deserialize it, and this is a trivial task that has been solved long time ago. (Well, maybe not with HTML, because somehow it's important to accept broken HTML and this may require a full-fledged parser and then some.)
http://www.youtube.com/watch?v=CjWYbcbpiWA
:)
<xsl:stylesheet version="1.0"
xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
<xsl:output
indent="no"
method="html"
doctype-system="about:legacy-compat"
encoding="utf-8"/>
<!-- Action parser that responds to HTTP requests. -->
<xsl:param name="action" select="'do.dhtml'"/>
<xsl:template match="/">
<html>
<head>
<meta charset='utf-8'/>
<title>Title</title>
<link rel='stylesheet' type='text/css' href='css/common.css'/>
</head>
<body>
<xsl:apply-templates/>
<script type='text/javascript' src='js/common.js'></script>
</body>
</html>
</xsl:template>
<!-- Make the document complete with div elements and classes. -->
<xsl:template match="*">
<div class="{local-name()}"><xsl:apply-templates select="node()|@*"/></div>
</xsl:template>
<!-- The 'id' attribute indicates a link. -->
<xsl:template match="*[@id]">
<div class="{local-name()}"><a
href="{$action}?action={local-name()}&id={@id}"><xsl:apply-templates
select="node()|*"/></a></div>
</xsl:template>
<!-- Retain the attributes (except if named "class"). -->
<xsl:template match="@*">
<xsl:if test="name() != 'class'"><xsl:copy-of select="."/></xsl:if>
</xsl:template>
</xsl:stylesheet>
CSS is then used to customize the presentation layer. For example, all the pages in the following website (click only "Next" and Safari is broken) are completely client-side XSLT, but use an XML document as a starting point:http://djarvis.bitbucket.org/xml/
The pie chart on the following page was also generated using client-side XSLT:
http://djarvis.bitbucket.org/xml/resources.xml
The only "trick" that was required was to load CSS (and some JavaScript) based on the current filename in the URL:
http://djarvis.bitbucket.org/xml/js/common.js
Most people write one XSLT page per web page to transform, but by using structured DIVs with classes that correspond to the XML element names, only one general-purpose XSL page is required.
This demonstrates the separation of content (XML) from presentation (XHTML/CSS).
I did once a less than 2k library in js that loaded the TTML markup in an hidden div and would show it on an HTML 5 video player with extra formatting.
Why would someone need XSLT in 2014? It was a failure in 2000.
xslt really seems to get acceptance from non-programmers who want it to perform a specific functions [like convert some hand crafted xml into html] and hatred from programmers who point out lots of problems with it as a language. I guess I would never even have thought of it as a language until I heard these complaints - I saw it more as a toolkit for, well, performing very specific functions.
Good article though.