Open Letter to XSLT Fans
snoyman.com
snoyman.com
For years I struggled to create 100-line XSLT files; it looked impossible. Then I realised what the problem was: I was trying to convert the source tree into the result tree in one big step. XSLT can do that, but it is not meant to do that and it is just bad software engineering. What I should have done (and I started to do) is to divide the transformation in smaller uniform transformation steps: for example, in one case I ended up with this series of XSTL files, each one taking care of a single step in the transformation pipeline:
* cleaning (remove duplicated and empty nodes);
* completion of partial data (where I go, fetch missing data and use it to make the original data more homogeneous);
* grouping and reordering (twist the data around to make it more similar to the destination format);
* data transformation (where the individual pieces of data are transformed into separate nodes in the destination format);
* layout integration (fill the holes in the layout template with the missing data).
Each file ended up being very small, readable and easily maintainable. The much shorter xpaths used in the templates are a clear indicator of how much better that project has become. Before the split I was doing things like fetching the missing data while putting the already transformed data in the layout template; that was forcing me to transform the newly retrieved data in place before being able to use it. It was a mess. It was a mess because I was not following one of the underlying and unspoken assumption of XSLT. Isn't it the same with any other language? Will an Haskell-based tree transformation language prevent me from writing huge single-file many-purposes-in-a-single-function spaghetti code?
That said, I do not like articles like this that do not explain why they do not like something and do not give examples of what is wrong and how it should be instead.
If you're writing an API that works with files (not an over the wire format), and you don't want to be locked into any one language or toolset, XML is pretty hard to beat. Being able to write just one XSLT file, and have it processed by any compliant processor is pretty impressive - no other templateing format works that way.
The problem for most people is that they never touch those other technologies that make XML great, because they're looking at basic data serialization, and don't care so much about validation and transform. Most web technologies fall into this category. Any tool can and will be abused - how many of us have seen a monstrous unmaintainable Excel file that performs a task that would be much better done in a scripting language?
TL;DR: XML tech has it's uses, but it's very frequently abused.
...there are a million oft-argued reasons why not, but that's one line of thinking that will lead you to do it.
These discussion always remind me of my favorite old Slashdot quote: "XML is like violence. If it's not solving your problem, you're not using enough of it."
No dynamic server-side templating language, easily transformable to other XML formats (docbook, atom (see front page of registry)). If the user/client has libxml, the processor command is already installed. The specifications are now provided in an application specific format that is language-agnostic instead of manually maintained ad hoc HTML. The extension registry index page is now automatically generated into static HTML instead of manual HTML+JavaScript for sorting the lists.
Finally, the restricted power of XSLT is an advantage for this application vs PHP, ASP, or JSP. If browsers actually implemented web technologies like XSLT, sending XML specs and generating HTML on the client side with cached XSLT would save bandwidth.
Note: You can generate HTML using JavaScript but it is not needed. You can manipulate the DOM to change what the user sees.
The other thing I like about XSLT is that it creates a hard code-boundary. But I hardly think its syntax makes that worth it.
In the same way ColdFusion is "an incredible single-purpose language for translating databases into web pages": it's not, the difference is that XSLT has no alternative whereas CF has a billion.
> and there is absolutely no substitute for what it does.
That's the issue, XSLT is not "incredible", there's just very little alternative.
It's also terribly opaque, and I'm afraid the community isn't very helpful. As a n00b I've worked with Apache, Bind, Python, bash, JavaScript, XSLT. None were harder to learn than XSLT. I've _used_ this tool, and it still takes me hours to figure out how sheets I've written actually work. The documentation is awful -- I have some of the community leading texts right here on my shelf, they are almost worse than useless. Appeals to various fora were as unhelpful as any I've made anywhere.
There's little point complaining about any of that, that community doesn't owe anything to me or anybody else. But after investing weeks learning that tool, I've moved on and use other stuff in places where XSLT should be the answer. And I'm by no means surprised to find other people expressing frustrations about the thing.
(XSLT does actually have a standard, syntax checking and built-in context-sensitive quoting)
lmao. Thanks, hadn't seen that one.
One trouble is that you know there's a cliff out there. There are some simple tasks that can be done with an XSLT that's just beautiful. But try to change what it does, and you reach this point where it becomes incomprehensible.
Back in the day we used to wonder if XSLT was Turing complete -- some guys wrote a paper and proved it, but that's the problem with XSLT. If it takes computer scientists half a decade to figure out if it's Turing complete or not, it's completely incomprehensible.
It's probably not quite a functional language by most definitions, but it has a lot of features common to functional languages, including immutable variables, no side effects, XSLT 2.0 does have first class functions, recursion, etc. The Saxon implementation of XSLT 2.0 even has lazy sequences (I'm not sure if that's a language thing, or just that implementation).
Without knowing what the OP is having to deal with, I can't really evaluate his complaint. I've certainly seen some hideous XSLT, but it doesn't have to be hideous, and it's great for the (pretty narrow) domain it was intended for.
You can abuse any programming language—I don't think XSLT is particularly special in that regard.
All joking aside, I like XSLT just fine, as long as it's use is limited to exactly what it's good at: transforming XML documents from one schema to another. Yes, it's verbose if you write it by hand, and yes it would be nice to have something with all the power of XSLT and with a lighter syntax, but XSLT is hardly the worst thing around.
That thing you are looking for is XQuery (http://www.w3.org/TR/xquery/) and it is fantastic.
Don't be fooled by the name: it can produce content, too.
Blind hate is hardly something desirable. I'd say, yes, even hate needs an argument.
edit: now that I've looked at it I remember seeing it before. I thought it was some crazy custom language invented by the authors of a script and it sent me running. Good to know that that's XSLT
"Oh, and the fact that you can call a language functional when it lacks first class functions makes my eye twitch. I'm tempted to upload a video of my eye twitching just to prove it."
Anyway, he's the initiator of Yesod, a screaming fast, type-safe web framework in Haskell. Some one called him the 'type safe version of DHH' last month :)
Anyone seen any good schemes?
I've been meaning to look at HXT, I've read it had something like that, but I have not needed to transform XML in a long time so it's fallen by the wayside. On the other hand, TFAA qualifies himself of "Haskell programmer" and does not use HXT so maybe it's not that good.
An alternative I've thought about (but not implemented on grounds of having absolutely no need for it these days, as noted above) is implementing what I consider the good part of XSLT (tree transformation via template matching through XPath selectors) in Python on top of lxml. Something akin to Flask, where the app would be a group of templates, and the routing would be a sequence of XPath assertions (instead of http PATHs + methods). Along with a few helper functions or methods (to easily recurse into the rest of the tree), this ought materialize most of XSLT's strengths in a general-purpose language (making extensibility trivial), and template groups would improve modularity significantly.
But I have to wonder if Python is the right tool for the problem. I get the sense that the XSLT transform engines are deployed to handle really big documents, and I wonder if a Python based tool could compete on speed with xalan or saxon.
The reason I prefer to work in non-declarative code land is that I usually must inject many service references into the translators. When converting an industry standard XML format into a company's internal domain model for quoting insurance policies, I had to employ a set of heuristics to create a valid policy from a set of coverage requirements which likely were ill-specified. For example, we didn't offer a $750 auto deductible. Should this be converted (with a note attached) to $500 or $1000? This decision varied by state, policy type, etc. We had a metamodel which I injected into the transformers at the points where such decisions were made.
That the source and/or target of a transformation are XML is a red herring, though. Most of my time is not spent on XML-ness itself but on solving fundamental impedance mismatch issues when converting between two different domain models, sets of assumptions, etc. Document formats don't matter for these problems although those formats with better surrounding toolsets certainly allow one to concentrate immediately on the part of the problem which is the hardest. I even prefer talking about these problems using terms like "model conversion" instead of "document translation" -- too many marketing folks have convinced IT managers that, through the magic of their overpriced ETL tools, "document conversion" problems are a trivial drag & drop matter.
One way I've thought of architecting these model transformations is through the invention of a few intermediate model definitions, each one becoming less source-like and more target-like. I think some stages of conversion are more compatible with declarative approaches. Perhaps attempting conversion in only a single pass has led me to throw the baby out with the bathwater w.r.t. declarative schemes?
For those of you who don't do corporate IT development, the sad reality is that a huge percentage of development effort is spent on data conversion/translation between systems. The ratio of glue:substance is highly skewed toward glue. Furthermore, the ratio gets worse as short-term benefits are prioritized over long-term ones, development is silo-ed between business groups, and data modeling takes a backseat to gettin' stuff "done".
Early browser bugs probably caused more harm than good for XSLT adoption. XSLT was available in browsers almost from the beginning, more powerful (at the time) than any other client side (javascript) solution. Too bad it didn't get off the ground some more, it might have actually changed the way we work today.
Yes, it's too verbose, and probably doesn't deserve the title of "programming language." But, it's a pretty powerful idea if kept in the scope for what it was invented; transforming one schema to another, with options for transformations occurring on the client or server side.
I found it extremely powerful.
I have never used it again, because the power it afforded me was not worth the heartache.
For getting simple data up on a web page it was fantastic. But for more complex stuff it got wickedly complicated to maintain at an exponential rate. XSLT just didn't scale (and/or we just weren't good enough at it to make it scale).
PAIN. That's what it is. Just PAIN.
A (difficult!) year of daily use behind me, I've found that all the above problems disappear. People come to XSLT expecting a procedural DSL for modifying an XML file which they can pick up in a day or two, and very quickly their code spirals out of control as they try to twist XSLT into what they expect. It very clearly isn't a procedural language, which is why I think it gets labelled as functional, because it doesn't accurately fit either model.
Once you develop good code style and know the pitfalls to avoid, it's perfectly pleasant to use. It's also significantly easier to get along with if you use a full blown XSLT IDE like oXygen, which allows you to breakpoint, see the call stack, and trace back which template generated a specific part of the output.
For some reason, even on teams with good engineering practice, all discipline goes out of the window when people write XSLT. So a dev who might write beautiful maintainable code in any other language suddenly reverts to the worst spaghetti style on XSLT with 1000 line functions, impenetrable naming, nary a comment to be seen.
I know. I've been that dev.
I think the fundamental issue may be more to do with a flawed distinction between read "code" and just a "stylesheet / data transform"
Sincerly, someone who understood that programming lanugages are tools.
At the time the only browser that didn't support it correctly was IE, but after some small work arounds it worked pretty well. The downside being that inserting ad networks didn't work because there is no document to .write() to being XML, and the fact that you had to host a valid DTD and your document would get verified against it each and every time.
The other big issue was that you have to write a lot of boiler plate code to get it to do what you want it to do, with a lot of recursion and the like.
I do not miss it.
it is awesome as a cross platform solution for transforming xml into some other stuff. for anything else - why on earth would you use it?!
i really enjoy playing with it, generating complex html out of xml. things like grouping with the muenchian method are the defintion of nerdy fun for me. to this day i have no idea what went on in steve muench's head to come up with that small snippet of code. i just love him to death for it. i pretty much build a career on top of it, as a lot of other coders never groked it.