Using XPath in 2023
denizaksimsek.com
denizaksimsek.com
I actually liked XML + XSLT + XPath combo a lot, and kept using them with my projects at Microsoft.
Thankfully it should be possible to emulate the whole thing with JS snippet, so old websites probably would work with minimal modifications.
I really liked it.
1/ To transform its Atom feed to pretty-looking HTML, if you access it from a browser: https://paul.fragara.com/feed.xml
2/ Similarly, to embed the feed dynamically on the homepage, using JavaScript XSLTProcessor (https://developer.mozilla.org/en-US/docs/Web/API/XSLTProcess...)
This works on all major browsers (and even on IE!) The XSL sheet is here: https://gitlab.com/PaulCapron/paul.fragara.com/-/blob/master...
In another context, I’ve leveraged XSLT 2.0, this time in a build process, to slightly transform (X)HTML pages before publishing (canonicalize URLs, embed JS & CSS directly in the page, etc.): https://github.com/PaulCapron/pwa2uwp/blob/master/postprod.x...
As to why URIs, I completely agree. I think URIs/URLs were introduced by TBL as arguably the major feat of HTML over generic SGML along with <a> and so phenomenally successful every W3C mechanism of name resolution and namespacing had to make use if it; RDF is littered with namespaces as well.
I once had the misfortune to have to use a third-party xml-based webservice whos developers thought that the xml namespace had to match the url used to access it. So their production and test systems had completely different namespaces. They maintained multiple xsd schemas for each system that only differed by namespace.
When I finally managed to explain namespaces to their "chief web architect" he got all embarrassed.
This was back in .net 4.x days, and I was generating serialisation classes from their schemas - so I had to create two different dlls and load the appropriate one at runtime. Very aggravating.
As an aside, I haven't used XPath in a few years, but use JSONPath / JSONPointer on multiple days each week, and highly recommend them. I find XML to be great for many things, but I find JSON much easier to work with.
I think in practice you can treat it as an opaque blob and if you want to just declare a namespace as "bar.foo.example.com" I think it would work just about everywhere. There's enough legacy, but still in use, namespaces that aren't full legal URIs that libraries will never be able to start enforcing it, and since there's effectively no utility in forcing them to be URIs anyhow I doubt it'll suddenly become an issue when all the libraries decide to become compliant. Given that it's a bit bizarre to have a URI to something that doesn't have to even so much as exist, I can't see the standards bodies every getting very fussy about this either. Backing it down to just being "some agreed-upon blob value" would be backwards compatible anyhow.
To be honest I think that just added to the confusion - the idea of using a URI that was just a unique name and not something that actually referred to anything wasn't the easiest thing to appreciate.
As an example of things that sometimes break:
* namespace redefinition: you have xmlns:ns1 on a parent node and xmlns:ns1 with different value on a child node.
* same namespace on sibling nodes (happens often with streaming): multiple child nodes have a xmls:ns1 with some value. That is only valid for that node.
* some processors expect namespaces only on a specific node (I think it was MS Navision ~ decade ago that failed when declaration wasn't on a specific node, but on a child node)
raw = Nokogiri::XML(File.read(uploaded_file).gsub(/ xmlns="[^"]\*"/, "")) { |config| config.noblanks }The annoying thing is namespaces are not that complicated. They're probably five times more complicated than they need to be because you're always dealing with some library or system's half-assed bizarre implementation of them instead of the rather simple spec.
My opinion is that XML would be much better without attributes. A lot of things would be simpler. Attributes are nice, but they add whole another dimension of problems everywhere.
Tags are always in the empty string namespace unless an explicit namespace prefix is given, in which case the tag's namespace in the one corresponding to the prefix.
That's it. That's all there is to namespaces and attributes. The idea is that the namespace would be redundant to the namespace on the tag, because the attribute's meaning is already 100% dependent on what tag it is on anyhow.
The big key to working with namespace documents in general is you have to understand that the names are now tuples (namespace, name), and that the string prefix you may see like "ns1" is an abbreviation, not the namespace. The namespace is the thing declared as the attribute of the xmlns tag. Get that right and most of the rest follows. What almost everything gets wrong is failing to treat names as (namespace, name) correctly, and trying to treat names as a single string like in non-namespaced document, and then writing in whatever misconceptions their particular implementation wrote in to get around the fundamental impossibility of handling a 2-tuple with an internal datastructure that is only a 1-tuple.
This is at best misleading (or worse: misled—and wrong). The semantics of the following snippet is completely different:
<foo:bar baz="xyzzy"/>
... from this: <foo:bar foo:baz="xyzzy"/>
(Assume the "foo" prefix is defined and mapped to a namespace.)This is before even getting into the fact that unprefixed elements will belong to the default namespace (despite being written with no "explicit namespace prefix").
> The namespace is the thing declared as the attribute of the xmlns tag.
The "xmlns tag"? No.
Have you ever seen tag name clashing in a XML document?
If I had to guess the chances of there being some element names clashing without namespaces would have been fairly high.
As the article states, the only tag that can't be searched for via CSS selectors is the new "hx-on".
And my initial thought about this is (which I'm open to reconsider), it seems like you probably shouldn't go looking for elements based on their event model. Show me every <div> with a certain value in its class attribute? Yes of course. But find me every <div> with a certain behavior on mouseUp? This isn't something you'd do in an OO language, you'd look for objects with a certain property value, not a specific event handler, which would likely be private anyway.
I'd also note that, for the same reason, standard HTML as well as React, etc have the same inability as HTMX here. The article notes that with HTMX you can't use CSS selectors to find "hx-on(something)" tags, the same handicap prevents you from searching for any element with an "on(something)" attribute, i.e. onClick, onMouseUp, onKeyPress etc.
Maybe I haven't thought through this entirely but I don't see any problem here.
(edit: to be clear, of course you can CSS-select a specific tag like onClick, and you can also select a specific HTMX tag like hx-on:click. What the article notes that you cannot do is search via wildcards or "starts-with" in CSS, i.e. hx-on:*, but you can't do on* either)
How does XPath deal with XML namespaces?
Generally the next stop after xquery, for me, is text mining, either on R+Python or on Orange ML. If a miner doesn't cut it, then LLM shenanigans.
Also, xpath? It's pretty great. XQuery? Does the job. XSLT? Ok, so NOW that's the feeling of a panic attack. I've been doing XSLT for literal decades, and I still don't know what I'm doing when wrenching on a giant pile of FOP generating funhouse madness. When I am tagged into a data transformation job, I always stress that xquery is the right tool, rather than a confounding nested directory of XSLT using different parsers and different passes like a figure-8 interstate off-ramp. For FOP, though, there's really just one game in town for that. Although, having said that, me and a bunch of others are doing our damndest to show that what you're trying to do with XSLT/FO can be done way way way easier with CSS and Paged Media (either via Paged.js or Vivliostyle or any of the other zillion PMM implementations). The downside is you have to wrench some CSS yourself, but honestly, that's probably going to be easier than wrenching on DocBook-XSL or the DITA-OT or one of the MIL-STD XSL packages.
It’s the problem from having built XSLT out of XML, that was completely unnecessary.
And then with XSLT being so verbose, having cheaped out and made the current node (“.”) so implicit. And the confusion from the dual use of templates as both match/patch and function constructs (it would have worked fine as a shortcut, but it makes grokking how things fit much harder than necessary).
I wouldn’t say jq is really comparable to xslt though, xslt has much more transformative flexibility. It’s closer to xquery.
jq isn't structured around data schema transformations, it's true, so writing complex transforms is not as easy as with XSLT, but it is very much possible. Essentially it's a reduction over [possibly paths to] elements of interest in `.` updating the reduction state with the new schema. You can organize your code into functions that do much what XSLs do.
Right, so here's the secret decoder ring of XSLT: Underneath all the complexity, it isn't doing ANYTHING you can't do in your language of choice armed with an XPath library. And it is often incapable of doing even some rather simple things you can do in your language of choice armed with an XPath library.
XSLT is just a terrible programming language. That's all it is. All of the magic is in the XPath part; once you've selected the nodes you're working with, XSLT is a horrifyingly awful way of manipulating them into doing what you want them to do.
XSLT is the intersection of the worst parts of declarative programming with the worst parts of functional programming, wrapped up in one of the worst ways of serializing a programming language. What confuses some people even to this day is that they see "declarative programming" and "functional" and even "standardized serialization" and accidentally credit XSLT with the benefits of such approaches, and then if XSLT doesn't work they blame themselves for failing declarative functional programming in such a wonderful serialization format. They're wrong. It's XSLT failing them, whose origin is also people who thought if they just create something declarative and functional and serialized through XML they were guaranteed to be producing something good because those things are just so Platonically good on their own that they couldn't possibly produce something useless and broken, so it was not necessary to analyze the resulting abomination to see whether it actually fulfilled the goals, because it simply by definition fulfilled the goals by virtue of being declarative and functional and in the bestest serialization ever.
Perhaps there is a declarative, functional XSLT-inspired language that could be written that would be as good as the people bedazzled by the buzzwords think XSLT is. (Though there's no world where such a language is helped by serializing into XML; serializing a language for manipulating XML into XML is actually the worst choice possible because of the nested encoding you inevitably produce!) However, in the meantime, you don't really need to wait around for someone to produce it because it turns your favorite general purpose language equipped with an XPath library is already 90%+ of the way there.
It’s not so much that XSLT is a terrible langauge, it’s a terrible syntax. At the end of the day, it’s just a really awkward way to express some pretty basic functions. With a bit of care, it’s easy to imagine transforming arbitrary XSLT to any arbitrary language expressing the same semantics. This is especially true because XSLT itself is so limited in capability.
Please don't tell anyone I said this, but the hardcore XML people can be a little culty. At untold meetings, I'd walk through all the functionality that they're getting - from a staggeringly gigantic XML system, priced in millions of dollars - and demonstrate, one bullet point after another, how the same functionality is available from a few lines of Asciidoc/ReStructuredText/<insert_your_tech_here>. Then, at the end of everything, I always have to contend with the last one: "But it's not XML". To which I shrug my shoulders: I can't solve existential problems. Why is this specific technology so important? My theory is that there's some very deep and possibly poisoned incentives at work here, which I don't really want to unpack because it would make me sound nasty, but the most innocent of these is the Sisyphean drive to impose labels on the chaos of business and on natural language. I think it might be a neurosis of the industry, but I've seen it outside of aero-def-landia too. Who knows? Maybe someday I too will rue the day I ever turned from the True Faith.
JSON is a pain in the ass for organizing and parsing structured data, particularly arrays of things.
It's so much more verbose and disassociates an object's definition from it's type. Let's say you have parents with children with names.
<parents>
<parent name="Bob">
<child name="Alice" age="12" />
</parent>
</parents>
The root object is 'parents', so you know a bunch of 'parent' elements are going to follow. When you read the <parent> object, you also know it's a parent because the tag say so. You don't even need a <children> element because it's fine to have a list of <child> elements directly after <parent>.Now here's what I think a typical JSON equivalent would be:
{
"parents": [
{
"name": "Bob",
"children": [
{
"name": "Alice",
"age": 12
}
]
}
]
}
Ok, so whitespace aside, it's less verbose, but look at all the info that's missing.What "type" is the root object? You'd say "parents", but how did you find that out? You have to know _a priori_ that a field called 'parents' would have to be there. Not a big deal on the root object because it's usually special, but how about a single parent?
Look at the 'parents' array. The only thing hinting at the fact that { "name": "Bob" } is a parent is the fact that it is part of an array, that's attached to the 'parents' field of the parent object. You have to do 'upwards' lookups to determine what this object is. The object itself doesn't have that information. Same thing with { "name": "Alice" }. How do you know that's a child object? You don't. You have to do an upwards lookup.
Now you might say "just tag the elements with their type so you can keep track of what these objects are". Let's try that:
{
"type": "parents",
"parents": [
{
"type": "parent",
"name": "Bob",
"children": [
{
"type": "child",
"name": "Alice",
"age": 12
}
]
}
]
}
Sweet, now we're reaching data representation parity, but if you get an object like this from a third party service, how do you validate it without kicking in the logic to process each element? You'd have to have a 'dry run' version of your logic.In fact, how do you formally describe the structure of these objects to another service so that the service could guarantee that it is only generating valid objects in the first place? XML Schema was a solid solution for that. JSON Schema had no support anywhere last time I looked at it. Where is it now? It looks like you could use it if you wanted to, but afaict most people generate fly-by-the-seat-of-your-pants JSON objects in code and no 'formal' validation is happening, other than the reply code from the service when the object is actually sent (if you think about it, that's just "testing in prod")
I think version 1 of XSLT, XPath and so on where pretty simple solutions to working with structured data, but people went overboard with trying to shoehorn XML into solving problems best suited for imperative code, so you got XSLT 2.0 (want for loops? no.), XQuery and XPath 2.0 abominations, various weird xlink solutions, imperative code in tags: <script>function foo() { }</script> which introduced a second syntax, and so on.
I understand why the XML world of nonsense had to be stopped, but we threw the baby with the bath water.
Don't get me wrong, I like JSON, but I also feel like we collectively took a step back and opted for 'the javascript of structured data representation'.
Maybe the rise of thick, JavaScript heavy clients had a lot to do with it? XML was never easy to work with in JavaScript, which is a shame, seeing as it has the same roots as HTML. I blame the DOM API - it has always been tedious to use in any language that implemented it.
ActionScript had that nice built-in @ syntax for selecting nodes and first-class support for XML in the language (E4X?). How we killed that first-class language support only to turn around and rediscover it in half-baked form as transpiled JSX is beyond me.
</rant>
A XML-lite that ditches all the noise and stupidity, but keeps the annotated text with user-defined sum-typed tags, and a type specification would be better than any standard we have today.
XML's creation (or definition) in the early 1990s was motivated mainly by a desire for an SGML-lite.
For those not familiar with the promise design controversy:
http://brianmckenna.org/blog/category_theory_promisesaplus
https://github.com/promises-aplus/constructor-spec/issues/24
EDIT: Thanks, sibling!
See https://devhints.io/xpath & https://devtoolstips.org/tips/en/evaluate-xpath/, and in this thread: https://news.ycombinator.com/item?id=36755041
I'm looking forward to those new things in htmx because it seems like it extends what we know of html. And I can't yet quite say if that is going to stick (by being integrated into IDEs, work with frameworks easily, be the go-to way to do this and that, etc). Time wil tell, but it's cool to see new proposals to improve the status quo out there.
sure, but more powerful and thus more able to handle situation than depending on css selectors.
Also consider Robula+ https://tsigalko18.github.io/assets/pdf/2016-Leotta-JSEP.pdf
https://tsigalko18.github.io/assets/pdf/2014-Leotta-ISSREW.p...
I wasn't recommending those either. roles or testids are much more robust than xpath and selectors.
Robula+
I had a very brief glance over the paper and while it looks a lot better than plain xpath it still has some of the problems that I think make tests flaky. One of the suggested examples is "[contains(text(),'123456789')]" where '123456789' is an input field with that value. If the example text is changed the path will break. It's testing the content of the field, not what the field is* or what the field should do. An alternative would be a selector like 'getByRole('button', {name: 'Telephone'})' (getByRole comes from Jest, but more testing libraries have something like it) means you can change practically everything about the input field (xcept the aria-label or aria-labelledBy value and the test won't break. And, as a bonus, your app will end up more accessible.
$x("//my/selector")
Very handy for testing XPath expressions for Selenium or Playwright tests, if there is no data-testid attribute.We use XPath at work from CSS with custom plugins, and use XPath in JavaScript for targeting elements that otherwise wouldn't be straightforward to select with CSS.
Another cool thing XPath does is have awareness of the text content of elements! //li[contains(.,"example")] would target all <li> elements with text content containing "example".
A huge part of XPath’s power though, and something you AFAIK can’t do in browsers, is extensibility.
For instance selecting an element on the basis of a class is absolute hell in XPath 1.0 (and not great in 2.0 either, XPath 3.1’s `contains-token` finally made that not hell). But server-side you don’t care because pretty much all implementations allow installing your own functions so you can add your own `contains-token` or even `has-class` predicate and be on your way.
I have never experienced that again. :(
I'd take it over YAML any day. And I speak from a YAML parser implementer perspective. Making a 1.2/1.3 YAML parser is hard. The number of edge cases and states is absurd, while the specification needs to be on HTML/XML level.
E.g. When I started working on YAML parser I didn't expect to be fixing errors in spec. But here I am.
And in terms of XPath, for folks using a JS stack, fontoxpath[2] supports a DOM facade adapter interface which allows for querying any arbitrary tree-like structure, so it could certainly handle the same use case.
1: https://tree-sitter.github.io/tree-sitter/using-parsers#patt...
we wrote a small package[1] (using pugixml) to transform XML to JSON using a custom Xpath template syntax. Make our job much easier.
<body>
<div><!-- you want this one -->
ok
<div id="the-child"></div>
</div>
<div><!-- you dont want this one -->
nok
<div id="the-wrong-child"></div>
</div>
</body>
css: div:has( > #the-child)
xpath: //div[@id="the-child"]/..
isnt that kinda the same? I think I'd even consider the css version to be easier to understand at a glance.If you've got full control over the Dom it'd be much better to just set an id attribute on the element in question though, as getElementById is still the best performing option
css:
div:has( > \* > \* > \* > \* > \* > \* > \* > #the-child)
(backslashes added by hn because of formatting conflict)
xpath: //div[@id="the-child"]/../../../../../../../..Now, everytime I've to scrape data from XML from the command line, I wish I had a tool to translate XML to json so that I could use jq on it, instead of having to wrangle with xmlstarlet.
> it addresses a specific shortcoming of HTML: the on* attributes do not support arbitrary events
> it is much less featureful & useful than, for example, @Alpine_JS, but if you have very basic needs for event handling it can help avoid another dependency
Yes, I saw that. But what does it have to do with XPath?
The author of htmx wrote a rationale (when and when not to use it). https://htmx.org/essays/when-to-use-hypermedia/