XPath is actually pretty useful once it stops being confusing
news.rapgenius.com
news.rapgenius.com
XPath : XML :: regex : text
(But along with relational algebra, they are among the few abstractions that work really well)
And that is what I've found most people who have trouble with it don't understand - like what exactly following-sibling or child means.
I spent about 2 months writing my own xpath evaluator once and it gets so much easier (to implement too) when you understand this is just a tree-traversal with an iterator following the axis.
Unfortunately the axis syntax makes it very verbose to read.
If you think 11 lines of code is a lot, you're overly focused on concision at the expense of readability. I've never (read: never) worked on any Ruby code, yet I find the posted example more readable than the supposedly more valuable xpath.
At the very least, they're the same. If you're writing code in Ruby, 11 lines is nothing. If you're writing code in Ruby and xpath is used nowhere else in the project, that single line of super-compact xpath might as well be 1000 lines of Ruby -- it doesn't matter.
If you're trying to compact 11 lines of code you're probably doing it wrong.
for this particular task, XPath is actually considerably slower than the pure-Ruby implementation. Interestingly, that's not true if you take out the <br> part and only look for text at the beginning of paragraphs. My guess is that the following-sibling axis is the culprit, since it has to select all the following siblings of the br tags, and then filter them down to only the first sibling.
I was hoping selectors were lazy, in which case, selecting all the following siblings but then immediately filtering that selection down to the first would be cheap. Lazy or not, can there really be no efficient way to do the equivalent of jQuery next()?
Just make sure that you document your test cases (i.e. what you should match) or your colleges will hate you.
On its own, .NET's XML libraries are really only good for consuming XML documents, but even that is a rather painful experience, especially as it forces a namespace on all documents, complicating the XPath expressions necessary to query it. Actually authoring documents is a nightmare. My XmlEdit project makes it almost as simple as key-value-pair config files.
https://github.com/capnmidnight/xml-stuff/blob/master/README...
However, just take note that the main concept of the library is "make it work". The idea was that, given an XPath expression with several attribute selectors, it would fill in any necessary nodes to just make it happen. So you can technically chain a ton of editing commands together, by using an appropriately complex XPath expression.
It came out of a need to repair thousands of broken XML documents. It's probably not very complete. It was written for one project--and though I took time to make it generalized--it didn't make it into a key role into any other projects; I just didn't ever again have the need to deal with XML documents on such a scale.
It's actually one of the first "big" things I wrote out of college. I'm not too happy with some of the design right now, but the functionality has held up over the years and it's not as shitty as some of the other code I wrote at the time. I guess I knew that a lot of the project was hinging on how easy it was to write XML documents, so I made sure I did a ton of testing to make it work.
Is there any reason to store the HTML version with <p>s and <br>s instead of a plain text and converting it to HTML with simple rules à la markdown? (single line break = <br>, double line break = <p>)
In every single task I do that involves munching on XML with xpath, the largest timesink is figuring out how the ns needs to be set up. (Looking you you, PHP simpleXML).
And more generally that's true of every single task involving muching on namespaced XML. Namespaces are a good idea implemented absolutely terribly.
XPath is a good idea well-implemented (no, XPath 2 does not exist, there is only one XPath). One of the few I've found in XML-land. I still hate that we have to use CSS selectors rather than XPath (although that's understandable considering CSS selectors predate XPath), most of the improvements since CSS1 were in XPath day 1, and the rest (pseudo-classes) could probably have been implemented using functions.
Also, that might have finally gotten us a non-eye-stabbing standard function for "match any item of a space-separated list in an attribute" (matching HTML classes in XPath without custom helpers is the worst)
CSS handles this rather nicely:
[class~=foo]
https://developer.mozilla.org/en-US/docs/Web/CSS/Attribute_s... [contains(concat(' ', normalize-space(@class), ' '), ' foo ')]
the normalize-space can be dropped IIF you're certain all spaces are normalized, the spaces around the needle not.xpath 2 does quite a bit better through `tokenize`:
//*[tokenize(@class, '\s+')='foo']
but still not great. And god forbid you need to match multiple classes in the same selector.[0] or xpath 1 + exslt if your xpath implementation provides it. exslt actually does slightly better as the pattern is optional and defaults to whitespace characters
I don't think it's a bad idea, most of the improvements in XPath 2 are the new standard functions which depending on your XPath implementation may be available as extensions (e.g. tokenize comes from exslt, an xpath 1.0 library) but along with that it brings significantly higher complexity and I think the spec has gone from "difficult to read" to "meaningless word-salad".
I really like XPath, but I can't say I was impressed by XPath 2, it loses much of xpath's simplicity with little to show for the added complexity.
doc = Nokogiri::XML(stream)
doc.xpath("//x:foo", {x: "http://somenamespace/"}).each ...
I always end up declaring the namespaces in a constant and passing it in, because there is no way to specify your mappings globally. It should have been something like: doc = Nokogiri::XML(stream)
doc.namespaces = {x: "http://somenamespace/"}
doc.xpath("//x:foo").each ...
or at least something functional like: doc = Nokogiri::XML(stream)
doc.using_namespaces(x: "http://somenamespace/").
xpath("//x:foo").each ...
I have seen this way of integrating namespaces and XPath in several libraries, it's not just Nokogiri.Is it really that bad to have 11 lines in Ruby?
Initially I didn't get the wrong tool part but after reading it all that did make more sense. I haven't used XPath more than a few times and they were pretty simple so can't complain. Just something I'll have to keep in mind.
http://hackpackers.lonelyplanet.com/2013/03/05/XML-Transform...
http://ofps.oreilly.com/titles/9780596155957/HerdingXMLInSca...
But as far as the OP, this seems like a case of worrying about the code instead of the data structure. This would be easier to address before the lines are transformed into HTML. Which I assume is not how they are stored.
Then there was XSLT, which was a pretty sweet way to turn a data format into a variety of "print" or "display" formats. Definitely been replaced by bigger and better things but it's a pretty awesome technology that does one thing really well.
Programming in XML is never a good idea. It isn't in XSLT, it isn't in Spring, it isn't in Maven. Anything that's XML and has elements or attributes with names like "if", "else" or "while", something went horribly, horribly wrong somewhere. It's horribly verbose, you can't reasonably debug it, and there's virtually no engineering best practices, which results in near-impossible maintenance tasks.
Any modern programming language with a good, consice, XML parsing library is a more effective tool to transform XML into something else than XSLT.
Don't code in XML.
Just curious; it's just syntax, after all, and I think that XSLT really has an advantage in transforming XML to other XML (or text) compared to a program in other languages.
It still can't turn a pile of XML into a pile of PDFs though, so XSLT is definitely king in some arenas.
There are a few issues I am having with my markdown editor, in comparison to githubs markdown support.
I'm not saying it's not useful. Actually, I believe that if you only have one use-case, then using xpath might be overkill because of all the added-complexity of maintaining a new library/technology/ideology. But if it's the sort of domain that xpath would be useful more than once, then sure use it.
It's conceptually the difference between:
count = 0
for row in db.query("select id from <table>"):
count += 1
and db.query("select count(*) from <table>")
I also don't find it at all confusing, you just have to understand the tree nature of XML. from lxml import html
doc = html.fromstring('<html><body><p class="text"></p></body></html>')
doc.xpath('//p[@class = "text"]') == doc.cssselect('p.text')
# or
doc = html.fromstring('<html><body><p>1</p><p>2</p><p>3</p></body></html>')
doc.xpath('//p[2]')[0] == doc.cssselect('p')[1]
# Note: My only annoyance is that .xpath() always returns a list, even
# when you know that it will return only a single item."hxselect - extract elements that match a (CSS) selector"
CSS is often a simpler way to extract data.
Pro Tip: the chrome inspector lets you right-click on an element and get its xpath.
Pro Warning: sometimes the xpath generated by chrome doesn't work when scraping with Nokogiri. I'm not sure why yet, I've just learned not to rely on it.
Describing an element as "the first child of the fifth child of the second child of the first child of the eighth child of the second child of HTML" is as much the right path to an element as if you described the way to your house as "Walk past the park then walk past the bus stop then walk past the hardware store then walk past the butchers then turn left then walk past the pizza shop then walk past the library"
And I'm not disagreeing with you–I'm only saying Chrome has this feature. I don't know what route they choose for you but I know they don't always work in tools that parse (HT/X)ML.
Here's a screenshot in case you don't believe me: http://imgur.com/9FZSMSt
The XPath Chrome returns for this page is: //*[@id="details"]/article/table[1]
For example, you could use TEI XML (http://www.tei-c.org/index.xml), and then use stanzas and lines. Then when you go to render your lyrics, you can capitalize the first letters in your presentation code.
- Why are the namespaces there in the first place?
- Do I really not care if the element is found in a namespace other than the one expected?
- Does my host environment have a way to specify the namespace of the element I want to find (hint: it probably does)?
- Is the reason that I want to remove the namespace that it's actually something I need to do or is it that I am ignorant of the method for specifying namespaces in my host environment?
And the answer is usually "no".
I have seen SOAP responses with 20+ namespaces, all of them being essentially implementation details -- every different section of their internal API getting its own namespace. Inevitably, the elements are also prefixed in a way that makes them distinct, or wrapped in a distinguishing element (i.e. Contact/NameInfo/FirstName rather than FirstName xmlns="contact-name").
In situations like that, your best case scenario is that you do the grunt work of setting up aliases for all the namespaces, putting them into your XPaths, and you're done. The worst case scenario (which I've encountered) is when a version update of the API changes the URIs for half the namespaces, even though the structure of the data hasn't changed. In a case like that, you're actually penalized for doing the 'right' thing and not just stripping the damn things off.
Long comment short, I agree. CSS selectors are easier to understand and read.
The `/` in an XPath expression is probably a better match for the space in CSS selectors.
I'm sure I'd have fun coming up with an XPath solution, but for me, the ultimate goal is maintainability. If I wasn't 90% sure that the next person to look at that code already knew XPath, then I'd go with the Ruby solution.
Dealing with 11 lines of code in a language you know is better than dealing with 1 line of code in a language you don't (which ends up forcing you to read 1000 lines of documentation and examples to understand it).