Htmlq: like jq, but for html
github.com
github.com
For reasoning about tree-based data such as HTML, I also highly recommend the declarative programming language Prolog. HTML documents map naturally to Prolog terms and can be readily reasoned about with built-in language mechanisms. For instance, here is the sample query from the htmlq README, fetching all elements with id get-help from https://www.rust-lang.org, using Scryer Prolog and its SGML and HTTP libraries in combination with the XPath-inspired query language from library(xpath):
?- http_open("https://www.rust-lang.org", Stream, []),
load_html(stream(Stream), DOM, []),
xpath(DOM, //(*(@id="get-help")), E).
Yielding: E = element(div,[class="flex flex-colum ...",id="get-help"],["\n ",element(h4,[],["Get help!"]),"\n ",element(ul,[],["\n ...",element(li,[],[element(a,[... = ...],[...])]),"\n ...",element(li,[],[...]),...|...]),"\n ...",element(div,[class="la ..."],["\n ...",element(label,[...],[...]),...|...]),"\n ..."])
; false.
The selector //(*(@id="get-help")) is used to obtain all HTML elements whose id attribute is get-help. On backtracking, all solutions are reported.The other example from the README, extracting all links from the page, can be obtained with Scryer Prolog like this:
?- http_open("https://www.rust-lang.org", Stream, []),
load_html(stream(Stream), DOM, []),
xpath(DOM, //a(@href), Link),
portray_clause(Link),
false.
This query uses forced backtracking to write all links on standard output, yielding: "/".
"/tools/install".
"/learn".
"https://play.rust-lang.org/".
"/tools".
"/governance".
"/community".
"https://blog.rust-lang.org/".
"/learn/get-started".
etc.I'm always looking for opportunities to dip my toes into Prolog; in hindsight it's clearly a good fit for tree-structured data structures.
(There was no Java, C++, etc. either. It was SML, Pascal, 68000, and Oracle Pascal-Embedded-SQL.)
...many languages (similar to regex / state-machine) can benefit greatly from offloading a portion to something prolog-ish, but it's unfortunate that prolog knowledge isn't as widely distributed.
>>> soup = BeautifulSoup(requests.get("https://www.rust-lang.org").text)
>>> [x["href"] for x in soup.find_all("a")]
['/', '/tools/install', '/learn', 'https://play.rust-lang.org/', '/tools', '/governance', '/community', 'https://blog.rust-lang.org/',...As I see it, a key attraction of Prolog is its simplicity: With a single language construct (Horn clauses), you are able to express all known computations, and the example queries I posted show that only a single language element, namely again Horn clauses to express a query, is needed to run the code. The Prolog query, and also every Prolog clause, is itself a Prolog term and can be inspected with built-in mechanisms.
As a consequence, an immediate benefit of using Prolog for such use cases is that you can easily reason about user-specified queries in your applications, and for example easily allow only a safe subset of code to be run by users, or execute a user-specified query with different execution strategies etc. In comparison, Python code is much harder to analyze and restrict to a particular subset due to the language's comparatively high syntactic complexity.
Honestly, if we didn't talk about the benefits of a language irrespective of how easy it is to hire for it, we'd never have introduced anything beyond FORTRAN, if we even made it that far. Bringing "X is easier to hire for" into a conversation about the language is, at best, a non-sequitur.
If we had just stuck with FORTRAN forever, how many problems would have been completely avoided!? There’d be better, and more, IDEs, since even if the language is hard to parse, it’s still just one parser that needs all the effort. So many unfortunate problems in education caused by language and ecosystem churn would have been avoided (the infamous “by the time you graduate, it’s always outdated” problem).
The only problem is that FORTRAN is too new. Should’ve stuck with the Hollerith tabulator.
?- use_module(library(sgml)).
true.
?- use_module(library(http/http_open)).
true.
?- use_module(library(xpath)).
true.
The second query also uses portray_clause/1 from library(format), which you can load with: ?- use_module(library(format)).
true.
After all these libraries are loaded, you can post the sample queries from above, and it should work.There are also other ways to load these libraries: A very common way to load a library is to use the use_module/1 directive in Prolog source files. In that case, you would put for example the following 4 directives in a Prolog source file, say sample.pl:
:- use_module(library(sgml)).
:- use_module(library(http/http_open)).
:- use_module(library(xpath)).
:- use_module(library(format)).
And then run sample.pl with: $ scryer-prolog sample.pl
You can then again post the goals from above on the toplevel, and it will work too.Another way is to put these directives in your ~/.scryerrc configuration file, which is automatically consulted when Scryer Prolog starts. I recommend to do this for libraries you frequently need. Common candidates for this are for example library(dcgs), library(lists) and library(reif).
Personally, I start Scryer Prolog from within Emacs, and I have set up Emacs so that I can consult a buffer with Prolog code, and also post queries and interact with the Prolog toplevel from within Emacs.
[0]https://en.wikipedia.org/wiki/XPath [1]https://github.com/benibela/xidel
When scraping and parsing (or writing integration test DSL), I always start out with CSS selectors. But always hit cases where they lack or require hoop-jumping and then fall back on Xpath. I then have a codebase with both CSS-Sel and Xpath, which is arguably worse then having only one method.
I suspect here, one uses this tool untill CSS selector limitations are getting in the way, after which one switches to another tool(chain)
Looks like that psuedo-class has not been implemented in the kuchiki library that htmlq uses though.
# Find all elements li and select the parent element for each
//li/..
# Find all element nodes with a child element named li
//*[li]
# Non-abbreviated queries
/descendant::li/parent::*
/descendant::*[child::li]
# CSS using :has
:has(> li)E.g. when you have a list of numbers on the website, XPath can calculate the sum or the maximum of the numbers
Or you have a list of names "Last name, First name", then you can remove the last name and sort the first names alphabetically. Or count how often each name occurs and return the most popular name.
Then it goes back to selection, e.g. select all numbers that are smaller than the average. Or calculate the most popular name, then select all elements containing that name
Element with src or href attr: //[@src or @href] or multiple conditions: //article[@state = "approved" and not(comments/comment)]
Element with more than two children: //ul[count(li) > 2] Element with matching descendents: //article[//video]
Element text containing substring: //p[contains(text(), "Foo")] Attribute containing substring: //a[ends-with(@href, ".jpg")]
Numerical attribute selection: //product[@price > round(2.5 @discount)] //product[sum(//[starts-with(name(), 'price-')]/@price) > 0]
Attribute values: //a/@href Text values with spaces normalised: //a/normalize-space(text())
Match all attributes or elements or text nodes: //user/@ or //user/node() or //user/text() or //user/comment()
Basically from any node in a document you can select its ancestors, children, descendants, siblings, attributes etc, and filtering has the same power as selecting does - in CSS there's :not() that can apply to selection or filtering, with :has() finally on the way and no :or(). CSS selectors match against HTML elements and they're great for that almost all of the time, but while you can filter by attribute value including substring and even by regular expression, for text there's :empty.
But for a query syntax you need to be able to select attributes and text content as well as elements. Either extend XPath to support #id and .class syntax
//#user-xyz//note/text() //code.language-js/@name
or extend CSS to at allow selecting attrs and text
#user-xyz note :text code.language-js @name
The former is more powerful, the latter a quick hack (if they only appear at the end of the selector anyway) with instant payoff.
(: XQuery comments are marked by mirrored smilie faces, like this. :)
An empty query is not valid. There needs to be something besides the comment
<users>
{
for $user in //users
let $comments = //comment[@uid = $user/@id]
where count($comments) > 0
order by $user/lastName, $user/firstName
return <user id="{ $user/@id }">
<name>{ concat($user.firstName, " ", $user.lastName) }</name>
<comments count="count($comments)">
{
for $c in $comments return <comment id="{ $c/@id }" />
}
</comments>
</user>
}
</users>
It's the bastard child of SQL and XPath 2 lol.It's also in nixpkgs, though for some reason the nixpkgs derivation is marked as linux-only (i.e. not Darwin). (Edit: probably because the fpc dependency is also Linux-only, with a linux-specific patch and a comment suggesting that supporting other platforms would require adding per-platform patches)
Comparing the two repos, it seems pup is dead, but cascadia may not be.
These tools, including htmlq, seem to sell themselves as "jq for html", which is far from the truth. Jq is closer to the awk where you can do just about everything with json. Cascadia, htmlq, and pup seem closer to grep for html. They can essentially only select data from a html source.
[0] https://github.com/EricChiang/pup [1] https://github.com/suntong/cascadia
But I don't think html has any need for a sed/awk tool, or at least not as much. Json output could very well be piped forward to the next CLI tool after you've changed it slightly with jq. I don't see this scenario as likely with html.
Exactly, and that is what I mean. If you want to compare, compare it with grep, not jq.
Someone else posted xidel[0] in this thread, which I've not used, but it seems to be the "jq but for html".
This is the kind of obvious tool that once it exists, you can’t really grok the fact it did not earlier, and that it took until now to exist.
A good opportunity to introduce `gron` to those unfamiliar!
▶ gron "https://api.github.com/repos/tomnomnom/gron/commits?per_page=1" | fgrep "commit.author"
json[0].commit.author = {};
json[0].commit.author.date = "2016-07-02T10:51:21Z";
json[0].commit.author.email = "mail@tomnomnom.com";
json[0].commit.author.name = "Tom Hudson";
https://github.com/tomnomnom/gronThank you - appreciated.
I haven't done much work with json but have had reasons recently to do so - and I immediately saw how difficult it was to pipeline to grep ...
But what I still don't understand is that some json outputs I see have multiple values with the exact same name (!) and that still seems "un-grep-able" to me ...
What am I missing ?
It seems both gron and jq only use the value that has been defined last:
~ echo '{"a":1,"a":2}' | gron
json = {};
json.a = 2;
~ echo '{"a":1,"a":2}' | jq
{
"a": 2
} > But what I still don't understand is that some json
> outputs I see have multiple values with the exact same name
This is neither explicitly allowed nor explicitly forbidden by the JSON spec. It is implementation dependent upon how to handle - does one value override the other? Should they be treated as an array?In practice, this situation is usually carefully avoided by services that produce JSON. If you are interfacing with a service that does produce duplicate values, I'd be interested in seeing it for curiosity's sake. If you are writing a service and this is the output, then I implore you to reconsider!
It doesn't handle malformed HTML that well but can be coaxed into working about 90% of the time, with the help of the other included package hxclean or something like html-tidy.
"jq is like sed for JSON data"
sed: "While in some ways similar to an editor which permits scripted edits (_such as ed_), sed works by making only one pass over the input(s)"
ed: "ed is a line-oriented text editor".
Software definition through a reference to another software is somewhat confusing. Potential users come from different backgrounds (I had no idea what is jq), and it is not clear what are the defining features of each project. Is jq line oriented? Is htmlq operating in a single pass?
As a job, computers were largely automated out of existence by solid-state transistor based automated computers and integrated circuit transistor automated computers in the 60s, 70s and 80s, which replaced the enormously expensive and often largely experimental electro-mechanical automated computers while radically reducing cost and improving performance both by several orders of magnitude.
You know Jimmy the famous mechanic? I'm Timmy, _his brother_ but an electrician.
IMO, at least `jq` has proven itself as the indispensable tool for json-data manipulation.
2nd sentence - Explaining the tool to folks in the general web domain what it can do for them.
3rd sentence - Explaining where to learn how to use the tool if you've stumbled across it but web is not your area of expertise.
All that info fits in nearly 25 words then it lists the options for the tool and jumps straight into multiple examples (with outputs!). If the only explanation had been "htmlq: like jq, but for HTML" I'd agree but having the comparison to explain what it does isn't a bad thing it's _only_ having the comparison that would be bad.
Personally I think this is a model example of a opening for a Github readme.
If you're going to write a minimal introduction, at least make sure it's not confusing.
I get the feeling the author felt compelled to write an introduction and did so with as little effort as possible.
0. https://www.google.com/search?hl=en&q=%22bits%20content%22
Honestly you could drop the "bits" which is a bit redundant and use the phrase "Uses CSS selectors to extract content from HTML files."
But if you haven't used Jq that I can see how that title is less than helpful.
So for general purposes, it's a terrible marketing pitch. And yet I think it's a very, very valuable demonstration of knowing some of their 'customers'.
I'll be trying it out next time I'm on a PC.
I can, and it's not illuminating at all.
I would expect that htmlq run the query a single time for a single html; just like jquery $('#something') or document.querySelector('#something')
Possibly, depending on background as you note, but not all promotion is intended at the same audience. When submitting to HN, "like jq, but for X" is short and conveys what it is to most the people that would care, I think. jq has been submitted and talked about here many times with lively discussion over the years.[1] At this point I think most those that are interested in what that is and what this is will understand fairly quickly from the title. Those that don't might be missed, or they might look it up like you, or they might see it through some other submission some other time with a different title which isn't based on a chain of references.
1: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
Invoke-WebRequest
eg. # what is the latest release of apache-tomcat?
$LINKS=$(Invoke-WebRequest -Uri 'https://tomcat.apache.org/download-80.cgi' | Select-Object -ExpandProperty Links)
$LATEST=$($Links | Where-Object -Property href -Match '#8.5.[0-9]+').href.substring(1)
$FETCH=$($Links | Where-Object -Property href -match "apache-tomcat-${LATEST}.zip$").hrefhxextract and hxselect perform similar extract functions.
hxclean and hxnormalize (combined) will pretty-print HTML.
Then I found out about jq because awscli was using it in example docs.
I guess `htmlq` makes sense if it has the exact same syntax as `jq`, and the user is already familiar with the latter?
XSLT can be an amazing tool when used properly and I've wondered about a JS equivalent over the years and started writing one on a couple of occasions. But JSON is just a data structure and not structured markup, and there's no sweet spot for a transformation tool like XSLT - you're more likely to be doing a "find items in JSON, filter() them, then map()/reduce() to output format" task that takes a minute or two in Node and then never gets used again, or doing a complete map from one domain to another where you'd need to do it in JS because of the complexity of processing and ability to handle errors, use third-party tools and even write tests.
An XQuery-esque language allowing selecting bits of JSON file(s) with filtering, grouping and ordering built-in, combined with a way of projecting results that's no worse than JS allows for i.e. not having to put quotes around everything and the like :)
#!/usr/bin/env ruby
require 'nokogiri'; p Nokogiri::HTML(STDIN.read).css(ARGV[0]).text
Just save it to a file in your /usr/local/bin/hq and do chmod +x !$Then you can do:
curl -s "https://news.ycombinator.com/news"|hq "tr:first-child .storylink"
It uses Nokogiri[0], which is much more battle tested and works with CSS and XPath selectors.[0] https://nokogiri.org/tutorials/parsing_an_html_xml_document....
curl -s "https://news.ycombinator.com/news"|hq "tr .storylink"
Deploy a website on imgur.comMy £4 a month server can handle 4.2M requests a dayFirst Edition...An xmlq that was really like jq would be fun, about 20 years ago.
There is also `xmlstarlet` for parsing XML in a similar fashion.
Is this the yq? https://kislyuk.github.io/yq/ It does contain an 'xq', as a literal wrapper for jq, piping output into it after transcoding XML to JSON using xmltodict https://github.com/martinblech/xmltodict (which explodes xml into separate JSON data structures).
This is a bash one-liner! But TBF it really is a 'jq for xml'. I think it would be horrible for some things, but you could also do a lot of useful things painlessly.
If you have any other suggestions for parsing XML for exploratory purposes I'm very happy to hear them.
Just installed xq. It's nice just seeing the pretty-printed json output, so thanks for the pointer. Probably better than xmlstarlet for my usage, which just queries and outputs text, not xml. hmmm, that's probably true for most commandline uses...
"jq is a lightweight and flexible command-line JSON processor"
Jq is very loosely inspired by that, I guess. Might come full circle here and use some XSL transformations ...
I made a tool that extracted parts of web pages 10-15 years ago, and it worked well. There are of course cases where the html is so unstructured that the results were unpredictable, but it worked well in general.
$ xmllint --html --xpath …
that doesn't choke on inline svg.
[0] https://github.com/bAndie91/tools/blob/master/usr/bin/parsel
Any pointers to something that exists? Interestingly I've also found very little for dom extraction in the OS ML space.
But it’s not using a full browser to back it, which I suspect is what’s really being asked.
https://stackoverflow.com/questions/1732348/regex-match-open...
> While arbitrary HTML with only a regex is impossible, it's sometimes appropriate to use them for parsing a limited, known set of HTML.
I hope we standardize on some jq query language, like we have with a base set of SQL syntax