Pup: Parsing HTML at the command line
github.com
github.com
Pup – Like Jq, but for HTML - https://news.ycombinator.com/item?id=24797697 - Oct 2020 (2 comments)
Show HN: Pup – A command-line HTML parser - https://news.ycombinator.com/item?id=8312249 - Sept 2014 (27 comments)
Random bit of history: that Show HN was a very early choice for what is now called the second-chance pool:
Ask HN: Why did three HN stories jump 100 ranking points in 5 mins? - https://news.ycombinator.com/item?id=8313505 - Sept 2014 (6 comments)
In fact, all of the source data for a project of mine, Baytyab[1] (couplet-finder) was scraped using bash + pup.
[1]: https://baytyab.com
I can basically write a .ts file anywhere and run it from the CLI with deno. Definitely prefer it over python at this point
I am interested in hearing more about this as maybe I should switch from Node to Deno for a scraping project I have.
You can probably try it in a few minutes if you already have Deno installed. Make a new file. Here's a quick example script
```ts
/* run with `deno run --allow-net demo.ts` */
import { DOMParser } from 'https://deno.land/x/deno_dom/deno-dom-wasm.ts';
const API_URL = 'https://subsidytracker.goodjobsfirst.org/parent/tesla-inc';
const main = async (args: string[]) => {
const resp = await fetch(API_URL);
const text = await resp.text();
const html = new DOMParser().parseFromString(text, 'text/html');
console.log(
'Tesla subsidies:',
[
...html.querySelectorAll('table:first-of-type > tbody:first-of-type > tr')
].map(tr => {
const [name, value] = [...tr.querySelectorAll('td')];
return {
name: name.textContent,
value: value.textContent
};
})
);
};
main(Deno.args);
```save that file and run:
deno run --allow-net demo.ts $ curl -s https://news.ycombinator.com/ | fq -r -d html 'grep_by(."@class"=="titleline").a."#text"'
Inkbase: Programmable Ink
New details on commercial spyware vendor Variston
How We Built Fly Postgres
...
$ curl -s https://news.ycombinator.com/ | fq -r -d html '{hosts: {host: [grep_by(."@class"=="titleline").a."@href" | fromurl.host]}} | toxml({indent:2})'
<hosts>
<host>www.inkandswitch.com</host>
<host>blog.google</host>
<host>fly.io</host>
...
</hosts>
See https://github.com/wader/fq/blob/master/doc/formats.md#xml and https://github.com/wader/fq/blob/master/doc/formats.md#html for examples and documentations.Last time I looked at using pup or similar I wanted to extract two values for each element. For example, let's say I have the following html:
<div class="image">
<p>Sunset in Hawaii</p>
<img src="../randomstring123.jpg">
</div>
<div class="image">
...
Now I'd like to get both the image description, and the source, for each similar image in the page. Preferably so I can pipe it to curl and do curl -o "$description.jpg" "$url"
I couldn't find an easy way of doing it, so I used Python instead.However we do have document.defaultView.getComputedStyle which probably will solve most of your needs. Not sure if that API is available with this tool though
A driver for an existing browser is the only reasonable option for the foreseeable future, with how complex modern web development has become.
You'll be stuck with GET requests only though, and very simple ones, unless you get creative with the js you pipe in.
Simple example, select all img tags without alt text, and insert given text into the alt tag. Or change the domain for all a.hrefs
Direct installation of brew scripts isn't supported anymore. `go get` installs aren't either.
It needs an update.
go install github.com/ericchiang/pup@latest
curl example.org | xmllint --html --xpath '//some/xpath/selector' -