it would be nice if it were a PDF I could download and save for later.
is there any way to turn a series of pages in to a PDF? like a recursive wget and then pipe through pandoc?
is there any way to turn a series of pages in to a PDF? like a recursive wget and then pipe through pandoc?
pdfunite page*.pdf output.pdf
I have had decent results using pdftk as well to do pdf surgery so that's another option.In this case, if you do a recursive wget I think it should "just work" because the files are named in a friendly way.
So, putting it all together:
wget -r 'https://dropbox.github.io/dbx-career-framework/overview.html'
cd dropbox.github.io/dbx-career-framework
ls ic*software*.html | sed 's/.html$//' | while read f ; do
pandoc --pdf-engine=wkhtmltopdf $f.html -o $f.pdf
done
pdfunite ic*.pdf output.pdf
[1] ie the ordering of the output of "ls" is the order you want the pages in the output pdfFirst, click the reader view in Firefox, then select all, then paste it into a new Obsidian page. It's really good at keeping a nice formatting and importing pictures etc. You can then export the result to PDF if so desired.
You can hack together some scripts to do the basics yourself, but archiving arbitrary pages is pretty difficult to get right.