#!/usr/bin/env python3
import requests
from bs4 import BeautifulSoup
for p in range(1, 17):
r = requests.get(f'https://news.ycombinator.com/favorites?id=app4soft&p={p}')
s = BeautifulSoup(r.text, 'html.parser')
print([{'title': a.text, 'url': a['href']} for a in s.select('a.storylink')])One more question: what is the best way stop it when it will reaches last page?
> for p in range(1, 17):
Actually p=17[0] is empty (as p=16 is maximum as for now).
Maybe, script should scrap pages from `1` to `infinity` UNTIL it detect next message on page[0]:
> app4soft hasn't added any favorite submissions yet.
A better way to solve it is to look at the `len()` of the results, and stop when it gets to 0:
p = 1
while True:
r = requests.get(f'https://news.ycombinator.com/favorites?id=app4soft&p={p}')
s = BeautifulSoup(r.text, 'html.parser')
faves = [{'title': a.text, 'url': a['href']} for a in s.select('a.storylink')]
print(faves)
p += 1
if len(faves) == 0:
break path = 'favorites?id=app4soft'
while path:
r = requests.get('https://news.ycombinator.com/' + path)
s = BeautifulSoup(r.text, 'html.parser')
print([{'title': a.text, 'url': a['href']} for a in s.select('a.storylink')])
more = s.select_one('a.morelink')
path = more['href'] if more else NoneI'm not sure if it still works as it too relied on HTML scraping. Perhaps I should update it to support favorites too.
Edit: Whoa, it's been 4 years already. I believe HN didn't have favorite feature at the time. That's why I used upvoting as my bookmarks system and created a script to export that data.
Edit: I see from the other PR it's called `upvoted`.
Edit 2: I changed it to `upvoted` and now I get a 200 OK, but the code crashed right afterwards on `tree.cssselect()`.
Btw, good work on your JS solution. It's great that it just works without requiring any download or installation. :)