Show HN: BBC Good Food Scraper in Go
github.com
github.com
Many sites don't consider user enumeration a bug/threat, but theoretically given enough sites, one could build a profile around a specific email address.
As an amateur web security researcher this project looks really interesting.
What makes this tool work is most sites (as an UX feature) will tell you if an account/email already exists. Whether that be an API call or a notice saying "Your password is incorrect", you'll be able to get the data you need. It was a learning experience for me to use Go to wrap each site check in its own goroutine to leverage concurrency. Quite nice.
[1] https://github.com/imwally/rfcsearch [2] https://duck.co/ia/view/request_for_comments
I was thinking today about making a more generic "when this thing on that website changes" notify me sort of thing. It looks like you're running on AWS. Would you be willing to share how much the bill for that runs to?
e.g. http://healthyeating.sfgate.com/difference-between-salt-sodi...
edit: oh actually the BBC website has it labeled as "Salt", but the HTML ID is "sodiumContent". Weird. Worth a comment then :P
Between "remember every bit of HTML", and "only remember parsed data" is perhaps "remember every bit of HTML, but notice base-html patterns so it can be massively compressed. "dynamic" content like java-script/AJAX content, rendered dates complicate this...
The BBC Food website, on the other hand, is funded by the licence fee. Recipes from many of their food programmes are published here, but much more too. It's grown to be much more than just a companion to their food programmes. But now, with little support from the public, they've decided to close the site.