Using YQL to grab HN links
query.yahooapis.com
query.yahooapis.com
select * from html where url="http://news.ycombinator.com/" and
xpath='//tr/td/a[substring(@href,1,4)="http"][@href!="http://ycombinator.com"]'
You can play with it yourself here (needs yahoo login):I think it's pretty crazy that you can now scrape well-marked pages with a SQL-like syntax.
Anyone know how to do this?
--
import urllib2
from BeautifulSoup import BeautifulSoup
ychtml = urllib2.urlopen('http://news.ycombinator.com/).read()
for tdtitle in BeautifulSoup(ychtml).findAll("td", "title"):
print tdtitle.a wget -O- news.ycombinator.com | grep -o http[^\"]*
Personally, I prefer curl because it writes to stdout per default: curl news.ycombinator.com | grep -o http[^\"]*
(After posting this, i noticed that HN cuts * signs at the end of a message. So I have to add this text here, or the last * would not be displayed.) wget -O- news.ycombinator.com | grep -o 'title"><a href="[^"]*' | grep -o http.*You are basically able to use Yahoo's cloud servers and huge internet pipe for free with this service.