Full Text RSS Feed: Get the whole feed and nothing but the feed
fulltextrssfeed.com
fulltextrssfeed.com
Word of warning, if it takes off, you basically start turning into someone who is both caching and harvesting the web every 15 minutes. There is an incredibly long tail on RSS feeds and it starts killing you to keep them all up to date. Storing and serving it is no big deal, but harvesting actually turns into real money when you figure out total bandwidth used. (we harvest about ~30,000 feeds every 15 minutes)
Just checked and we have ~25k feeds in the system, though not all are deep harvesting as we call it.
Note we do a few things over just extracting the full content as well, we also try to grab out images and create a pleasing thumbnail using face detection etc.. So that probably slows things down a good deal as well.
What's the difference between a Cease & Desist letter from a big company, and free advertising for a startup?
I think this is a great idea and very similar to a lot of stuff I have worked on recently. It's cool to see so much interest in these text-related services.
EDIT: The date and title are in the RSS feed already! No further analysis needed.
btw I know that at Techmeme, Gabe spent years perfecting his story parsing for the 50k+ sites he tracks. Even something that would seem simple such as parsing the date of a story from a webpage has a ridiculous number of permutations that you have to grep for.
I have a lot of experience with fetching and parsing feeds and pages, so I'm not trivializing the problem, just observing issues I'm seeing with this solution.
I've tough of this idea since 2 years, but I am so ineffective at building my own ideas that it doesn't surprise me that someone else built it, as the idea was really floating more and more since instapaper mobilizer.
Considering the legal aspect I had more ideas about that. It is to hide behind the DMCA takedown, and provide an email address to take-down a feed. But do not map the www.example.com/feed.xml to http://fulltextrssfeed.com/www.example.com/feed.xml , but use an alias, so the take-down just remove the alias not the whole * .example.com*.
http://fulltextrssfeed.com/www.aaronsw.com/2002/feeds/pgessa...
And yeah, it was built in a weekend :) But now you don't have to.
"We understand you'd like to delete your account. If you delete your account all of your information including your comments, messages, posts, and friends and followers associations will be removed from our system. Please consider the following options before clicking delete."
Yikes! =X
(also, that needs to be fixed asap, lest anyone get the wrong idea)
Yet another service that requires me to hand over information on what I read.
Why couldn't this be made as a privacy-respecting application I can run from my own machine?
Works well for keeping up with HN too :)
Setting up a quick script to rip all the content from another site is trivial. There's also wget -m
Add /opt/local/bin to your PATH (bashrc or zshrc or whatever you use).
sudo port install libxml2 and sudo port install libxslt
Then sudo gem install nokogiri --no-rdoc --no-ri should run with no issues. That's all I had to do for the system ruby (1.8.7 on OSX 10.6) and 1.9.2 via rvm.
Im currently using "Readable Feeds" Nirmal J. Patel (http://www.nirmalpatel.com/hacks/hnrss.html) and Andrew Trusty (http://andrewtrusty.com/2009/06/29/readable-feeds/)
I like it -- but its really inconsistent!
Cheers,