Show HN: gi.st – the gist of the web
thegi.st
thegi.st
For example:
bbc.com/news thewashingtonpost.com huffingtonpost.com reuters.com cnn.com theglobeandmail.com theguardian.com en.wikipedia.org
Go to any of those as a start and submit one of the articles you find and everything suggested by @gist was generated by an algorithm.
Our hope is that will help get the ball rolling. We are certainly very subject to network effects, so focusing on a niche and providing tremendous value to that one vertical is a good strategy.
I tried a couple of articles on Washington Post and the algorithm did a decent job. I then read the gist first before reading an article, and while some sentences were a little hard to understand w/o the surrounding context, I felt I still got a decent summary. I can see myself skimming the article through this service instead of relying solely on my eye balls when scanning through.
Couple of suggestions:
- Linking the extracted sentences back to the original page if I want the surrounding context - A tool/browser plug-in which can allow me to select a representative sentence from the story and submit to gist
Seems like a useful stand alone service as it is; having people vote on and submit gists themselves would be cherry on top.
In addition, we'll release the api so that people can build their own tools if they wish.
First tests are pretty damn impressive though. Gist'd (is that even the right verb?) a couple of wiki pages and found the summary quite useful, and surprisingly, grammatically correct. Will stay tuned for sure.
I wonder what the feedback will be from big publications? —maybe they will take the hint and write more concise content...
We'll probably change the ratio to scale in a smarter way rather than a constant fraction.
goto one of the following sites: www.bbc.com/news, washingtonpost.com, huffintonpost.com, reuters.com, theglobeandmail.com
On those sites, find any text article that's medium+ size in length, and submit the article (not the home page of the news site) to gist.
You should see that it's not just reiterating the title.
On many sites it's quite buggy, and what you're seeing is likely the result of gist either a.) attempting to gist something when it shouldn't (like say bbc news home page), or b.) attempting to gist when the scrape/parse failed and there was too few or poorly parsed sentences.
We do extract the html title and stick it in at the top of the page as a reference, but that title, like each gist, is editable by the community, and serves only to act as a jumping off point.