Keep the good work. I love it :)
My only suggestion is to offer a ranking and order the results by # of quotes
Keep the good work. I love it :)
My only suggestion is to offer a ranking and order the results by # of quotes
On parsing: it's actually a somewhat fastidious process that involve digesting a couple of GB of data, but here is the bottom line.
I look for amazon.com links in the content in general - I will broaden that to other publishers and full-text extraction too later on.
The content itself comes from the StackOverflow dump (for SO) and a mixture of a crawler allowed by PG + the previous database dump that was available at some point.
I extract all the books, quotes, users data from both, conform these into a common schema, and index the whole result.
Hope I answered your question properly - feel free to ask again if you'd wish.
On ranking: I know what you mean! I need to find some way to balance number of quotes with textual relevance, which requires me to dive a bit more into solr. I currently use textual relevance first because it gives more useful results so far.
We'll see how it goes.
For ranking, instead of textual relevance (which will be hard to achieve :/) and # of quotes (it can be easily hacked/spammed or new announced books will have a huge weight on the ranking), I suggest you to check timelines: if a book is quoted once/twice/... a month regularly, I'm pretty sure it worths reading it
I'll love to read about the architecture behind the site... yep, technically curious :)
On the architecture: I'll create a side-blog that will outline all I learned while working on this. It's been a crazy ride actually (especially because I started using chef and vagrant full speed).
I'll post it back here in all cases.
Questions: Apart from the "Quoted by" section, is any of the content from SO or HN displayed on the site? E.g. are any comments incorporated in the book descriptions, or are those written by you/wife?
Currently, no content from HN/SO is displayed apart from the Quoted by area.
We're not editing anything manually; what I will do is display the actual conversations in the "Quoted by" area, either when you click on a conversation.
I may ove the quotes above to make them stand out more.
Did I properly answer your question ?
Was justing wondering if the content might be touched by any copyright issues etc.
I really love what you got here. I'd be happy to help you try out IndexTank and make it better. It would really take the Solr configuration burden off of you.
well I considered using IndexTank earlier on, especially because I didn't know yet how to deploy Solr. The relevance is mostly done already, I just needed to learn to use formulas :)
One thing that put me off is your cap in queries per day. The smallest paid plan (50k items in index) is capped at 1,000 queries per day.
Isn't that an issue for most sites ? Do people usually cache your results ?
I can't really tell yet, but my guess is that at least around the number of indexed documents could give a more usable subscription.
It would also be nice for people to know what happens if you go beyond the cap: do you offer some tolerance ?