On parsing: it's actually a somewhat fastidious process that involve digesting a couple of GB of data, but here is the bottom line.
I look for amazon.com links in the content in general - I will broaden that to other publishers and full-text extraction too later on.
The content itself comes from the StackOverflow dump (for SO) and a mixture of a crawler allowed by PG + the previous database dump that was available at some point.
I extract all the books, quotes, users data from both, conform these into a common schema, and index the whole result.
Hope I answered your question properly - feel free to ask again if you'd wish.
On ranking: I know what you mean! I need to find some way to balance number of quotes with textual relevance, which requires me to dive a bit more into solr. I currently use textual relevance first because it gives more useful results so far.