Show HN: Cocktail search written in Python using scrapy and sphinx
cocktails.p24dev.de
cocktails.p24dev.de
1. Some stemming/normalizing such that "créme" and "creme" (as in, say, creme de menthe) are not two distinct result sets. (To your credit, whisky/whiskey appear normalized)
2. A list of sources available apart from mouse-scrubbing the results. Any reason Esquire (considered borderline definitive by certain cocktail snobs) isn't in there?
Wishlist:
-A way to make a search for "brandy" return anything containing "cognac" or "calvados," but not vice versa. Ditto for whisky vs scotch/rye/bourbon.
-Ability to exclude certain sources from all searches
At the moment I only crawl Wikipedia, liquor.com and seriouseats.com for cocktail recipes. But thanks for mentioning Esquire. I didn't knew that website yet. Maybe I will crawl them as well in the future.
One direction mapping of words is unfortunately not possible with sphinx. However I could expand "brandy" to "brandy OR cognac OR calvados", just before I run the query. I will consider that approach.
Things I'd like to see: * I prefer milliliters over oz (I didn't even know what it was before I googled it, since I live in Europe), please make this an option * Some form of autocomplete, preferably with a dropdown list (but make sure it doesn't obstruct the next textbox). There's a jquery plugin that does dropdown autocomplete well. * Pressing enter should make the next textbox active (seems more intuitive than tab in this case) * It would be nice if the ingredients in the ingredient list below a cocktail were clickable. Clicking them should add them in the list of selected ingredients and update the cocktails visible.
While I appreciate the use of history, I promise you I don't need to hold my position on the page for every single letter I type in the search box: http://screencast.com/t/1qFwxlz5
I ended up with a history entry for "search results for 'c'", "search results for 'co'", "search results for 'cog'", etc.
Maybe trim this down to only firing after 10s on inactivity in a text input or use onblur? I'm not too sure what the answer is but it (page history) definitely loses any usefulness when architected in this manner.
However when updating the history only every 10s, you might have to wait until you can copy the link to your search from the address bar. And if you aren't typing fast enough there still can be incomplete searches in your history.
Using onblur would be better. But what would you think about, updating the history as soon, as the cursor is moved? I think I like that idea pretty much.
However I have just implemented another approach to deal with that issue. Now, the history is updated either when the current field lost focus or when you move the mouse, after you entered something. Please try it out and let me know what you think.
I would also have a couple suggestions for improvement. (1) include the recipe for how to make the cocktail, as opposed to simply the ingredients. This can make a big difference in the result. (2) include an AND function for the ingredients search (eg. a search for pisco AND lime, I didn't see that this exists currently).
Searching the introductions in additions to the list of ingredients, will rank recipes higher, when the ingredients you were searching for, occur more often in the instructions. But a recipe isn't necessarily more relevant, because of its instructions are more verbose.
I'm not sure if an AND function would add much benefit. Yes, at the moment only one of the given ingredients must be part of the cocktail, in order that it appears in the search result. But the search results are also sorted by relevance, so that cocktails that contain more of the searched ingredients are ranked higher, than cocktails that contain fewer of the searched ingredients.
I'm pretty sure that there is no way to PM other users, here at Hacker News and you don't have an email address in your profile. However if you want have a further discussion on that, feel free to write me an email. You'll find my email address in my Hacker News and Github profile.
There is as you have already discovered, a crawler implemented with scrapy. However I don't use the scrapy server and pipelines. Instead I have a script that lets scrapy generate JSON files with the crawled recipes and builds the sphinx index from the crawled data. There is no RDBMS. Basically sphinx is my database. :)
Well and than there is the website. Its server-side is implemented with werkzeug and its UI with jQuery.
There are alternatives to Sphinx? ;)
* Sphinx is ridiculous fast, as you can see when searching. But even building the index takes only 340ms (from which 220ms are spend by the python script that generates the XML) for 1699 recipes on my 3 years old notebook.
* Sphinx don't require a RDBMS to index documents from. It can index documents from any source. You just need to write a simple script that brings the documents in the XML format expected by sphinx.
* Sphinx is not only a full text search engine. It is also a multi-value store. You can add extra information like the title and url to indexed documents. And so you don't need an additional database.
* I need the ability to limit a fulltext search to sentence boundaries. I don't know if there are other fulltext search engines that can do that.
@ingredients (rye SENTENCE whiskey) | vermouth
The SENTENCE operator, basically works like the & operator, just that both operands must occur in the same sentence. That is very helpful, since the index field "ingredients" contains a list of all ingredients of the recipe separated by an "!".
However Sphinx can not only do full text search. You can also filter and sort by attributes and complex expressions, that involve any attribute, the relevance from the fulll-text search, arithmetics and some built-in functions. And Sphinx does that much faster than any RDBMS does. I never managed to generate a sphinx query that took a measurable amount of time, on my 3 years old notebook. ;)
The sphinx index doesn't contain any HTML. However I return HTML, from the WSGI app, that serves the AJAX calls. So I guess that is what you are talking about. It just seemed to me, it would be simpler to generate the HTML for the search results on the server-side with Python than in Javascript.
By the way if you want to use Sphinx for your own project, and you have the data to index already in a MySQL or PostgreSQL database, there is no need to write a script to generate XML. Sphinx can index data from MySQL and PostgreSQL databases directly. However I was just saying that, thanks to xmlpipe support, you don't need an RDBMS just to use Sphinx. And that thanks to index attributes, Sphinx can completely replace an RDBMS in a lot of cases.
Back button behavior is pretty annoying :)