Data Mining: Finding Similar Items and Users
bionicspirit.com
bionicspirit.com
This is a vast area of research. Diving head on might result in serious injury. For example, R statistical package simba (http://cran.r-project.org/web/packages/simba/index.html) lists 56 different similarity/dissimilarity measures just for binary data.
Many data mining approaches first use a feature selection or feature extraction approach. That is, an approach which finds the relevant feature subsets, or discovers the underlying features of the data set.
Inverse Image search and the solution to the Netflix prize both used feature extraction approaches.
My startup provides a very fast similarity engine (in a DB of 100 million objects I can find similar objects in under 20 millisec. with one CPU) in case you worry about scalability. URL: http://simmachines.com
Unfortunately not yet, they will be released but I am not sure when.
I don't know the python libraries off the top of my head, but the Python implementation of R (RPy) is decent and R is heavily used in most circles to have great literature written for it.
Two good first steps to look into, depending on your needs, are Bayesian classifiers and SVD (reduction of high dimensionality, the application to text processing was patented as Latent Semantic Indexing/Analysis, LSI or LSA, by IBM, I don't knnow if that's lapsed).
Basically, you need a reasonable feature to match similarity on. N-words are pretty easy to construct, a 2-gram would be every pair of words used in a document.
Tf-idf is a good metric with that kind of feature, because it handles well the bias of frequent words like "the"
I am somewhat like many people who are unfamiliar with data mining for this type of matching.
For example, I want to provide a "similar items" for Vacation Rentals, where the "dimensions" or attributes, could be "location", "bedrooms", "price", etc. It's hard to quantitfy anything to show something which might be more relevant to someone else based on the previous properties that they have currently been viewing.
Instead I have just taken to approach of creating a bounding box based on the Geo coordinates, and then offer up similar properties within their search price range. But I would really love to eventually implement something like your original article. (Suggestions welcome).
Their suggestions are like: customers that viewed this item also viewed; customers that viewed this item ended up buying; customers that bought this product, also bought these other products.
That last metric in particular is interesting, because it tells you for a product what are the complementary products that customers may be interested in. So you don't actually have to measure somehow the physical properties of the objects getting sold to discover relationships.
In your case I don't have knowledge about the problem domain to give advice, but "customers that viewed this deal also viewed ..." is always a great addition. Also add ratings and follow-up on people with emails to rate on their vacation, after coming back from the trip. I don't know how well it will work - there's no general solution, you try something and if it doesn't work, try something else.
I could very easily create something that takes every visit, MapReduce it, and then track the entropy between potential matches to provide the "best" match based on user visits of that property also.
To take the example, it would be really great to also know that people from Germany aren't interested in the slightest in our Italian properties based on trends of their national behaviour.