A journey in e-commerce search
danpalmer.me
danpalmer.me
When we get to the last step, of implementing bitmaps for filters, and then trying to optimize to deal with the RAM implications... why not just use Solr or ElasticSearch instead of reinventing the wheel? That's basically what lucene (which Solr and ES are built on) is doing, something like that implementation you came up with, but with lots of optimizations and tunings on top of the naive implementation.
Not quite. Rather than the free text search being over products, it's over combinations of filters. These filters can then be applied to produce "exact" results, rather than "fuzzy" results based on textual matches.
> why not just use Solr or ElasticSearch instead of reinventing the wheel
A few reasons. Firstly, the faceted search in these took a fair bit of effort to expose to users in a way that wouldn't result in bad results (like "blue shirt" returning blue non-shirts, and non-blue shirts)
Secondly, these are new infrastructure components that take a lot of work to run. ElasticSearch is notorious for being a pain to operate. We were a small engineering team, only 3-5 backend engineers in the company at the time (depending on how you count), running the whole stack, infra, everything.
But this means there are then non-blue shirts.
One of the issues for us is that we wanted to rank by item suitability for the user, not by relevance. When using FTS, this didn't work because you'd get a non-blue shirt that would be better for the user and that showed up first. But when using the filter-based search, that couldn't happen because the recommendation algorithm was only getting blue shirts to rank, so all of them were relevant, and the same level of relevance.
I think sticking with Postgres was a mistake, though. Denormalizing into something like ElasticSearch would probably have gotten you further (and it's so fast).
Searching the filters is pretty clever, though.
And then fine tune what to return by putting "and" or "or" in between keywords based on that search engine's behavior.
If I understood correctly, a user query is first applied to the `SearchItem` table (using the `search_vector` field?). Then the returned `filters` are used to generate a normal (not full-text search) SQL query that would retrieve items from the product table?
But putting that aside for a moment, most of the effort here actually has nothing to do with which engine you are using to power your search. For this kind of challenge, it doesn't matter if you use Elasticsearch, Algolia, or Postgres Full Text Search! Instead, the problem space has a lot more to do with your data and how we can teach a computer to interpret it the way a human would.
In the e-commerce/search space, one of the biggest challenges is what I call "query understanding," which is breaking down the user's query and trying to understand what they meant. This isn't trying to be "smart" and show you things you aren't looking for, but rather understand what the underlying intent is and filter on _that_.
In OP's scenario, you aren't searching for things with the word "blue" and "shirt" in it. You're actually searching for one thing: a shirt, with a (strong) preference for blue!
Most search engines, by default, will do something to the effect of "SELECT * FROM products WHERE anyfield CONTAINS 'blue' AND anyfield CONTAINS 'shirt'" and just return everything. But as OP discovered, that might return a product where "blue pants made by the people who brought you your favorite pink shirt," which definitely isn't what you want.
Instead, you look at the query: what are we looking for here? We're looking for a shirt. So the entity is "shirt" and the modifier is "blue." Typically "entities" represent "categories" or "product types", but often are more granular than what we would put in those filters. Thus, when we construct the information architecture for an e-commerce platform, we need to include specific types of data that represent higher-order abstractions of the product that we're looking at. Cheezits aren't just Cheezits -- they're "crackers", and more specifically, "cheese crackers." Milk isn't just milk, it's a beverage and it's dairy.
We now have this concept of an entity ("shirt"), and a modifier of "blue", which through various processes we have determined is a color, a theme, or perhaps a pattern. Then we can return a list of all of the shirts, prioritizing the blue ones.
Complicating things is this: is navy considered "blue"? What about teal? Turquoise? Cyan? It's important to be precise about the colors for categorization, thumbnail, filtering/faceting, and UX purposes, but when people type in search queries, they often are not as precise as their brains are.
When sorting the results, should we prioritize proper "blue" over navy or teal? If someone types in "shirt" and then clicks the "blue" filter, what are we showing and hiding based on that?
We obviously need to be a little bit smarter here....
...and now we're down this giant rabbit hole of how to optimize search. Clearly, throwing things into a Postgres Full Text Search isn't going to cut it. We're going to need to leverage more annotated data (garbage in, garbage out -- make sure your product data is properly annotated!), better query understanding, better sorting, and faceting/filtering!
And, honestly, most of this isn't even about whether we're using Elasticsearch or Postgres or Algolia -- a lot of this is taking your meatspace brains and thinking about what words really mean, how people type them in, and how to interpret those words into something that returns something meaningful.
This ain't Google, but it also ain't a phone book. There are tons and tons of weird curve balls that will make you think you'll never figure it out, even to the very end of your days.
Here are a few random examples of how this gets confusing pretty fast:
- "chocolate milk" and "milk chocolate" consist of the same words but are two totally different things! (thank you, n-grams)
- all "roses" are "flowers", but not all flowers are roses! Conversely, "garbage bags" and "trash bags" are the same thing! synonyms can make things weird
- "cream cheese" is not cream-flavored cheese; it's something totally different. If you show me "sharp cheddar cheese" next to "cream cheese," I might not be happy with those results.
- "strawberry yogurt" is entity "yogurt" (not "strawberry" -- even though strawberries are nouns/entities, this time it's "flavor"), and "strawberry yogurt made with organic milk" is still yogurt, not milk!
- people make typos -- how do you detect them? if you get zero results, try a fuzzy search? That might work, but "coke" and "cake" are one transformation away, yet two totally different things
- you may not sell "blue shirts with a whale in an ocean," but you do sell "blue shirts" and "shirts with whales," which ones do you show and in what order?
One could spend years and years trying to perfect query understanding for any particular business. And perhaps even worse, there are no generalizable strategies to solve for everything! My examples above, if it weren't obvious, are food-themed because most of my search experience has revolved around data sets involving food. Searching for clothing presents an entirely different set of challenges! For example, a box of Cheerios is a box of Cheerios, but a T-shirt can come in eleven different sizes and sixteen different colors, and each of them are valid search results. Messy!
This is a pretty fun topic, IMO. Love to see posts about it on HN and love to see conversations about it!
Thanks for taking the time to write this, comments like these are why I'm still a daily Hacker News reader after 10 years :-)
I'm interested how this compares to the second phase of our search implementation. I realise that searching filters was a somewhat limited implementation and not a general solution to search, but it seems like it was pretty much an implementation of what you've described. Certainly, my understanding of our problem was essentially what you've described.
Where I might go next with this would be to investigate the search results and evaluate the quality of those results, and then improve that over time by increasing your query understanding of what the users are asking for.
One question to ask yourself is "What is the purpose of search?" My answer is "to find what I'm looking for as fast as possible."
Remember, time is money, and if it takes customers longer to find what they're looking for on your site versus another e-commerce site (or worse yet, an actual physical store), then they'll leave and go to the other site. Speed and low effort is key!
(Note: If you are happy with conversion rates, then you can stop here -- don't go down rabbit holes you don't need to go down. There are other cool projects to work on, too!)
So, given this idea of making this as efficient for customers as possible, what's the least amount of effort (and therefore the fastest) way for a customer to purchase something?
The answer, of course, is to find it on the first try! (This reminds me of that scene in Happy Gilmore where Happy hits a hole in one, and he says to his coach, "I should just do that every time.")
Imagine I search for `blue shirt`, and the shirt I wind up buying is the 12th item on the list. From a quantitative standpoint, the conversion rate is 100%. I searched for something, and then I bought it! But qualitatively, the user experience is sub-par: you showed me eleven things I wasn't interested in before the 12th thing that I was interested in! This is cognitive load, and every time I have to look at the _next_ item in the list (or worse, go to the second -- or god forbid, the third! -- page), that's a chance that I'll give up and leave, or conduct another search and be equally frustrated.
The goal here, then, is to try and get the right item in the first slot 100% of the time. This is a bit of a holy grail, but it's a good target to aim for.
If your users' queries are simple and straightforward and your catalog is sufficiently small and focused, simply doing ANDs against filters might be good enough, but if you have a huge variety, then you may wind up having the "right" blue shirt be pretty far down the list.
How do you fix this? Well, first you need data. Track conversion rates, track which index position the conversion, and look at keywords that people are looking for. Also look at search queries that result in zero hits! Consider typo correction and synonyms to ensure maximum coverage of those queries.
If you start identifying keywords that folks are searching for but finding zero or low conversion, then you have some interesting information! Either start carrying products that match those terms, or start tagging relevant products with those terms to increase conversion rates and reduce zero-hit results.
Then you possibly need some sort of popularity metric (sales volume over a period of time is a good one) or some sort of personalization engine (looks like you have something, but I generally don't recommend it for folks just starting out). That way the most popular "blue shirt" will be first -- the odds are fantastic that your customer's preferences are in line with the majority -- and this should boost the conversion index to be closer to that first slot.
Also, as sort of a tangent, IMO it's also better to say "Hey, we don't have that," than to return bad stuff. Again, imagine that I search for `blue shirt` and you don't actually have any blue shirts. In that case, you may drop `blue` and return `shirts`, or maybe you drop `shirts` and return `blue`, and now I see red shirts and blue pants. It's not clear to me as a user that there are no blue shirts, so now I waste time looking through potentially pages of results!
Other things to consider, which I think you already discovered:
1) Autocomplete is a powerful way to search. Naively, you autocomplete the `products` table, but as you discovered, it's almost better to autocomplete search queries (or filters). This probably means you need to roll your own autocomplete system, tracking the most popular search terms/filters over time, and dropping the less popular searches from the list. I have a lot of thoughts about this, but it's not insanely complicated, so I'll stop here.
2) Sorting/Filtering is really powerful, but it's also really easy to screw up. In general, my rule of thumb is to sort by things that are either objectively useful (filters) or by something where I control the definition. For example, sorting by "price" is something that's highly requested, but sorting by price often returns unintuitive results.
Consider Amazon and searching for "dslr camera" and then sort by "Price - Low to High." What you really want to see is cameras -- but what you see at the top of the list (below the ads) isn't cameras, it's camera accessories. Lens caps, straps, batteries. That's not what you meant.
Instead, we would want to sort by "Best Deal", which means the top item might not be the objective lowest price, but it's the lowest priced camera, which is what you're asking for. (And your query understanding model needs to recognize that the entity is "camera" in order for this to work!) If you promise "Price - Low to High" and the first item (an actual camera) is $300 and the second item (battery) is $25, people might feel like either the filter is broken or that you don't know what you're doing. If you use subjective language like "Best Deals" or "Trending," you can play around with the results and build trust more easily.
3) Temporality is introducing the concept of time. Over time, preferences chance. Certain search terms fade from use at certain times of year. People stop searching for "sandals" in the fall and start searching for "boots," and so if you type in "brown shoes" in November, you probably want to ensure that boots are at the top, and not Birkenstocks. This is a pretty fun problem, but you also have to be careful because maybe they did want sandals, and you don't want to mess that up.
I'm just rambling at this point. Very passionate about this topic. Good luck with your system!
Essentially we forced all queries to be either only faceted search, or to fall back to full text across descriptions. For our use case, almost all queries could be satisfied by facets, so it didn't make sense to include any non-faceted search results in the results for those queries.