Evolution of Search Engines Architecture – Algolia Search Architecture Part 1
highscalability.com
highscalability.com
They are struggling to sell their techno to people who need them deeply, for a lot of reasons. But one of them is that they are a tricky choice. It is not a database technology, so not a developer choice but also their technology is only useful to developers.
As a result they have to try to sell their product when you need a search but no developers are working on it. That's how you end up powering external and internal documentation portals. That's really a waste of resource
How so? (What you follow that statement with doesn't seem to explain it.)
When I took a close look at Algolia for a project it seemed straightforward as a choice, but the cost would've been completely out-of-wack in comparison to what I was spending for the rest of the tech stack. This aligns with what a contributor to Typesense mentions in the comments, which is that what they hear from many potential Algolia customers is that it's "a great product but can get quite expensive at even moderate scale".
Search is really hard even with the best elasticsearch libraries. IMO the biggest blocker with algolia is the price. It's really hard to get company buy-in because leadership don't get it: "Just build it"
It feels tailored to first party use.
This is a really interesting side-effect of what Algolia was probably trying to do: use documentation search to spread brand awareness, but then ironically, because of their (successful IMO) strategy, the product is being perceived as mostly being used primarily for documentation search.
Is anyone else using them? What are your impressions so far?
Much appreciated
Algolia is great to get started but it doesn't make sense at scale. If you have large indexes it's just too expensive.
worth giving a look
Algolia is a great product but can get quite expensive at even moderate scale. If I had a dollar for every time I’ve heard this from Algolia users switching over…
I recently put together this comparison page, comparing a few search engines, including Algolia, you might find interesting: https://typesense.org/typesense-vs-algolia-vs-elasticsearch-...
I was thinking about adding a row about speed to the comparison matrix, but couldn't find a way to express the comparison clearly... Imagine a row that said:
Search Speed | Super-fast | Super-fast | Slow? ...
That felt a little off. So I resorted to just mentioning primary index location as a proxy.
Open to suggestions on how to express this succinctly.
I've had great success at a client where simply upgrading a DB to an instance with enough RAM to fit 80% of the entire data set fixed all performance problems and significantly reduced I/O "pressure" at least for reads (writes were never a problem).
But one thing I would add is ElasticSearch is quite versatile and flexible, so I wouldn’t be surprised if you can contort it to get it to work for a wide variety of use cases. This is a blessing and a curse - blessing because it’s so flexible, curse because the flexibility breeds complexity and brings with it a steep learning curve and operational complexity.
Where I think Algolia / Typesense help is that things work out of the box without the learning curve or operational overhead.
Does it OOM?
Commercially available RAM today goes all the way to 24TB these days, which should be sufficient for a good number of search use cases. Beyond that you’d have to shard data across multiple clusters.
Similarly with Algolia, they use 128GB RAM clusters, and recommend you keep your Algolia index size below 100GB: https://www.algolia.com/doc/guides/sending-and-managing-data...
> Exact Keyword Search ("query")
Any plans on adding it in future?
In the meantime, we introduced a way to turn off typo tolerance and prefix-search on a per-field basis. This has helped some users search for fields containing model numbers for eg.
We don’t have comparative relevancy benchmarks. But we fo have performance benchmarks here: https://typesense.org/docs/overview/benchmarks.html
One of our goals is similar to yours: browsing an online store should be like walking around in a physical store. The navigation system on the site should be as adept as a knowledgeable store employee in helping you find exactly what you're looking for.
At Loop54 many of our customers come from Algolia. It's very popular, and nobody ever gets fired for buying Algolia. In that sense, it's a safe option.
On the other hand, customers come to us from Algolia because Algolia requires a bit of hand-holding and it still doesn't quite seem to get what users are really looking for. When our prospects run randomised controlled trials, our search consistently seems to give users what they want better than Algolia does, with less effort. I can ask about specific numbers if you want.
However, another strength of Algolia that Loop54 is currently behind in is in the surrounding tooling. For better or worse, with Algolia, you'll have more knobs and levers to play with (and you'll need them much more often!)
We do have one or two customers that have a majority of books in their product catalogues, and we know there are some unique challenges that come with that domain.
Loop54 is a very competent, but smaller player. If you think it's interesting, it's worth talking to us. I can't evaluate how good a fit your site would be for us, but that's why we have people who do that for a living!
Edit: I should also say that yes, Loop54 is even more expensive. You shouldn't blindly trust us (or any other provider.) I would strongly suggest running a randomised controlled trial to see whether any expense at all is worth it in your case.
I say this in part because I'm a man of science and believe in experiments to measure things, but also out of self-interest; anyone can throw out impressive marketing, but our search truly shines when put to the test against the alternatives.
The TL;DR is that we're offering customers more flexibility than the tiered model suggests. Consequently there's a large variability in what customers pay. A number on the pricing page turned out to be more misleading than helpful for our target audience.
The long-term solution is creating a more flexible pricing model that can be very transparent. We are working on that. (Though as you know, generic, modular solutions take a while to get just right, unfortunately.)
We didn't think a short-term solution was needed, but based on your comment, that might be worth reconsidering. The best I can think of is a 90 % range of what customers actually end up paying for each tier. What suggestions do you have?
Really happy with the service it provides and the ease of implementation. Note that because the docs can take a few seconds to load, the their crawler times out and misses some content some of the time. With better page performance this should not be an issue.
(We are actively working on some cool ideas to make Lowdefy apps super fast)
What results do you expect more than keywords search ranked by upvote on HN? I find it great honestly, it's fast and don't do magics
I also wish it searched both stories and comments by default. I guess you can set your own defaults but meh, I use many different computers and don't change defaults as a general rule because it's a hassle to keep all the computers in sync.
TLS prevents your run of the mill MITM scenarios. Like ISPs inserting ads (something Comcast actually did), or public wifi doing the same. Or worse, more malicious scripts.
You could argue that all I'm really looking for in most cases is message integrity (signing), but if you're going to do that, you might as well just encrypt it too and avoid accidents where sensitive information is sent over encrypted channels.
It doesn't matter what is supposed to be on these sites. From security perspective they contain MITM attacker's content. They are effectively an API for issuing arbitrary commands to the browser. To shut down this attack API, all sites have to stop using HTTP, no exceptions.
Not saying you’re wrong, but it’s just a different audience that would be interested in the actual search algorithms.
Elastic is great on-disk, especially on SSDs and avoiding issues like write amplification.
Loading large indexes in memory isn't simple/cheap and when it comes to vectors we're talking apples/oranges I feel. Modern search architectures need to embrace ensemble approaches but boolean-based content searches is often the primary util in enterprise (and search is supplemented by a customizable td-idf). Using vector-based retrieval & similarity is still useful but not something you necessarily need elastic to do for you or couldn't co-exist together.
I'm also suggesting that use cases that aren't enterprise type search problems are more common than you'd think these days.
Edit: Additionally the thing here is you have classical boolean search systems like ES and vector search solutions like Milvus, but no one's gotten around to making something that does both well, from what I can see a lot of the players in the space are trying to go in that direction but it's a slow painful crawl that results in this type of situation where we had to do a lot of custom gluing of these systems together and keeping that parity that was super annoying and expensive, and time consuming, but not necessarily performance inhibiting.
It looks a lot like this: https://huggingface.co/blog/bert-cpu-scaling-part-1
We have to store large "index" embeddings on SSDs and use leveldb for value retrievals of the lucene results.
At sajari.com we have been working on an experiment that uses a 1 cpu machine on cloud run to serve a neural network generated, hash based index of an old BestBuy catalog (25k products). Retrieval uses an approximate nearest neighbour (ANN) look up which typically takes ~1msec. Speed and relevancy are already pretty good.
But we have also learned that there is no one silver bullet and we have seen the best results when combining neural search with traditional keyword search and reinforcement learning.
You can take a look at the demo here: http://neural-hashes.sajari.com
Be gentle, this is an experiment and not a production scale implementation.