How Algolia Built Their Realtime Search as a Service
stackshare.io
stackshare.io
Part 1: https://blog.algolia.com/inside-the-algolia-engine-part-1-in...
No affiliation, just thought they were really interesting.
For each document, we extract the list of words and build a hash-table that associates words to documents
When all documents are processed, we compute an on-disk binary data-structure containing the mapping of words to documents. This data-structure is the index we will use to process queries.They are one of those great companies that you want to emulate, their entire setup is polished and perfect.
I created an Algolia account once with my personal github account, just for testing purposes. Starting the next day and continuing for a few weeks, I received a stream of emails from their sales people requesting times to schedule phone calls, and telling me how much they can help my company. I had no idea how they knew which company I worked for until I checked my linkedin and found Algoia people were visiting my profile. They have some seriously stalkerish sales practices.
Totally agree that sales teams can overreach though. Perhaps there should be a little “why are you signing up? Personal/work/etc” quiz when you sign up, then emails reduced accordingly.
Using Algolia in a full-page-load HTML request is using a Ferrari do drive to the grocery store. You can, but what a waste
Algolia also implemented a linkedin contact search infinitely superior to Linkedin own search.
I remember thinking that both demos were really brilliant growth strategy because it showed clearly, by contrast, how much the status quo was painful
There is no way to search across multiple comments on the same "story"; I use Google site search for this.
It only counts replies/"comments" for stories, not comments (which would probably be a lot more work).
It differentiates between plural and singular search terms as very different searches.
The only ranking options are by popularity and date (would like to be able to sort results by number of replies too).
Thats a bold move. Afaik the timeouts in Zookeeper can be tuned, no?
If you're looking for a front-end library for autocompletion, Twitter's typeahead.js is a nice one: https://github.com/twitter/typeahead.js
If you want one that works seamlessly with your React setup, React Autosuggest is pretty neat: https://github.com/moroshko/react-autosuggest
Have you considered using FoundationDB as a storage layer to match that feature?
Disclaimer: I have no idea how to build a distributed search engine.
Having said that, exploring a FoundationDB integration is definitely an interesting idea. However, quite a lot of use cases can be served perfectly well with a simple master+slave set-up, so my primary focus is on that until there is enough demand for horizontal scalability. For e.g. Typesense is not a great fit for things like log data that typically need large amounts of storage.
That seems a good recommendation.
Querying on Firestore is really limited though and quite surprising really. They even have geopoint data types but you can't query on them.
I'm sure they'll add a lot more soon but Algolia has been great for filling in those gaps and adding robust search.
I mean, there's the "locate" command, but many people disable it because it's a performance hog. Shouldn't "search" be an integral part of OS and/or filesystem architecture?
There's also been a few attempts at integrating search with the various desktop projects, generally backed by xapian or some other full-term search library. I'm not sure what are currently the best/best maintained options.
I seem to recall gnome "tracker" had the most traction last I checked, not sure if eg the kde project has something similar. Looks like canonical booted tracker from the default install in 18.04 lts:
https://community.ubuntu.com/t/install-tracker-by-default-in...
Microsoft had its aborted attempt at a new fs built on top of sql server, for "proper" search at the fs level. I'm not aware of any real file systems that do full search out of the box. And I guess it's not clear that it'd be any better than initial indexing+index refresh on change/on a schedule.