You can decompose a "search engine" into multiple big components and figure out what you want to look at first:
(1) web crawler/spiders
(2) database cache of web content -- aka building the "search index"
(3) algorithm of scoring/weighing/ranking of pages -- e.g. "PageRank"
(4) query engine -- translating user inputs into returning the most "relevant" pages
Each technical topic is a sub-specialty and can be staffed by dedicated engineers. There are also more topics such as lexical analysis, distributed computing (for all 4 areas), etc.
If you're mainly focused on experimenting with programming another ranking algorithm, you can skip part (1) by leveraging the dataset from Common Crawl: https://index.commoncrawl.org/
Here are some videos about PageRank:
https://www.youtube.com/watch?v=JGQe4kiPnrU , https://www.youtube.com/watch?v=qxEkY8OScYY
... but keep in mind that the scope of those videos omits all of (1), (2), and (4).
A different "information retrieval" perspective overall would be to take a deep dive into the career of Susan Dumais, http://susandumais.com/
I think a major effort, if its not web, is defining a common data structure to represent the varying source data sets.
Edit to add: OP if you want to study this, start smaller than "search the web" and get the major component concepts down. Expanding to the web will involve distributed systems stuff to scale but there's only like a handful of companies doing that yet there are many more companies that need a search engine on in-house data.
https://en.wikipedia.org/wiki/Apache_Lucene
https://en.wikipedia.org/wiki/Apache_Solr
https://en.wikipedia.org/wiki/Elasticsearch
those who dont know history will repeat it, etc
then maybe the newer stuff like https://typesense.org/
to be clear i dont know these things but thats what I would do and I'd happily read fly.io style blogposts drip feeding out knowledge over time
is what I saw as the primary difference. Whether that's going to pan out in reality as well as it does in HN comments is "the devil's in the details" though
Also, have you thought about partnering with commoncrawl.org? I could see that relationship benefiting both sides: they get fresher indices, you get access to the historical web snaps
The problem is that if you were to build a new search engine from the ground up it will take millions in infrastructure, and a lot of time for you to test one idea. And there are multiple attack vectors to Google's business model (privacy, subscription model, modality, etc.) however you might get the change of testing one of them, and if that fails, starting again is super expensive so you might not be able to get funds to do it.
My approach then became to build something that others can build on top of.
I'm currently using common crawl but my main problem is that I need to build a small toy to test it and even processing common crawl is crazy expensive. Just a single snap are 150 Tb, so this needs to be process on metal, or you're gonna pay a hefty AWS bill.
For that specific search I would start at Wikipedia, but for more general "data search" I lean towards Wolfram Alpha, which has some usability issues, but interesting maths engine for queries. https://www.wolframalpha.com/input?i=Barack+Obama+vs+Donald+...
I’ve been dreaming of an open web index and social graph for more than a decade.
Any company having the data + the algorithm + the presentation layer is way too much power. We can and should split that problem into its separate domains.
I hope you succeed, keep us posted.
You can get a taste of the kinds of things covered in this field by looking at this class site (among many others)
https://web.stanford.edu/class/cs276/
That said, to build a modern competitive search engine, you'll need to look into software engineering, distributed systems, artificial intelligence, machine learning, linguistics, computer vision, signal processing, graph theory, databases, scheduling, machine translation, and FSM-only-knows what else.
Smartness (NLP): https://web.stanford.edu/~jurafsky/slp3/
There are also newer techniques, like "deep learning" for search. Not sure what's a good resource for that (using ML to learn the scoring functions is more than a decade old field (learning to rank), it just surfaced again because of the new ML techniques).
Check the table of contents of those books. Courses which mention chapters or sections from them are probably a good choice.
You can also check Apache Lucene and do a deep dive to see state-of-the-art implementation (Lucene in Action is a good introduction).
Edge and link data were stored in a graph database called dgraph which I would run a spark job to recompute a weighted pagerank once a day to use as a signal in the search process. At some point I was working on moving from statically computed pagerank to dynamic pagerank that would update on the fly as indexing occurred.
All the text content was stored in a special format in s3 to avoid as much data growth as we could while maintaining the content for indexing. An identifier was then inserted into another SQS queue which was processed by our vectorizer pipeline for semantic search, which was another stateless python project using a finetuned version of sentence bert and some very neat lower level execution graph optimizations to get that to run ~very~ fast (which could be several blog posts of its own) wound up being able to pull off ~2m vectorizations an hour for 128 token chunks which was then stored in a vector database called Milvus which uses some very nice algos for approximate nearest neighbor matches.
Then there was another queue that would process the text content and bulk batch insert to a very heavily tuned Elasticsearch cluster (long term the idea was to replace Elasticsearch, in fact I'm writing a rust based search engine but I've since fallen off that onto other things, I'd like to get back to it at some point though.) ES is very good but it requires very good tuning and some clever pre and post processing combined with their inbuilt ranking methods to get compelling results over a very diverse set of data like we had assembled.
We used the semantic search as a prefilter for an actual text search over that Elasticsearch cluster, say return 100k ids from the semantic search and then run full text search over those 100k to reduce the time. At the end with those hundreds of millions of pages we were keeping mostly < 1s response times. One of the tricky things with ES there are the cold starts for non cached queries so you could see periodic spikes for certain classes of queries and I was working on smoothing that out.
We were lucky in that we wrangled up something like 35k in AWS credits so while this was running us about 3k a month we had a bout 10months of lead time and we shut down with like 3 months to go. Both me and my cofounder are very technical and on top of all this we were talking to users, getting feedback and looking for avenues to monetize that weren't just based on ads. We wound up pivoting to strictly technical search for businesses throughout this and were talking with those businesses to provide dev tools with specialized search and documentation capabilities based around our search engine and their codebases. Alas neat tech doesn't pay the bills, so after a few near misses on funding and some early contracts that fell through we shuttered it.
I'm super sure I'm missing details here since this was like 1 year of part time work and 6 months of full time work but hopefully that satisfies the curiosity!
I'm working at an open-source project that builds an AI-powered search framework [0], and I've built some examples in very few lines of code (for searching fashion products via image or text [1], PDF text/images/tables search [2]) and one of our community members built a protein search engine [3].
A good place to start might be with a no-code solution like (shameless self-plug time) Jina NOW [4], which lets you build a search engine and GUI with just one CLI command.
[0] https://github.com/jina-ai/jina/
[1] https://examples.jina.ai/fashion/
[2] https://colab.research.google.com/github/jina-ai/workshops/b...
https://www.cs.virginia.edu/~evans/courses/
I highly recommend digging deeper and trying to find the course materials - I'm pretty sure they should be available somewhere on the Internet. Perhaps the author could point you in the right direction.
Introduction to Information Retrieval: https://www.amazon.com/Introduction-Information-Retrieval-Ch...
Search User Interfaces: https://www.amazon.com/Search-User-Interfaces-Marti-Hearst-e...
IME, analysing alternative "search engines" will reveal most are a waste of time.
Using a text-only browser I get only one search result for each query.
[0] https://faculty.nps.edu/ncrowe/book/book.html [1] https://faculty.nps.edu/ncrowe/index.html
https://blogs.cornell.edu/info2040/2011/09/20/pagerank-backb...
Google's algorithm is indecipherably complex today but in the early days the way search engines more or less worked was they crawled the web and the way they decided to rank the pages was by how many other pages had a URL reference pointing to it.
You can apply this idea today in private (or I suppose public) search engines in the same way to interesting results.
For example a search engine for for scientific papers might use page rank to sort papers that are cited by the most other papers.
Or if you were going to make a search engine for open source projects you could create a page rank algorithm based on what projects had dependencies to other projects.
Part of why Google's algorithm today is more complex than this is that people try to game whatever algorithm search engines commonly use. You may remember back in the 90s and 2000s people would do stuff like put back links to other websites in the source to try to game page rank. Today that kind of behavior has expanded into a whole cottage industry (unfortunately).
What's interesting though is that for a lot of more limited data sets is you have less of that SEO type problem.
Whichever kind you're building good luck! Search engines usually wind up being petty cool (and profitable).
Truth is, this track you'll set up a search engine in about 20 minutes:
Document Store/NoSQL: https://elastic.co https://solr.apache.org
Classical AI: https://youtube.com/playlist?list=PLUl4u3cNGP63gFHB6xb-kVBiQ...
Stat ML: https://www.tensorflow.org/tutorials https://youtube.com/playlist?list=PLoROMvodv4rMiGQp3WXShtMGg...
Frontiers: https://www.ycombinator.com/companies/?query=Search
It's pretty simple: You crawl a website and search for all links, then you crawl all the links from these linked websites and so on.
You can still see some of my code here: https://github.com/Wronnay/search-lib
[0] https://www.tbray.org/ongoing/When/200x/2003/07/30/OnSearchT...
I wrote a small but somewhat complete search engine some time ago.
The steps are basically:
Have a queue with urls.
Download a page from the queue
Use a html parser to remove markup and get the links of the page. Add the links to a queue without the already visited ones.
Use a stemmer to clean up the text (porter stemmer or whatever).
Calculate an inverted index: https://en.wikipedia.org/wiki/Inverted_index
Save the stuff in an appropriate data structure (Hash or Tree).
Write a query engine for AND, OR and (all the words).
Calculate a simple page rank by counting the links to a page.
For just learning purposes it is not that hard but if you want to get all the crazy corner cases of the "real" web you will go insane.
The easy alternative:
Use a lib to crawl a page.
Plonk all the documents text into postgres with full text search or elastic search.
[1] https://github.com/czekster/markov
[2] https://doi.org/10.1137/050623280
EDIT: admitedly there's much more to it
Search engines are a fairly broad topic, and a lot of it depends on the _type of data_ that you want to build a search engine for. If you're looking towards more traditional, Google/Yahoo-like search, Elasticsearch's learning center (https://www.elastic.co/learn) has quite a few good resources that can point you in the right direction. Many enterprise search solutions are built on top of Apache Lucene (including Elasticsearch), and videos/blogs discussing Lucene's architecture (https://www.endava.com/en/blog/Engineering/2021/Elasticsearc...) is a great starting point as well.
Opposite text/web search is _unstructured data_ search, i.e. searching across images, video, audio, etc... based on their semantics (https://www.deepset.ai/blog/semantic-search-with-milvus-know...). Work in this space has been ongoing for decades, but an emerging way of doing this is via a _vector database_ (https://frankzliu.com/blog/a-gentle-introduction-to-vector-d...) such as Zilliz Cloud (https://zilliz.com/cloud) or Milvus (https://milvus.io/). The idea here is to turn the type of data that you're searching for into an high-dimensional vector called an embedding, and to perform nearest-neighbor search on the embeddings themselves.
Disclaimer: I'm a part of the Zilliz team and Milvus community.
I think that Facebook got around this back in the day by having the user's device do the initial scraping. A side effect of this is that sometimes I'd post an article I found, but the preview image was blank because I submitted it too quickly or something.
Until we have a free publicly downloadable cache of all websites, similar to CoralCDN (is this defunct?), then building our own search engines is probably a nonstarter.
For IP blocks, try this: https://github.com/kordless/mitta-screenshot
I asked a question about my ideal search engine a few days ago: https://news.ycombinator.com/item?id=32452318
Basically, a conversational semantic search engine like the ones in Star Trek. I also listed some technologies that can possibly help implement it. You might be interested in that list.
I've converged on a design that is actually almost eerily similar to Google's original design. ( http://infolab.stanford.edu/~backrub/google.html )
https://www.amazon.com/Managing-Gigabytes-Compressing-Multim...
Maybe instead, we need people to answer questions as real people... and then score/weigh/rank those answers/people.
I just think Google is doing an amazing job for the web, and where it is failing is not with web search, but high quality answers from humans (and not marketing trash which Google is trying to fix this week)
Let's start right away!
urls_old = []
urls_new = [ 'news.ycombinator.com' ]
while urls_new.length>0:
spider(urls_new[0])
... to be continued ... who writes the next 5 lines?I actually did kickstart projects that became pretty big lifestyle businesses by saying "No, you will not figure out how to raise venture capital. Let's open a text editor and start coding".
And it would be funny, if the massive, collective coding power of HN could collaborate and write a search engine right here on the spot. By everyone contributing 5 lines :)
Is perfectly reasonable conversation.
The simplest and the best way to get xp is to get hands dirty
- Tries (patricia, radix, etc...)
- Trees (b-trees, b+trees, merkle trees, log-structured merge-tree, etc..)
- Consensus (raft, paxos, etc..)
- Block storage (disk block size optimizations, mmap files, delta storage, etc..)
- Probabilistic filters (hyperloloog, bloom filters, etc...)
- Binary Search (sstables, sorted inverted indexes, roaring bitmaps)
- Ranking (pagerank, tf/idf, bm25, etc...)
- NLP (stemming, POS tagging, subject identification, sentiment analysis etc...)
- HTML (document parsing/lexing)
- Images (exif extraction, removal, resizing / proxying, etc...)
- Queues (SQS, NATS, Apollo, etc...)
- Clustering (k-means, density, hierarchical, gaussian distributions, etc...)
- Rate limiting (leaky bucket, windowed, etc...)
- Compression
- Applied linear algebra
- Text processing (unicode-normalization, slugify, sanitation, lossless and lossy hashing like metaphone and document fingerprinting)
- etc...
I'm sure there is plenty more I've missed. There are lots of generic structures involved like hashes, linked-lists, skip-lists, heaps and priority queues and this is just to get 2000's level basic tech.
If you are comfortable with Go or Rust you should look at the latest projects in this space:
- https://github.com/quickwit-oss/tantivy
- https://github.com/valeriansaliou/sonic
- https://github.com/mosuka/phalanx
- https://github.com/meilisearch/MeiliSearch
- https://github.com/blevesearch/bleve
- https://github.com/thomasjungblut/go-sstables
A lot of people new to this space mistakenly think you can just throw elastic search or postgres fulltext search in front of terabytes of records and have something decent. The problem is that search with good rankings often requires custom storage so calculations can be sharded among multiple nodes and you can do layered ranking without passing huge blobs of results between systems.
Source: I'm currently working on the third version of my search engine and I've been studying this for 10 years.
This might be something to build on or explore.
Books
* Lucene in Action - A bit out of date, but it has a lot of the core concepts of Lucene. 10 years ago, this was the go-to book even if you’re working with Solr or Elasticsearch, to understand the core data structures
* Relevant Search (my book) - Introduction to optimizing the relevance of full text search engines. Getting a bit old now, but still relevant for classic search engines.
* AI Powered Search (disclaimer - I contributed) - Author Trey Grainger is brilliant, and has been a long-time colleague of mine. He managed the search team at Careerbuilder (where we did some work together). This is in some ways his perspective on how machine learning and search work together.
* Elasticsearch The Definitive Guide - online free, very comprehensive, book from Elastic
---
Blogs
* OpenSource Connections (my old company) Blog (http://o19s.com/blog) - lots of meaty search and relevance info
* Query Understanding (https://queryunderstanding.com/) - Daniel Tunkenlang is a long term very smart search person. Has worked at Google, etc
* James Rubenstein’s Blog (https://jamesrubinstein.medium.com/) - I worked closely with James at LexisNexis. He has helped work on search and search evaluation at Ebay, Pinterest, LexisNexis, and Apple
* Sematext (https://sematext.com/blog/) - Sematext are probably best known for being really in the weeds search engine engineers, a fair amount of scaling, etc. But some relevance.
* Sease Query Blog (https://sease.io/blog-2/our-blog) - Sease are London based Information Retrieval Experts
* My Blog (http://softwaredoug.com)
---
Paid Trainings
* OpenSource Connections Training (https://opensourceconnections.com/training)
* CoRise (https://corise.com/course/search-with-machine-learning)
* ML Powered Search (My Sphere Class ) - https://www.getsphere.com/ml-engineering/ml-powered-search?s...
* Sease Training - https://sease.io/training
* Sematext Training - https://sematext.com/training/
---
Conferences
* Haystack - http://haystackconf.com
* MICES - http://mices.co (e-commerce search)
* Berlin Buzzwords - search, scale, stream, etc - https://2022.berlinbuzzwords.de/