Writing your own search engine is hard (2004)
queue.acm.org
queue.acm.org
Important sites have a bunch of anti-crawling detection set up (especially news sites). It's even worse that the best user-generated content is behind walled gardens in facebook groups, slack channels, quora threads, etc...
The rest of the good sites are javascript-heavy and you often have to run chrome headless to render the page and find the content - but that is detectable so you end up renting IP's from mobile number farms or trying to build your own 4G network.
On the upside, https://commoncrawl.org/ now exists and makes the prototype crawling work much easier. It's not the full internet, but gives you plenty to work with and test against so you can skip to the part where you figure out if you can produce anything useful should you actually try to crawl the whole internet.
Something is deeply wrong with such an adversarial ecosystem. If sites don't want to be found and indexed why go to any effort to include them? On the other hand there are millions of small sites out there keen to be found.
The established idea of a "search engine" seems stuck, limited and based on some 90's technology that worked on a 90's web that no longer exists. Surely after 30 years we can build some kind of content discovery layer on top of what's out there?
I work on a small - medium ecommerce website and my code just... sucks. I kind of don't want to admit it but it is true. When there is some Chinese search engine that tries to crawl all the product detail pages during the day (presumably at night for them?), it slows down the site to a crawl. I mean technically I should have the pages set up so they can't pierce through the cloudflare cache but it is easier to just ask cloudflare to challenge user (captcha?) if there are more than n (I think currently set to something small like ten) requests per second from any single source.
I don't understand all the business decisions but yeah, I'd suspect the biggest reason is we simply have poor codebases and can't spend too much time fixing this while we have so many backlog items from marketing to work on...
Webserver was a single VM running a Java + Spring webserver in Tomcat, connecting to an overworked Solr cluster to do the actual faceted searching.
Caches kept most page loads for organic traffic within respectable bounds, but the crawler destroyed our cache hit rate when it was scraping our site and at one point did exhaust a concurrent connection limit of some kind because there were so many slow/timing-out requests in progress at the same time.
Does the served HTML contain everything: the product, related products, comments, reviews, etc? If so, caching the entire page might be counterproductive.
But if the page is designed in such a way where it's broken up into fragments, then caching heavily is your friend.
Those fragments can either be loaded client-side (Javascript) or server-side (ESI, SSR, etc) -- but page fragments are key to increasing how much and how long you can cache.
And with this, you've added a layer of complexity that a small to medium e-commerce site may not be able to fulfill with their in-house talent.
It could look something like this: https://web.archive.org/web/20000302042007/http://www1.yahoo...
You could even add the ability to discuss these links, and add a similar voting system to those discussions.
Feels like accepting user content no longer flies on the web in 2022 due to the effectiveness of bot spam and lack of viable countermeasures :-/
Could probably back it off some human-readable format in a git repo or something, pull requests might help with moderation, and recruiting volunteers might be easier if the data is open and available/forkable for other projects.
Then they should treat all bots equally and block Google as well. If they block Google as well, then yes, we should leave them alone.
Why give unfair treatment to Google? That's anti-competitive behavior and it just prevents new search engines from being created.
What do you think they have against smaller search engines? I can't quite fathom the motives.
Small search engines getting blocked is just collateral damage from websites trying to solve this problem with a blunt ban-hammer.
I think this is the inevitable result of surveillance capitalism. Once you can reliably be identified as holding certain views or meeting some arbitrary criteria pages can be dynamically altered to restrict what you have access to. A racist social platform can look completely benign to anyone who isn't part of the in-group, forums for teens can automatically lock out anyone suspected of being over a certain age, companies can show products or services only to people within a certain class, etc.
We already restrict a lot of access to content based on guesses about location. Youtube doesn't even let you see that certain content on their platform exists on their site unless you're logged in from an IP in specific counties (as opposed to showing you the videos exist and throwing up an error saying it's not available for you when you try to view it) and netflix blocks VPNs just to try to keep a bunch of their content away from the wrong kinds of people.
These are the kinds of barriers the internet should have freed us from, but instead it's being used to put up more gates and to force people into them.
I think it is not about being found. It is more about being copied.
These sites are afraid their content is stolen, so they only allow Google to crawl them.
The news sites etc. want to be indexed by Google, and they whitelist Google so they show up in search results. They don't want to be crawled by random aggregators and scrapers which don't bring them ad views and subscriptions. Indexing for me and not for thee.
Other places like Slack channels are private or semi-private and shouldn't be indexed, I agree. There was some kind of scheme a while back (before Freenode went down the tubes in a completely different and unrelated way) to index Freenode. People rightfully freaked out about it and the scheme was shot down.
Also an interesting blog post here: https://fulmicoton.com/posts/commoncrawl/
- Tries (patricia, radix, etc...)
- Trees (b-trees, b+trees, merkle trees, log-structured merge-tree, etc..)
- Consensus (raft, paxos, etc..)
- Block storage (disk block size optimizations, mmap files, delta storage, etc..)
- Probabilistic filters (hyperloloog, bloom filters, etc...)
- Binary Search (sstables, sorted inverted indexes)
- Ranking (pagerank, tf/idf, bm25, etc...)
- NLP (stemming, POS tagging, subject identification, etc...)
- HTML (document parsing/lexing)
- Images (exif extraction, removal, resizing / proxying, etc...)
- Queues (SQS, NATS, Apollo, etc...)
- Clustering (k-means, density, hierarchical, gaussian distributions, etc...)
- Rate limiting (leaky bucket, windowed, etc...)
- text processing (unicode-normalization, slugify, sanitation, lossless and lossy hashing like metaphone and document fingerprinting)
- etc...
I'm sure there is plenty more I've missed. There are lots of generic structures involved like hashes, linked-lists, skip-lists, heaps and priority queues and this is just to get 2000's level basic tech.
- https://github.com/quickwit-oss/tantivy
- https://github.com/valeriansaliou/sonic
- https://github.com/mosuka/phalanx
- https://github.com/meilisearch/MeiliSearch
- https://github.com/blevesearch/bleve
- https://github.com/thomasjungblut/go-sstables
A lot of people new to this space mistakenly think you can just throw elastic search or postgres fulltext search in front of terabytes of records and have something decent. That might work for something small like a curated collection of a few hundred sites.
- sentiment analysis
- roaring bitmaps
- compression
- applied linear algebra
- ai
In a vent diagram intersecting all of these topics, is search. Coding a search engine from scratch is a beautiful way to spend ones days, if you're into programming.
Probably more like a few million but otherwise 100% true. Once you really need to scale you have to start losing some accuracy or correctness.
It helps that the goal of a search engine is not to find all the results but instead delight the user by finding the things they want.
So Kubernetes or Borg and scalable storage and load balancing and global traffic management (glsb)
This does go some way toward explaining why Google interviews looked the way they've looked. It's just a shame everywhere else has copied their homework, without actually needing the same skills.
He was dumbfounded that i would want to spend two weeks 'tunning Solr queries' for a project. He asked( nay stated)
"Why ? Search is a solved problem ?"
We no longer talk to each other.
The next search engine will never be made by a few kids playing around in the shed.
The latest Common Crawl is roughly 120TB, enough to fit on 48 LTO-6 tapes. LTO-6 is the perfect middle ground expense-wise if you did want to play with the entire dataset at home. The tapes cost ~$30 (AUD) each, a cheap loader from eBay can be had for around ~$1000-4000 (like a Dell PowerVault or Fujitsu Eternus).
Your looking at around $8k just to have your data close to your computer. Either that, or you are going to just run everything on AWS/Google Cloud/Azure. Two of these are your competitors.
No one is going to copy 48 tapes for you, either :)
120TB would take around three months to download - if my ISP didn't cut me off first for probably being the heaviest consumer user of bandwidth.
https://www.chatnoir.eu/ has a search engine that is built from the CommonCrawl data running on Elasticsearch, and it runs on 130 nodes [0].
I would love to be able to run my own search engine, but unless there are a number of breakthrough algorithms, I don't see how it can be easily achieved.
I did ask some guys from the Internet Archive if they would be happy to copy some data for me onto some tapes. I wonder if there is a service for this?
[0] https://webis.de/downloads/publications/papers/bevendorff_20...
I'm also not entirely sure which problem CC is intending to solve. It's not like their data is in any way more complete than my own crawl sets, what I can't access to crawl, they don't seem to crawl either.
ok im ready to be judged!
Sadly it’s a little out of date. I’d love to see a more modern post by someone. Perhaps the authors of mojeek, right dao or someone Elise running their own custom index. Heck I’d pay for some by Matt Wells of Gigablast or those behind Blekko. The whole space is so secretive that for those really interested in the space only crumbs of information are ever really released.
If you are into this space or just curious the videos about bitfunnel which forms parts of the bing index are an excellent watch https://www.youtube.com/watch?v=1-Xoy5w5ydM and https://www.clsp.jhu.edu/events/mike-hopcroft-microsoft/#.YT...
While writing a search engine is hard, it is also incredibly rewarding. Over the past two years, we have brought up a meaningful crawl / index / serve pipeline for Neeva. Being able to create pages like https://neeva.com/search?q=tomato%20soup or https://neeva.com/search?q=golang+struct+split which are so much better than what is out there in commercial search engines is so worth it.
We are private, ads free and customer paid.
We wrote an op-ed last month in Fast Company re: the challenges of crawling at scale in today's world: https://www.fastcompany.com/90759792/with-google-dominating-...
tldr; there are two big challenges in crawling the web:
* quantity -- how do you crawl the web at O(B) pages per day
* quality -- how do you make sure you are crawling the high quality parts
On the quantity side, the web is a much trickier place to crawl than 10-15 years ago. Even if you build a system that is well behaved, respects rate limits and work w/ webmasters and CDNs to make sure you are treated as a good bot, the following things will bite you:
* some sites only allow googlebot/bingbot via robots
* even when sites allow all search crawlers, you'll see 429s, 503s, crawl delay directives that limit you to very low qps, and other mysterious errors. (most retailers and aggregators)
* working w/ webmasters works only if you manage to get a response from them (good luck)
* many sites require JS rendering (which is 100x more expensive in terms of the number of assets you are crawling)
(OTOH, Cloudflare is a force for good: https://radar.cloudflare.com/verified-bots)
On the quality side, crawl prioritization works best when you have click data, which most small search engines don't have enough of. In the absence of that, good seed sets and high quality inlinks are your friend.
In other news, eng blogs are being written up as we speak. Will post here on HN as they roll out. Also, like marginalia_nu pointed out, we share whatever we've been working on on a weekly basis on our Twitter account.
Who pays the Tom's Guide and New York Times "experts" of mattress from your landing page bed query demo?
> One Cuil = One level of abstraction away from the reality of a situation.
> Example: You ask me for a Hamburger.
> 1 Cuil: if you asked me for a hamburger, and I gave you a raccoon.
> 2 Cuils: If you asked me for a hamburger, but it turns out I don't really exist. Where I was originally standing, a picture of a hamburger rests on the ground.
> 3 Cuils: You awake as a hamburger. You start screaming only to have special sauce fly from your lips. The world is in sepia.
> 4 Cuils: Why are we speaking German? A mime cries softly as he cradles a young cow. Your grandfather stares at you as the cow falls apart into patties. You awake only to see me with pickles for eyes, I am singing the song that gives birth to the universe.
http://cuiltheory.wikidot.com/
Two years later, in 02010, the founders shut down the search engine, laid off all the employees, sold its patents to Google, and became Google employees again. Now all that remains of Cuil is the Cuil Theory Wiki.
I'd be really interested to see a retrospective on what went wrong. I guess writing your own search engine really is hard, but I'd like to know what turned out to be so much harder than they expected.
Or put differently, if you are going to replicate what existing search engines already do, you are probably not going to be as good initially and you are going to struggle making money. Fixing the money part is the hard part.
Don’t do page rank initially. Actually don’t do it at all. For this observation I risk being inundated with hate mail, but nonetheless don’t do page rank. If you four guys in your garage can’t get something decent-looking up without page rank, you’re not going to get anything decent up with page rank. Use the source, Luke—the HTML source, that is. Page rank is lengthy analysis of a global nature and will cause you to buy more machines and get bogged down on this one complicated step—this one factor in ranking. Start by exploiting everything else you can think of: Is the word in the title? Is it in bold? etc. Spend your time thinking about anything you can exploit and try it out.
The web is full of sites that want to rank, since traffic makes money and appearing high in search results gets you traffic. Simple handling of HTML source is incredibly gameable, and while it might have worked okay on the web of 2004 it definitely is not enough now. It's you and your team verse an enormous number of SEO people.
1. Synonym collapsing isn't even necessarily a good thing, As Clay Shirky noted in "Ontology is Overrated -- Categories, Links, and Tags", people who enjoy watching movies and people into cinema are different crowds. https://oc.ac.ge/file.php/16/_1_Shirky_2005_Ontology_is_Over...
Bandwidth: This is now also cheap; my residential service is 1 Gbit. However, the suggestion to wait until you’ve got indexing working well before optimizing crawling is IMO still spot-on; trying to make a polite, performant crawler that can deal with all the bizzare edge cases (https://memex.marginalia.nu/log/32-bot-apologetics.gmi) on the Web will drag you down. (I bypassed this problem by starting with the Stack Exchange data dumps and Wikipedia crawls, which are a lot more consistent than trying to deal with random websites.)
CPU: Computers are really fast now; I’m using a 2-core computer from 2014 and it does what I need just fine.
Disk: SATA is the new thing now, of course, but the difference these days is HDD vs SSD. SSD is faster: but you can design your architecture so that this mostly doesn’t matter, and even a “slow” HDD will be running at capacity. (The trick is to do linear streaming as much as possible, and avoid seeks at all costs.) Still, it’s probably a good idea to store your production index on an SSD, and it’s useful for intermediate data as well; by happenstance more than design I have a large HDD and a small SSD and they balance each other nicely.
Storing files: 100% agree with this section, for the disk-seek reasons I mention above. Also, pages from the same website often compress very well against each other (since they’re using the same templates, large chunks of HTML can be squished down considerably), so if you’re pressed for space consider storing one GZIPped file per domain. (The tradeoff with zipping is that you can’t arbitrarily seek, but ideally you’ve designed things so you don’t need to do that anyway.) Also, WARC is a standard file format that has a lot of tooling for this exact use case.
Networking: I skipped this by just storing everything on one computer; I expect to be able to continue doing this for a long time, since vertical scaling can get you very far these days.
Indexing: You basically don’t need to write anything to get started with this these days! I’m just using bog-standard Elasticsearch with some glue code to do html2text; it’s working fine and took all of an afternoon to set up from scratch. (That said, I’m not sure I’ll continue using Elastic: it has a ton of features I don’t need, which makes it very hard to understand and work with since there’s so much that’s irrelevant to me. I’m probably going to switch to either straight Lucene or Bleve soon.)
Page rank: I added pagerank very early on in the hopes that it would improve my results, and I’m not really sure how helpful it is if your results aren’t decent to begin with. However, the march of Moore’s law has made it an easy experiment: what Page and Brin’s server could compute in a week with carefully optimized C code, mine can do in less than 5 minutes (!) with a bit of JavaScript.
Serving: Again, ElasticSearch will solve this entire problem for you (at least to start with); all your frontend has to do is take the JSON result and poke it into an HTML template.
It’s easier than ever to start building a search engine in your own home; the recent explosion of such services (as seen on HN) is an indicator of the feasibility, and the rising complaints about Google show that the demand is there. Come and join us, the water’s fine!
It sounds like this same data is not only compressible, but it also has zero useful signal in it and it can be filtered away.
Have you tried that? Are there articles/approaches that talk about it?
Given this is from 2004 I'm not surprised.
Often those types of pages are difficult to link to as well, as they're often highly stateful, applications rather than documents, and what you're seeing may not be what you get when you click the link. Not really what you want in a search engine.
[1] https://vespa.ai/ [2] https://docs.vespa.ai/en/getting-started.html
Text search on the web will slowly die. People will search video based content, and use the fact that a human spoke the information, as well as comments/upvotes to vet it as trustworthy material. Google search as we know it will slowly die, and then will decline like Facebook. TikTok will steal search marketshare as their video clips span all of human life.