Show HN: I'm building a non-profit search engine
github.com
github.com
I've included some search terms below that I've tried - I've not cherrypicked these and believe they are indicative of current performance. Some of these might be the size of the index - however I suspect it's actually how the search is being parsed/ranked (in particular I think the top two examples show that).
> Search "best car brands"
Expected: Car Reviews
Returns a page showing the best mobile phone brands.
then...
> Then searching "Best Mobile Phone"
Expected: The article from the search above.
Returns a gizmodo page showing the best apps to buy... "App Deals: Discounted iOS iPhone, iPad, Android, Windows Phone Apps"
> Searching "What is a test?"
Expected result: Some page describing what a test is, maybe wikipedia?
Returns "Test could confirm if Brad Pitt does suffer from face blindness"
> Searching "Duck Duck Go"
Expected result: DDG.com
Returns "There be dragons? Why net neutrality groups won't go to Congress"
> Searching "Google"
Expected result: Google.com
Returns: An article from the independent, "Google has just created the world’s bluest jeans"
e.g. you get the same results for “London” as you do for “London cats”, “London cat rescue” and “London test”.
* Okapi BM25 for determining the relevance of a result to a query.
* TF-IDF for determining the relevance of a term to a document.
* PageRank for ranking domains.
It seems really hard to produce quality search results. Takes a lot of investment. Makes it an expensive product. But no one wants to pay. So selling ads it's the only way forward.
Maybe there's a way to convince people to pay what it takes? I dunno...
I wonder how much money Google search makes per the average user. Is it more than $5/mo?
I think they are $4.95/mo or something? haven't payed a cent yet since there are a few discounts that they do to prompt you to learn how to use it (I really liked that, and def made me more likely to stick with it!)
Or maybe they're average, but you only see the ones where DDG fails.
Next time also try Yandex and Baidu.
Similarly, I do have DDG as my main search on all machines and devices just out of principle, but its region-aware searching (I'm in NZ, and often only want NZ results) are very close to useless in my experience (with NZ as the region ticked, it will still return results from .ca and .co.uk domains, which I would have hoped would be almost trivial to remove), and Google seems much better in this area (but not perfect).
Similarly, there's often technical/programming things I'll search for that DDG doesn't have indexed at all, and Google does.
Google also seems a lot better at ignoring spelling differences (color/colour, favorite/favourite) than DDG, which is often (but not always!) useful.
In contrast, Kagi provides Google-quality results mosts of the time, better-than-Google semi-often, and worse-than-google rarely. They support g!, but I only use it a couple of times a week, usually for site-specific searches.
Additionally, I really like that I am their customer and not their product - incentives are aligned for them to continue respecting my privacy and preferences.
The other 20% I resort to Google are mostly things with a geographical/country context, which DDG really sucks and Google excels.
Many here would pay $5/mo, but probably only a handful would pay $180/year.
Now imagine the hundreds of millions they serve. Safe to say most wouldn't pay even $1/year.
Also, if Google loses even 10% or 20% of their audience, the overall value lost will probably higher due to network effects.
Say for laughs you are only interested in yourself. You put every page by you and about you in the crawler. It will obviously render fantastic results. Using google you would have to start every query with your full name OR user names, whatever you type behind it doesn't even matter, it wont return pages with all keywords. With yacy you just type the query and it will return EVERYTHING. To compare the 2 would be to compare useless with perfection.
"langlands program" (pure mathematics thing): yup, top result is indeed related to the Langlands program, though it isn't obviously what anyone would want as their first result for that search. Not bad.
"asmodeus" (evil spirit in one of the deuterocanonical books of the Bible, features extensively in later demonology, name used for an evil god in Dungeons & Dragons, etc.): completely blank page, no results, no "sorry, we have no results" message, nothing. Not good.
"clerihew" (a kind of comic biographical short poem popular in the late 19th / early 20th century): completely blank page. Not good.
"marlon brando" (Hollywood actor): first few results are at least related to the actor -- good! -- but I'd have expected to see something like his Wikipedia or IMDB page near the top, rather than the tangentially related things I actually god.
"b minor mass" (one of J S Bach's major compositions): nothing to do with Bach anywhere in the results; putting quotation marks around the search string doesn't help.
"top quark" (fundamental particle): results -- of which there were only 7 -- do seem to be about particle physics, and in some cases about the top quark, but as with Marlon Brando they're not exactly the results one would expect.
"ferrucio busoni" (composer and pianist): blank page.
"dry brine goose" (a thing one might be interested in doing at this time of year): five results, none relevant; top two were about Untitled Goose Game.
"alphazero" (game-playing AI made by Google): blank page. Putting a space in results in lots of results related to the word "alpha", none of which has anything to do with AlphaZero.
OK, let's try some more mainstream things.
"harry potter": blank page. Wat. Tried again; did give some results this time. They are indeed relevant to Harry Potter, though the unexpected first-place hit is Eric Raymond's rave review of Eliezer Yudkowsky's "Harry Potter and the Methods of Rationality", which I am fairly sure is not what Google gives as its first result for "harry potter" :-).
"iphone 12" (confession: I couldn't remember what the current generation was, and actually this is last year's): top results are all iPhone-related, but first one is about the iPhone 6, second is from 2007, this is about the iPhone 6, fourth is from 2007, fifth is about the iPhone 4S, etc.
"pfizer vaccine": does give fairly relevant-looking results, yay.
Is there a concern that volunteers could manipulate results through their crawler?
You already mentioned distributed search engines have their own set of issues. I'm wondering if a simple centralised non-profit fund a la wikipedia could work better to fund crawling without these concerns. One anecdote: Personally I would not install a crawler extensions, not because I don't want to help, but because my internet connection is pitifully slow. I'd rather donate a small sum that would go way further in a datacenter... although I realise the broader community might be the other way around.
[edit]
Unless, the crawler was clever enough to merely feed off the sites i'm already visiting and use minimal upload bandwidth. The only concern then would be privacy. oh the irony, but trust goes a long way.
This might even be intentional through robots.txt ... A browser extension that passively crawls visited sites could easily download robots.txt as the single extra but minimal download requirement.
The new fly in the crawler ointment is Cloudflare: If you're not the googlebot and you hit a Cloudflare customer you need to be running javascript so they can verify you're not a bot. It's a continual arms race.
https://developers.google.com/search/docs/advanced/structure...
That little thought experiment is true for many online services, from social networking to (marginally) publishing. But nowhere is it more true than for search results, which differ in two fundamental ways: being text-only, they don't bother me to anywhere near the degree of other ads. And, second, they are an order of magnitude more valuable than drive-by display ads, because people have indicated a need and a willingness to visit a website that isn't among their bookmarks. These two, combined, make this the worst possible case for replacing an ad-based business with a donation model.
The idea mentioned in this readme that "Google intentionally degrades search results to make you also view the second page" is also wrong, bordering on self-delusion. The typical answer to conspiracy theories works here: there are tens of thousands of people at Google. Such self-sabotage would be obvious to many people on the inside, far too many to keep something like this secret.
Are DDG results inferior, for 95% of users no.
Do you have a source for this figure? Maybe it's mostly true for the "average non-tech-savvy user in an English-speaking country", but I've found DuckDuckGo and everything other than Google inferior in many cases, especially when looking for Hungarian content.
Yes, I know verbatim mode exists, but I always forget to enable it, and the setting eventually gets lost when my cookies are cleared or something.
Unfortunately I can't switch to another search engine because in my experience every other search engine has far inferior results, despite not having the annoying behaviors Google does. DuckDuckGo is only useful for !bangs for me.
Is this relevant for non-profit project? Do you pay $30/year for Wikipedia?
And yes.
1. What is the rationale behind choosing Python as a implementation language? Performance and efficiency are paramount in keeping operational costs low and ensuring a good user experience even if the search engine will be used by many users. I guess Python is not the best choice for this, compared to C, Rust or Java.
2. What is the rationale behind implementing a search engine from scratch versus using existing Open Source search engine libraries like Apache Lucene, Apache Solr and Apache Nutch (crawler)?
Apart from that, the misconception that "python is slow" should die :-)
Yeah it's not python that is slow, it's the interpreter.
But I agree, get it working first, then re-implement it in another language if it turns out to be necessary.
Is keeping performance in mind and choosing the tech stack accordingly really a premature optimization?
This might be the most abused phrase in CS history. Perhaps we should add "Premature optimization fallacy" to the list of cognitive errors programmers use as an excuse to not seriously think about performance.
And it's often misquoted, and/or taken out of context. The full quote goes like this: "The real problem is that programmers have spent far too much time worrying about efficiency in the wrong places and at the wrong times; premature optimization is the root of all evil (or at least most of it) in programming."
Back then - this was the 60s, remember - CPU cycles were costly.
But with all theories, it stands its time in the sun. We have the same "problem" today, but in a different form. We call it "agile" today, though; make sure that the customer is happy before the programmer is happy. If the programmer is allowed to spend too much time on trying to become happy, the customer is either gone, or someone else came up with a better solution.
In regards to your specific "is keeping performance in mind and choosing the tech stack accordingly really a premature optimization" question, and keeping in mind OP's endevour, you're on the right spot. But the real question is how programmers _get there_.
By experience.
And in turn, by relating to clients more directly these days, programmers have to adhere to the laws that were separated from them back in the days. And your business isn't worth shit without customers, even if you have the best programmers that can create the best code from day 1.
Hence Knuth's quote, translated: "if you spend so much time on planning your journey that you don't reach your flight, you are getting nowhere."
I remember seeing one more non-profit search engine on HN but can't seem to find it right now.
"Ecosia is a search engine based in Berlin, Germany. It donates 80% of its profits to nonprofit organizations that focus on reforestation" [1]
"80% of profits will be distributed among charities and non-profit organizations. The remaining 20% will be put aside for a rainy day." [2]
"Ekoru.org is a search engine dedicated to saving the planet. The company donates 60% of revenue generated from clicks on sponsored search results to partner organizations who work on climate change issues" [3]
[1] https://en.wikipedia.org/wiki/Ecosia [2] https://ask.moe/ [3] https://www.forbes.com/sites/meimeifox/2020/01/19/how-the-se...
> Ecosia says that it was built on the premise that profits wouldn’t be taken out of the company. In 2018 this commitment was made legally binding when the company sold a 1% share to The Purpose Foundation, entering into a ‘steward-ownership’ relationship.
> The Purpose Foundation's steward-ownership of Ecosia legally binds Ecosia in the following ways: - Shares can't be sold at a profit or owned by people outside of the company and - No profits can be taken out of the company.
https://www.ethicalconsumer.org/technology/how-ethical-searc...
Probably this one? "A search engine that favors text-heavy sites and punishes modern web design" https://news.ycombinator.com/item?id=28550764 (3 months ago, 717 comments)
"To Google" has entered the English lexicon as a verb, but I don't think anybody will ever say they "mwmbled" something.
I get your point but I think we should normalise things that don’t come completely naturally for English speakers.
I wonder if it's possible to take advanage of that type of search by putting a facade in front of the "search engine" and based on the search term and the private local user history, then go direct to a known site, or if it seems a search is needed, go to a specific search engine. This may open up opportunities for say program language specific search engines, or error messages from a program specific search, or shopping for X sites.
One way is we send crawl_task to different to N random nodes and accept one that most similar?
another way could be build messy network to solve messy problem. What we do is build reputation bashed graph network and you accept index from nodes you trust. so people will start un following misbehaving nodes. there is not universalroot view of network instead its dynamic and different from prospective of each node. or it could have one root view if we store reputation data in bchain and with some type quadratic voting to modify the chain. ?
yeah Bitcoin showed us way to build mathematically secure system without any trusted party but it could do that cz problem it was solving is mathematically provable. Problem like collecting indexing crawl data you have to trust somebody.
If the miner, they would be able to deliver any crap so content deliveries would have to be judged in some way and awarded differently.
If the pool provides addresses to crawl, the miner could be given crafted/dedicated URIs from time to time and lack of delivery of Proof Of Crawl could result in a penalty chosen in a way rendering "cheating" unprofitable. But then fresh URIs have to come from somewhere.
The second associated problem is how would one prevent them from appropriating the work of others that would just re-sign it.
One way would be to allow the worker to introduce a few voluntary errors but have a secret joker that allow him to pass the challenge of a failed verification.
One alternative way is based on data malleability. The worker pick a secret one way function and compute is F( data + secretFunction(data,epsilon) ) ~ F(data) and publish the values of the secretFunction(data,epsilon) but not the secretFunction. Only someone with knowledge of the secretFunction can make a claim on the work done. If there is a challenge only the real worker will be able to publish the secret of the secretFunction (Or use some zero knowledge proof to convince you they know it).
A web page isn't an immutable piece of text. It can change on every visit and it can sometimes returns errors.
Once a reference snapshot has been crawled, the indexing task is more easily verifiable.
The crawling task is harder to verify, because external website could lie to the crawler. So the sensible thing to do is have multiple people crawl the same site and compare their results. Every crawler will publish its snapshots (which may contain some errors or not), and then that's the job of the indexer to combine multiple snapshots of various crawler and filter the errors out and do the de-duplication.
The crawling task is less necessary now than it was a few years ago, because there is already plenty of available data. Also most of the valuable data is locked in walled garden, and companies like Cloudflare make the crawling difficult for the rest of the fat tail. So it's better to only have data submitted to you, and outsource the crawling.
If every contributor maintained their own index, then you could reward contributors based on how many hits their index generated.
This would open up the possibility of people maintaining indices for specialized topics that they were experts in, and give the federated search engine a shot at taking on Google.
For most people, the cost of creating and maintaining a website is high. This is why products like Wix and Squarespace exist (and are not cheap).
I am thinking a simple dashboard where anyone could go and curate a list of content they find useful. They could share this with the world.
The interface should be so simple that my parents could use it - and they aren't going to be putting up websites anytime soon.
(But of course it wasn't a volunteer public benefit effort like you describe.)
I also have been wondering how this would play out with some kind of decentralized indexes. The nodes could automatically cluster with other nodes of users sharing the same interests, using some notion of distances between query distributions. The caching and crawling tasks could then be distributed between neighbors.
Crawling is also not as resource consuming as you might think. Sure you can distribute it, but there isn't a huge benefit to this.
Also setup a foundation to guide its development and be able to hire a management team.
The real challenge is not the code development but setting up an organization that will outlast all the challenges that will appear. Wikipedia is the model to follow.
Ideally with explainable AI (XAI) that can tell me WHY is result A ranked higher than result B. I would even pay a monthly subscription to use it.
This way you can escape spammers, powertripping moderators, and the tyranny of the hive mind; it doesn't matter if there's a large population of spammers, shills, and idiots upvoting crap because you set their weights to zero (or negative). In fact, that becomes a feature, because by upvoting crap, they generate a crap filter for you. If the weights are also public, then you can automatically & algorithmically seed your web of trust (simplest algo for sake of example: give positive weight to identities who upvoted and downvoted the same way you did) but you could still override the algo with manually set values if it gave too much weight to bad actors.
Obviously this has privacy implications (all your votes and your network becomes public), and can generate a large dataset (performance challenge, how do you distribute it / give access to it?), so it's far from a trivial project. For the privacy angle, I'd start by keeping identities pseudonymous (e.g. a public key or random id -- you don't know who's behind the identity unless they blurt it out). Furthermore, I think it'd be useful to automagically split your actions across multiple identities so it's harder to link all your activity. I think the system should also explicitly allow switching identities, for privacy but also because sometimes you just want a different "filter bubble" which helps tailor the content you get to what you're looking for. Maybe the network that yields best shopping results isn't the same network that yields best cooking recipes or technical docs.
With this model, everyone is a moderator and everyone can defer moderation to identities they trust, but neither the hive mind nor individuals have the ultimate power to dictate what you see. If you want to read spam or conspiracy theories, you just switch to your identity which upvotes such content and has positive weights towards other identities with similar votes.
I doubt you're going to build this; I doubt people want this. I certainly want it. Maybe one day I'll try, but it probably won't work well without network effects (=reasonably large quantity of users). I just wanted to let you know about the idea because your project is inspiring and inspiring things inspire me to share ideas.. :)
This topic is interesting to me because I'm building a faster search engine for programming queries and trying to solve the core issues that got us stuck with crappy engines.
The "fairest" solution for both sides I can think of is ads which no not send tracking information, and are shown primarily based on search terms and country, or even other parameters that the visitor has set explicitly. Any other ideas on how to finance such an engine so that incentives are aligned?
[0]: EDIT: off-topic because the page clearly states that this project will be financed with donations only.
I think donations are probably workable. It works in the private tracker scene; the larger ones have "donation meters" and never seem to fall behind.
It could also work on a subscription model which is essentially just formalizing the donations and making it easier to plan cash flow.
You'd be surprised how cheap a search engine can be to operate. My search.marginalia.nu has a burn rate of less than $100/month.
1. They somewhat get around this with their maps feature, but their regular search doesn't actually search by area; you always get national websites that optimize the best. That would be a nice feature to have starting out without having to type in the specific area you're looking for.
2. Search results for hotels that actually work! Not only if they're set up on OTA's! This could actually get your search engine some traction as the search engine to go to when making travel plans which would give you a nice niche to start out in.
I've built something similar called Teclis [1] and in my experience a new search engine should focus on a niche and try to be really, really good at it (I focused on non-commercial content for example).
The reason is to be able to narrow down the scope of content to crawl/index/rank and hopefully with enough specialization to be able to offer better results than Google for that niche. This could open doors to additional monetization path, API access. Newscatcher [2] is an example of where this approach worked (they specialized on "news").
Building search engines are cool and fun! They have what seems like an endless source of hard problems that have to be solved before they are even close to useful!
As a result people who start on this journey often end up crushed by the lack of successes between the start and the point where there is something useful. So if I may, allow me to suggest some alternatives which have all the fun of building a search engine and yet can get you to a useful place sooner.
Consider a 'spam' search engine. Which is to say a crawler that you work to train on finding spammy useless web sites. Trust me when I say the current web is a "target rich environment" here. The purpose would be to not so much provide a search engine in total here, as it would be to provide something like the realtime black hole list did for email spam, come up with a list of URLs that could be easily checked with a modified DNS type server (using DNS protocol but expressly for the purpose of doing the query 'Is this URI hosting spam?' in a rapid fashion.
There are two "go to market" strategies for such a site. One is a web browser plugin that would either pop up an interstitial page that said, "Don't go here, it is just spam" when someone clicked on a link. Or a monkey-script kind of thing which would add an indication to a displayed page that a link was spammy (like set the anchor display tag to blinking red or something). The second is to sell access to this service to web proxies, web filters, and Bing which could in the course of their operation simply ignore sites that appeared on your list as if they didn't exist.
You will know you are successful when you are approached by shady people trying to buy you out.
Another might be a "fact finding" search engine. This would be something like Wolfram Alpha but for "facts." There are lots of good AI problems here, one which develops a knowledge tree based on crawled and parsed data, and one which answers factual queries like 'capital of alaska' or 'recipe for baked alaska'. The nice things about facts is they are well protected against the claim of copyright infringement and so people really can't come after you for reproducing the fact that the speed of light is 300Mkps, even if they can prove you crawled their web site to get that fact.
How is it to be funded?
You'd have to find a way to verify reputation to make sure no bad actors could contribute.
Not knowing their implementation details, I’m guessing it could be doable without reinventing much. An oracle could dispatch a P2P archive job to a pool of clients randomly assigned tasks, with both the first to archive and the first to validate being recognized by the swarm somehow, with periodic re-archiving and re-verification, rate adjusted by popularity of site and of search keywords.
my dream, is a distributed/p2p index. each browser contribute to storing part of the overall index, and handle queries coming from other users so that how to fund huge data centers never become a question.
> is the size of the common Web already way too large to play catch up against google/bing at this point?
Probably. But I would prefer a search engine that didn't search the whole web. I would prefer a search engine that searched the sites related to fields that I'm interested in.So I would pay for or donate to a search engine that provided me good results in e.g. software development. They could add additional fields as demand warrants, so long as quality as maintained. I would even like to see a faceting feature, so I could search for e.g. Matrix and get results on the mathematical concept when need be without having to wade through movie review or fiddle with magic search keywords.
> should indexing skip a blog page because a the author usually write about algorithm but here goes on and on about business while sporadically mentioning algorithmic technical aspects?
Yes, because that author's pages wouldn't even be fetched at the point of development that we are discussing. > and what if you want to search about pottery that Sunday morning, turn to Google?
Yes, Google still exists. Why not?I'm somewhat fine with very commercial search engines when searching for something within a field I'm familiar with, e.g I if search for "kubernetes configmap best practice", skimming infomercial sites, ads, junk is trivial. if I search for "baby milk safety" I have no clue how to digest the results list, I'm totally unfamiliar with the media outlets catering to parenting, nutrition or food health, I know absolutely no authors in the field, no renowned media outlets for tips, not even brand reputations to make a somewhat informed decision on what to skim and what to read with attention.
But I don't disagree an engine focusing on specific fields provides great value. and perhaps that's where it should start to have a chạnce to win.
edit / full disclosure: I work at google but nothing related to search
I think OP's point is, assume you only have 10/100 terabytes of space and limited compute ability - how would you approach the problem? I assume 90% of google's searches probably come from less than 1% of their total index, not to mention that Google is also keeping full cached versions of the whole website including images.
The challenge isn't to index the entire web, it's to index the useful parts of it, and I think an index covering most of the useful web can be seeded quite easily with some community effort.
The challenge then moves to the curation, but it's no longer infeasible.
It's got a fairly small index, but yeah, it's not particularly hardware-hungry.
I'm confident 100mn is doable with the current code, maybe .5bn if I did some additional space optimization. There are some low hanging fruit that seem very promising. Sorted integers are highly compressable, and right now I'm not doing that at all.
> What is your current max rps?
It depends on the complexity of the request, and repeated retrievals are cached, so I'm not even sure there is a good answer to this.
I honestly don't know what the actual limit is, all I know is it dealt with 2 QPS without affecting response times. But 2 QPS for a search engine is actually kind of a lot. Most people don't actually search that much. Like you get a few queries per day. Put it this way: 2 QPS is what you'd expect if you had around a million regular users. That's not half bad for consumer hardware.
(1) Investigate if running a DNS server will help me get a more robust picture of what websites exist.
(2) Investigate if supplying a custom browser would help me to leverage client PCs to do the crawling / processing for me.
(3) Investigate if there is any point in building a search engine with the data gathered in a non-proffit way... Non-proffits are not as sustainable as for proffit corporations.
Free and open software is a great ideal, but the reality is that people need money to live - and ads are the way to make that money on web2 platforms, which is why Google is in such a sad state. Why not do something similar to Brave? You can add tokenomics to the search engine and make money while keeping it 100% useable and open source.
Do you think there's an alternative to tokens in order to fund the project without degrading the search algorithm? If so, I'm all ears.
In fact, I think we all are. Please enlighten us. But you haven't proposed a solution while crypto devs have been working on one since 2008.
And where's the community behind mwmbl project?
goal was/is to include all charities in the world, based on open data, and open software.
It's been on pause for a while, but still works, and open for new sources of data to incorporate.
Is there an objective way to measure this? Do we just compare the output to Google or DDG? This seems like one of the many big hurdles in creating a competitive product in this space
And have the large server instances (lots of ram)?
i was spending 2500 a month just on indexers and had query traffic taken off my costs would have shot through the roof since you want query nodes to all be in memory cache and that was expensive back then. today i would have used some modern in memory distributed doc dbs instead of query masters with heavy block buffer caches. i learned a lot but lost my shirt :)
Wondering out loud if there's a crawl that will produce real-time updates via common news outlets..
The 20tb drives would be too slow for query though. Query really needs everything in memory to be fast enough.