Alexandria Search
alexandria.org
alexandria.org
I've realized my searching is basically optimized for google and the web that has grown up around it. Also, in 1998 I wasn't as aware of what was out there as I am now. It's pretty rare (even if its possible) that I do a search and come across a completely new site that I haven't heard of before, for anything nontrivial. That was different when search began.
Google is now almost a convenience. If I have a coding question, I search for "turn list of tensors into tensor" or whatever but I'm really looking for SO or the pytorch documentation, and I'll ignore the geeksforgeeks and other seo spam that finds it's way it. It's almost like google is a statistical "portal" page, like Yahoo or one of those old cluttered sites were, that lets me quickly get through the menus by just searching. That's different from a blank slate search like we might have done 25 years ago.
I think what's really lacking now is uncorrupted search for anything that can be monetized. Like I tried to search for a frying pan once on google and it was unusable. I'm not sure any better search engine can fix that, that's why everyone appends "reddit" to queries they are looking for a real opinion on, again, because they are optimizing for the current state of the web.
Anyway, all that to say I think there are a lot of problems with (google dominated) search, but they are basically reflected in the current web overall, so just a better search engine, outside of stripping out the ads, can only do so much. Real improved search efforts need to somehow change the content that's out there at the same time as they improve they experience, and let us know how to, in a simple way, get the most out of it. I think google has a much deeper moat than most people realize
I believe people will at least start looking for alternatives. For example, I have been collecting search engines, and whenever I encounter a page with too many commercial-laden SEO-porked results, I use a different search engine in Firefox.
I have enabled the Search Bar, I can do Alt+D, Tab, Tab, enter my query, then click a different search engine, which searches instantly, unlike the main bar, where you have to press Enter once more after clicking.
I just added this one also. See my collection: https://0bin.net/paste/ZSCRYVx1#sxD+jBIpScJismXBYwaoJPh75TH9...
I had to do a `pip install lz4` then apply and run this change: https://gist.github.com/Tblue/62ff47bef7f894e92ed5?permalink...
And it did not work. I hit a brick wall. I completely lost trust in Firefox. I want a browser created by a non-profit. Thank you Google for corrupting everything you touch.
I then found this blessed soul: https://www.jeffersonscher.com/ffu/searchjson.html
And here are the ones not built-in:
https://metager.org/meta/meta.ger3?eingabe={searchTerms}
https://wiki.archlinux.org/index.php?title=Special:Search&se...
https://www.mojeek.com/search?q={searchTerms}
https://www.openstreetmap.org/search
https://pypi.org/search/?q={searchTerms}
https://www.qwant.com/?q={searchTerms}&client=opensearch
https://swisscows.com/opensearch.xml
https://yandex.com/search/?text={searchTerms}&from=os&clid=1...
https://whoogle.sdf.org/search
https://en.wiktionary.org/w/api.php?action=opensearch&format...
https://github.com/search?q={searchTerms}&ref=opensearch
http://www.urbandictionary.com/define.php?term={searchTerms}
http://alternativeto.net/browse/search/?q={searchTerms}
https://search.marginalia.nu/search?query={searchTerms}&ref=...
https://kagi.com/search?q={searchTerms}
https://you.com/search?q={searchTerms}
https://millionshort.com/search?keywords={searchTerms}&remov...
https://search.f-droid.org/?q={searchTerms}&opensearch=1
http://teclis.com:1333/?topics={searchTerms}
Many more installable engines are available at https://mycroftproject.com as OpenSearch XML plugins, compatible with Firefox and discoverable by Chromium.
Is it just me, or I feel like Google does not provide anymore good results for me.
Like every time I search something completely out of my knowledge, like "How to purchase a property in Mexico", it will give me 100+ results of some results with autogenerated content like "10 best places to buy property in Mexico". And the only way to fix that would be to add something like `site:reddit.com`
If all websites try to optimise for SEO, they undermine the assumption that the evaluation of a search engine is the pure consequence of how well a site satisfies a query.
I think it's very possible that we have effectively raised the noise floor so high that there is no signal, but also likely that perverse incentives from trying to profit off of search engines have made them our enemies instead of our friends.
For instance, does Google favor sites that run their own tools on them? I've stopped paying attention but recall hearing mutterings to that effect. If so then running the tools is a protection racket.
For other perverse incentives: if you try to rank sites by how long someone stays on them before backing out, or searching again, then you end up favoring rabbit-hole sites, that either string you along or suck you into a tangent. "Oh, this must have answered their question about keeping bees," no, they're reading gossip about the Queen of England and have forgotten all about beekeeping.
I am starting to suspect that there might be nothing to find.
I just don't think people (other then the tech oriented) are creating websites and running forums - and why would they? Reddit might be be only place you _can_ find that type of content. What should search engines do then?
With a tiny number of exceptions, it might be that people chat on reddit, read Wikipedia, ask questions on the stackexchange network/Quora, local communities use facebook groups, and businesses have a wordpress site with nothing more then a bit of fluff, a phone number and an email address.
Speaking about myself; I cold turkey migrated to DDG ~2 months ago. So far I've had to resort to Google search 10 times or so.
One thing I miss though is Google's nice visualisation of fast changing results e.g., match scores. For example: https://imgur.com/a/Q5nZkjo
I'd like to try them out, could you mention which?
Try comparing results with a Bing-based engine (e.g. DuckDuckGo) or a Google-based one (e.g. StartPage, GMX) to see if they differ. (Don't use Google or Bing directly, since results will be personalized based on factors like location, device, your fingerprint, etc.).
The biggest thing I've found is that when doing technical searches it always turns up the sources I'm actually looking for, actively filtering all the GitHub/StackOverflow copycat sites.
It also seems to up-weight official docs compared to Google. For example, "how to read a json file in python" turns up the Python docs as the second result, where in Google they're nowhere to be found.
I first discovered Alexandria in early February: https://git.sr.ht/~seirdy/seirdy.one/commit/935b55f10f9024ee...
Around the same time, I also discovered sengine.info, Artado, Entfer, and Siik. By sheer coincidence they all were mentioned to me or decided to crawl my site within the same couple weeks. So yes, from my perspective there have been more than a few new smaller engines getting active on the heels of bigger names like Neeva, Kagi, Brave Search, etc.
1. free version is curated topic filters that sit on top of Google -- best on laptop or desktop at moment / iterating mobile due to ad splash; it's unclear due to TOU that we can ever make that fully server-side legally, though do so have some things testing to see if free version can be ad-free or more tracking free
2. premium version will be mix of our scraping and Bing, depending on topic - standard web search via Bing + same curation as (1), closer to real-time for us
3. have tested out most other indie indexes or sites, for full web scale, money is on ahrefs or brave giving bing / google run for it
4. our primary emphasis is on bringing back some of the Yahoo! directory or Alta Vista look & feel of drilling into topics, so balancing 1-3 (^) as best we can atm, small team, fully bootstrapped modulo tiny F&F round
5. also @DotDotJames on twitter, still iterating when add team info to site
Instances of the engine indexing locally on the machine where a user is browsing the web.
I'm regularly reminded of this concept of a distributed search index, but at the moment it doesn't seem very likely that it'll gain traction.
Are there search benchmarks to be found somewhere?
There must be. If you want to write a search engine, you need a way to validate the results.
For some standard corpus.
The phrase I was looking for. Thx a bunch! Gonna marginalia that now.
Here is an example (2014) Web track paper of the 23th TREC: https://trec.nist.gov/pubs/trec23/papers/overview-web.pdf (TREC has a plentitude of difference benchmark tasks and you can submit your own: https://trec.nist.gov/pubs/trec29/trec2020.html - recent TREC 2020 papers)
Just to expand, if I want the api reference, say I search for defaultdict (for some reason I like using them but always have to look at the reference), I want the python documentation. I definitely don't want a third party telling me about it.
And if I search a "make list of tensor into tensor" type question, I want SO where someone had asked the same question and got "tensor.stack" as the reply, so I can understand the answer and follow up by looking at the tensor.stack pytorch reference it I want.
Anything else is wasting my time, I think most users with similarly specific queries are not looking for tutorials, they are looking for the names of functions they hypothesize exist, or references. That's why intermediary sites that try to give an explanation are annoying, at least for me.
In other words, I'd rather see these at the bottom of the SERP than the top, but I wouldn't want to completely eliminate them.
To paint with a broad brush, I look at three criteria:
1. Infoboxes ("instant answers") should focus on site previews rather than trying to intelligently answer my question. Most DuckDuckGo infoboxes are good examples of this; Bing and Google ones are too "clever".
2. Organic results should be unique; most engines are powered by a commercial Bing API or use Google Custom Search. Compare results with a Bing or Google proxy (duckduckgo, startpage, etc) to avoid personalized results. Monitor queries over time to see if SERPs change in ways that diverge from Google/Bing/Yandex.
3. "other" stuff. Common features I find appealing include area-specific search (Kagi has a "non-commercial lens" mostly powered by its Teclis index; Brave is rolling out "goggles"), displaying additional info about each result (Marginalia and Kagi highlight results with heavy JS or tracking), user-driven SERP personalization (Neeva and Kagi allow promoting/demoting domains), etc.
And always check privacy policies, TOS, GDPR/CCPA compliance, etc.
> Google is now almost a convenience. If I have a coding question, I search for "turn list of tensors into tensor" or whatever but I'm really looking for SO or the pytorch documentation, and I'll ignore the geeksforgeeks and other seo spam that finds it's way it. It's almost like google is a statistical "portal" page,
I like engines like Neeva and Kagi that allow customizing SERPs by demoting irrelevant results; I demote crap like GFG, w3schools, tutorialspoint, dev(.)to, etc. and promote official documentation. Alternatively, you can use an adblocker to block results matching a pattern: https://reddit.com/hgqi5o
My name is Josef Cullhed. I am the programmer of alexandria.org and one of two founders. We want to build an open source and non profit search engine and right now we are developing in our spare time and are funding the servers ourselves. We are indexing commoncrawl and the search engine is in a really early stage.
We would be super happy to find more developers who want to help us.
I'd consider contributing. Seems you have something here.
It takes us a couple of days to build the index but we have been coding this for about 1 year.
All the indexes are on disk.
Love it. Makes for a cheaper infrastructure, since SSD is cheaper than RAM.
>> It takes us a couple of days to build the index
It's hard for me to see how that could be done much faster unless you find a way to parallelize the process, which in itself is a terrifyingly hard problem.
I haven't read your code yet, obviously, but could you give us a hint as to what kind of data structure you use for indexing? According to you, what kind of data structure allows for the fastest indexing and how do you represent it on disk so that you can read your on-disk index in a forward-only mode or "as fast as possible"?
>> It's hard for me to see how that could be done much faster unless you find a way to parallelize the process
We actually parallelize the process. We do it by separating the URLs to three different servers and indexing them separately. Then we just make the searches on all three servers and merges the result URLs.
>> I haven't read your code yet, obviously, but could you give us a hint as to what kind of data structure you use for indexing?
It is not very complicated, we use hashes a lot to simplify things. The index is basically a really large hash table with the word_hash -> [list of url hashes] Then if you search for "The lazy fox" we just take the intersection between the three lists of url hashes to get all the urls which have all words in them. This is the basic idea that is implemented right now but we will of course try to improve.
details are here: https://github.com/alexandria-org/alexandria/blob/main/src/i...
I actually don't know what roaring bitmaps are, please enlighten me :)
There are some algorithms that have been optimized for intersect, union, remove (OR, AND, NOT) that work extremely well for sorted lists but the problem is usually: how to efficiently sort the lists that you wish to perform boolean operations on, so that you can then apply the roaring bitmap algorithms on them.
Yes our documentation is probably pretty confusing. It works like this, the base score for all URLs to a specific domain is the harmonic centrality (hc). Then we have two indexes, one with URLs and one with links (we index the link text). Then we first make a search on the links, then on the URLs. We then update the score of the urls based on the links with this formula: domain_score = expm1(5 * link.m_score) + 0.1; url_score = expm1(10 * link.m_score) + 0.1;
then we add the domain and url score to url.m_score
where link.m_score is the HC of the source domain.
We hope we can become a useful search engine powered by open source and donations instead of ads.
Does this mean we’re not in Commoncrawl? Or are there any factors you weight much more heavily than Google might?
1. Do you have any plans to support the parsing of any additional metadata (e.g. semantic HTML, microformats, schema.org structured data, open graph, dublin core, etc)?
2. How do you plan to address duplicate content? Engines like Google and Bing filter out pages containing the same content, which is welcome due to the amount of syndication that occurs online. `rel="canonical"` is a start, but it alone is not enough.
3. With the ranking algorithm being open-source, is there a plan to address SEO spam that takes advantage of Alexandria's ranking algo? I know this was an issue for Gigablast, which is why some parts of the repo fell out of sync with the live engine.
4. What are some of your favorite search engines? Have you considered collaboration with any?
1. Yes, any structured data could definitely help improve the results, I personally like the Wikidata dataset. It's just a matter of time and resources :)
2. The first step will probably be to handle this in our "post processing". We query several servers when doing a search and often get many more results than we need and in this step we could quite easily remove identical results.
3. The ranking is currently heavily based on links (same as Google) so we will have similar issues. But hopefully we will find some ways to better determine what sites are actually trustworthy, perhaps with more manually verified sites if enough people would want to contribute.
4. I think that Gigablast and Marginalia Search are really cool and interesting to see how much can be done with a very small team.
Which syntaxes and vocabularies do you prefer? microformats, as well as schema.org vocabs represented Microdata or JSON-LD, seem to be the most common acc to the latest Web Data Commons Extraction Report[0]. The report is also powered by the Common Crawl.
[0]: http://webdatacommons.org/structureddata/2021-12/stats/stats...
Apologies if I missed it (and solely out of curiosity), but how roughly much does hosting Alexandria Search cost (per month)? (I'm assuming you've optimized for cost to avoid spending your own money!)
I have some other questions (around crawlers, parsing, and dependencies), but I need to read the other comments first (to see if my questions have already been answered).
The active index is running on 4 servers and we have one server for hosting the frontend and the api (the API is what is used by the frontend, ex: https://api.alexandria.org/?q=hacker%20news)
Then we have one fileserver storing raw data to be indexed. The cost for those 6 servers are around 520 USD per month.
1. How do you plan to finance?
2. How will you avoid SEO?
3. What kind of help would be most welcome?
1. We would prefer to be funded with donations like Wikipedia.
2. I don't think we can avoid it completely, perhaps with volunteers helping us determine the trustworthiness of websites. Do you have any suggestions?
3. I think programmers and people with experience raising money for nonprofits could help the most right now. But if you see some other way you would want to contribute, please let us know!
Google search is a lot better than people give it credit for.
What Google arguably struggles with is surfacing relevant documents, that is... search.
also adding premium tier that's alerts + ad free + feeling lucky that would take user to top result, which is a UTC page, re: https://breezethat.com/?q=UTC+time
4 of 6 are google's and have to include -- iterating some designs internally that refactor how they're presented on mobile
The ranking would have to vary over an infinite spread of purposes for webpages, and it would have to converge almost perfectly to what is actually most helpful. Among all the technical problems, Google will not optimize correctly against ads for the same reason that websites trying to drum up affiliate purchases and ad revenue won't put content quality above SEO.
When recipes return to having the recipe and ingredients first, followed by an optional life story,I'll revisit my assessment.
Simply put, I believe that Google sucks at search, in the modern context. It is great at indexing, it has solved phenomenal technical challenges, but search it has not solved. Why do I have to write site:stackoverflow.com or site:reddit.com to skip the crap and go to actual content? Why can my brain detect blogspam garbage in 0.5 seconds of looking but billion dollar company Google will happily recommend it as the most relevant result above a legitimate website?
I feel this 12 year old XKCD is still relevant: https://xkcd.com/810/ .
Because the site most likely is laden with Google Ads, it's in their interest to show you that garbage and not what you're actually looking for.
This technology is now part of Google search.
I find it absolutely great to use.
I find its results to be of much better quality than Google.
Google search shows a lot of SEOd results which are absolutely horrible in quality, filled with Adsense ads, and have Amazon affiliate links in them.
Kagi is a breath of fresh air for me.
I also use You.com sometimes. When I am exploring something for the first time, you.com is my place to go. It gives one a good lay of the land which is missing from others.
I still find Google to be the best for looking up code syntax, simple solutions and so on.
But when I am looking for something that depends on opinion, I find Kagi to show much better results and not SEO vomit. (When one wants direct facts even Bing is sufficient.)
On some days my Kagi usage surpasses my Google usage.
I hope this is a project that grows to solve real needs in this space. However, even if it never makes it past this point, there is a chance someone will be inspired by this to construct their own version. Maybe in a different language with a different storage format or a different way of ranking results.
Thank you for sharing your work.
I noticed that if a term can't be found, there will be a random number of results that it says were found, but nothing is actually displayed. Eg: https://www.alexandria.org/?c=&r=&q=moonmusiq
I'll keep trying this out. It seems really promising
https://www.brown.edu/Departments/Portuguese_Brazilian_Studi...
- a Github issue for dask
- an article about panda populations
- some coronavirus article that happens to have an unrelated snippet of pandas code
Google obviously picks the relevant stackoverflow thread as the first response.
For my second I searched my name and got a Wikipedia article about a show I've never heard of which didn't have my name anywhere in it.
For my third I searched "GFlowNetworks" again and it said Found 2,656,844 results in 1.61s, but showed no results again
[0]: https://seirdy.one/2021/03/10/search-engines-with-own-indexe...
It's earned a spot in my search bookmarks alongside Right Dao and Marginalia for when I want a bit of serendipity in my results.
I also appreciate the minimal UI, without any JavScript or other unnecessary bells and whistles.
https://www.alexandria.org/?c=&r=&q=real+estates+puerto+esco... - 3 results only :( If you correct it to "real estate puerto Escondido" - that works better https://www.alexandria.org/?c=&r=&q=real+estate+puerto+escon...
A lot to improve. But a good start
Searching is screamingly fast.
The index seems stale, though. Alexandria, how old is your index?
How long did it take you to create your current index? Is that your bottleneck, perhaps, that it takes you a long time (and lots of money?) to create a Common Crawl index?
Common crawl indexes about once every 40 days, the current crawl's data is through January 2022, so it's 1.5 months old at best.
Alexandria.org is a non-profit, ad free search engine. Our goal is to provide the best available information without compromise.
The index is built on data from Common Crawl and the engine is written in C++. The source code is available here.
We are still at an early stage of development and running the search engine on a shoestring budget.
Please contact us at -email- if you want to get involved, want to support this initiative or have any questions.
> Alexandria.org is a non-profit, ad free search engine. Our goal is to provide the best available information without compromise.
> The index is built on data from Common Crawl and the engine is written in C++. The source code is available (at https://github.com/alexandria-org#).
Edit: formatting
> Common Crawl is a nonprofit 501(c)(3) organization that crawls the web and freely provides its archives and datasets to the public. Common Crawl's web archive consists of petabytes of data collected since 2011. It completes crawls generally every month. ...
I have a test search string that I use to try out search engines, this one didn't do very well.
You might get away with having like a custom micro-index where your search basically does a hidden site:-search for your favorite domain and related domains, but that's not quite going to do what you want it to do.
Ah.
If you try to rent that sort of compute, you're probably looking at like $100-200/month.
So my answer is: if you can do all you need in the base language, something more modern like D-lang is preferred, but if you need some particular library, you either have to add all the bindings yourself, or use C++.
I think, I found a minor bug, while (of course) searching for my homepage:
https://www.alexandria.org/?c=&r=&q=www.hadjian.com
The status line below the search box says, that it found many results, but the results are empty. Also, when hitting F5 a couple of times, the number jumps around.
Keep up the great work. I think there is a lot potential to Common Crawl and things built on top of it.
i am (and most of us are) trying to solve my own issue. not google's.
that's why you get vague answers to your questions. it's not because it doesn't happen. it's because at that moment we care much more about solving our problem. that's what brought us to google search in the first place.
I just searched for python3 join string and I didn't get the Python docs in the first page. Both DDG and Google got them at position 9 which is way too low. At least I got a different set of random websites and not the usual tutorialpoint, w3schools, geeksforgeeks etc that I usually see in these cases.
A few of my test searches came up with very useful results. However, one disappointment was searching for a javascript function, for example "javascript array splice", and the MDN site was not in the results. Adding "MDN" or "Mozilla" to the search did not help either.
I suggest you start by not implementing a crawler but use commoncrawl.org instead. The problem with starting a web crawler is you will need a lot of money and almost all big websites are behind cloudflare so you will be blocked pretty quickly. Crawling is a big issue and most of the issues are non-technical.
Some sort of partnership between crawlers could go a long way. Have you considered contributing content back towards the Common Crawl?
This seems like a reasonable fallback option but it's also a weaker one. By "most of the issues are non-technical", do you mean that you need special permission from someone like cloudflare to get "crawl rights"?
eTools.ch uses commercial APIs so it doesn't get blocked, but it might block you instead (very sensitive bot detection).
Dogpile is one of the older metasearch engines, but I think it only uses Bing- and Google-powered engines.
also part of why going with topic filter approach at Breeze -- so if you search all the web, we'll give you option to open others with indie indexes in new tab - say Mojeek or Yandex -- similar to what airline search engines do
if you switch to say code search, you can use ours or redirect to any one of say PublicWWW, Nerdy Data, or Builtwith
pretty much no other legal way to do it on the main
At the moment we primarily need help with development and funding. But if you have suggestions or want to help in some other ways, please let us know!
https://stackoverflow.com/questions/44077294/encounter-this-...