The Web is missing an essential part of infrastructure: an open web index
arxiv.org
arxiv.org
> What is Common Crawl?
> Common Crawl is a 501(c)(3) non-profit organization dedicated to providing a copy of the internet to internet researchers, companies and individuals at no cost for the purpose of research and analysis.
> What can you do with a copy of the web?
> The possibilities are endless, but people have used the data to improve language translation software, predict trends, track the disease propagation and much more.
> Can’t Google or Microsoft just do that?
>Our goal is to democratize the data so everyone, not just big companies, can do high quality research and analysis.
Also DuckDuckGo founder Gabriel Weinberg expressed the sentiment that the index should be separate from the search engine many years ago:
> Our approach was to treat the “copy the Internet” part as the commodity. You could get it from multiple places. When I started, Google, Yahoo, Yandex and Microsoft were all building indexes. We focused on doing things the other guys couldn’t do. [2]
From what I remember reading once DuckDuckGo doesn't use Common Crawl though.
[2] https://www.japantimes.co.jp/news/2013/07/28/business/duckdu...
While you say "news" really that covers any information about current-ish events. It's not just "what happened today" but background on things like the Muller report right now.
Any technology release, or update.
Reviews of any hardware or software.
Information about security vulnerabilities.
Film reviews.
Game reviews.
Book reviews.
New scientific publications.
What you’re describing is just news on the latest updates representing a small slice of the market. Making the remander far from useless.
Although, that said, with Google personalising search results the top result is very likely to be the user's preferred site anyway. We can't have people seeing outside their filter bubble after all.
edit: yep here it is https://doi.org/10.1145/3184558.3191636
A site that gets lots of links quickly (and is therefore important) will likely garner them from sites you are already frequently visiting.
There are a significant number of crawlers out there that don't respect robots.txt. The usual response to them isn't to roll over dead, it's to get CloudFlare (on the technological end) and/or sic the lawyers on them (for CFAA, IP, or ToS violations).
We've been around for about a decade. IBM watson used us as their social data provider during Jeopardy. We provide data to tons of companies and you're probably using our services - just that it's not obvious where we're used since it's SaaS B2B and not B2C.
We're not free but the primary reason we exist is that other vendors charge borderline extortionate pricing and I fundamentally believe that the web MUST remain open.
We've also been providing data for very affordable pricing to researchers for more than a decade.
Search for us as Spinn3r under Google Scholar (our previous name) and we have hundreds and hundreds of PhDs who have access to our data.
We do charge for research usage now but it's very very very affordable.
The entire point is that we're trying to enable innovation.
Yet the company first mentioned does it for free, lol:
I've checked Datastreamer.io for 5 seconds, I don't see any link to their repo. If not "open source" then what does "open" mean?
> Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize.
Your comment makes zero sense in this context, because it's just marketing.
> we're trying to enable innovation.
You're trying to make profit, like every other company in the world and that's OK.
Common Crawl (non-profit): Stores regular, broad, monthly crawls as WARC files. Provides a separate index that can be used to look data up (no a fulltext index though). Used mostly in academia.
Mixnode (for-profit): Regularly crawls the web and lets users write SQL queries against the data. Not sure who the primary users are since it's in private beta.
There are some search engine APIs, but I don't think the conflict of interest would allow for cost-effective large-scale access and pricing...
Not for existing search machine providers, but I think there is room for new players to do this large scale. Imagine an AWS service that high performance access to crawled data as well as a number of indexes and a fairly simple search engine using this data. That would commoditize one of Google's biggest advantages, and anyone could, at least in principle, run their own search engine from the data. Because the market for this is much wider than traditional search engines just providing the data and indices for a pay-as-you-go fee could still be very profitable.
I'm sure there are tons of obstacles to that path, but it also would be far ahead of any new initiative in at least two ways: it already has a huge index and ingestion pipeline, and it is a trusted organization.
I like a modified version of this. I think that it should be a p2p technology and not try to create one meta-index but rather be many domain-specific ones, with one or more tools or DBs to select which indices to search given a query/context.
Are there any decentralized alternatives to Google out there already?
I think that also this overlaps with the idea of moving from a server-centric internet to a content-centric internet.
EDIT: I waited a few minutes and now the results are MUCH better! I think I just needed to let it connect to more peers or something.
in fact it is an aggregate or meta-search that sends proxy requests to user-selected search engines (with defaults varying from instance to instance).
a list of instances[2] is available via the source git repository. i would recommend a few[3] myself.
as of yet searx does not do some things we might want done:
a) original indexing b) federation between cooperative instances c) offer a spec for archiving data[4]
i think we'll get to something like this soon. there are a lot of pieces in play and it falls to all of us - users, hackers, developers - to participate in development, curate adjoining projects, donate time to test & halcyon & on & on.
as always, it will be interesting to see what we all come up with.
[01.0] https://asciimoo.github.io/searx/ [01.1] searx is copylefted floss via GNU Affero GPL3
[02.0] https://stats.searx.xyz/
[03.0] https://search.disroot.org/ [03.1] this organization respect's EFF's Do Not Track [03.5] https://searx.prvcy.eu [03.6] a secondary useful for reasons indicated by the URI
[04.0] this is where the submitted comes into play. e.g. [04.0] should we develop some sort of open API for domains [04.0] to request archiving? this could take multiple forms [04.0] as a project but as long as it's floss and has RFC...
1. spam
2. child pornography
3. content against the laws
The only thing that is easy to define as policy is #2. No one likes child porn. But even then, there are grey areas with differing legal status - lolicon on the anime side and "barely legal" on the realistic side, plus CGI.
Spam - for me I'd flag all commercial advertisings as spam, others would heasitate to block Viagra spammers.
Then the final category: illegal content. The US doesn't like nipples. Germany has no problem with nipples. Swastikas and other NS insignia? Other way around. Some post-Soviet states have banned Hammer and Sickle or the Red Star. Some countries have extremely strict libel laws, others have non-existing libel laws. In some countries (hello Germany) even linking to illegal content can get you thrown into jail, in others not.
And finally: who should pay for operational costs of such an index? Wikipedia only works out because the contributors worldwide donate enormous amounts of time to it, and Wikipedia has only a fraction of the amount of content that Youtube and Twitter create, and Facebook is orders of magnitude bigger.
I hope you do realize the contradiction.
There are far more than a 'few thousand pedophiles'; that number is more reflective of the number of convictions each year. While drawing statistical inferences is difficult, the stats in the appendices to this report suggest there's perhaps ~100k tips a year to police about child sexual abuse across the US.
Filtering out spam, pornography, and other undesirable or illegal content would be done at the service level, i.e., but companies/organizations building user-facing search applications on top of the index.
No, it has to be done before, at the infrastructure level. There are jurisdictions (Germany, for one!) where even the storage or publication of links can be illegal under certain circumstances. With the new GDPR law and whatever is coming up in the US, the situation is even more unclear as it is trivial to embed protected personal data into URLs.
A few of us out there are also working on small directories:
* https://href.cool (mine)
The thought is that you can actually navigate a small directory - they don't need to be five levels deep - and a network of these would rival a huge directory, avoid centralization, editor wars, single point of failure.
The benefits of this would be a standard for Wysiwyg editors (goodbye million rich text editor projects, Markdown and even Microsoft Word), and more semantic markup for both search engines and accessibility.
Right now it takes millions of man hours to create a performant browser, which is limiting those engines to only the largest organizations. Even Microsoft gave up making their own. And even with all that effort, I still can't create a clean HTML document with an interface as rich as MS Word, or even add bold or color formatting to a Twitter post, or update a Wikipedia page without knowing wiki markup.
We need to pull the dynamic, JS powered side of the web out from the core, limit CSS to non-dynamic properties, and standardize on an efficient in-document binary storage akin to MIME email attachments so HTML docs can be self-contained like a Word or PDF doc.
This document-centric web could be marked off within a standard web page, so you could combine it in regular interfaces for things like social network posts. Or it could be self standing, allowing relatively large sites to be created with indexes, footnotes, etc., but served from a basic static server.
This isn't a technical challenge, it's an organizational one. I've thought for years that Mozilla should be doing this, instead of messing with IoT and phones, etc. It's such an obvious problem that needs addressing, and would have a huge payback in terms of advancing the web as we know it.
No, it's an economical one. Who will use that web? You mention Twitter, yet are they not dependent on JS for analytics and ad-tracking? The few sites not dependent on such features are already usable on Lynx and Elinks, and the others simply won't use them.
For the advantages, you mention having a good WYSIWYG editor, but the reason you can't add bold or color to a Twitter post is obviously not because they are unable to add those functions, but because they don't want you to do that. Which raises the question: what happens when that editor lets you create something the site doesn't allow you to use?
(By the way, Wikipedia has had a visual editor since 2012, you just have to switch using the "pencil" button: https://en.wikipedia.org/wiki/Wikipedia:VisualEditor)
If you look at the original Yahoo Page when Yahoo first started out it attempted to solve this problem.
I believe this index could be regionally or language based...
In the United States one could use
Dewey Decimal
https://en.wikipedia.org/wiki/Dewey_Decimal_Classification
Library of Congress
https://en.wikipedia.org/wiki/Library_of_Congress_Classifica...
I'm not saying it shouldn't be done but I think it will be way more work than expected and there will be all kinds of issues.
Our project LearnAwesome[1] currently relies on volunteers to curate topics, but classfication / ontology engineering is in fact seems to be a hard problem.
I tried to simplify all data into ~30 categories. My own interests fit into 16, so I drew a visual representation of them. https://github.com/peterburk/sortlikes
Next, I need to figure out the sub-categories. Genres for music, countries for travel, etc.
What interests me most is the cross-cultural connections. For example, Taiwanese punk rock (Fire Ex), or Mongolian folk metal (Nine Treasures, Hanggai). I like that music because it's the same sub-category I'm interested in (Music/Rock).
It's also possible to model the flow of finance around the world through this categorisation. Some of the categories are innately human and don't seem to exist in animals (music, cooking).
Email me if you'd like to chat more about how to categorise culture - I think it's important and I've got lots of ideas about it, but I haven't yet met any other people with this same passion.
Page-ranking would remain an issue, likely outside this scope.
I'd like to see some sort of cache-and-forward structure.
And you'd be relying on good-faith actors, which means heavily penalising bad actors.
XML sitemaps are a microcosm of putting the indexing onus on websites instead of the search engines - they are basically ignored by search engines because they have been abused and are not a useful signal. If pages aren't important enough to be linked to throughout your website then they aren't interpreted as being important enough to return to users. The optimistic case is that sitemaps/indices will send parallel signals to the search engines in which case they are redundant. The pessimistic case is that the sitemaps/indices will send signals orthogonal to the content provided to users in which case the website is either being deceitful or incompetent. In any case, the search engine will not want to use the sitemap/index as a signal as it either doesn't provide value or provides negative value.
It would be pretty easy to verify whether or not the index is accurate with a small random sample of pages on the site, and then penalize / exclude (or do a de-novo crawl) for those sites not providing a legit index.
Google has been purging large swaths of data from the indexes and they won't say how or why or exactly what criteria they are using. It's difficult to imagine a worse solution for the web than this current model.
Wow, interesting, this is the first I've heard of this. Might you have some link or citations about this? Thanks.
https://searchengineland.com/4-of-the-google-index-hit-by-de...
The top-level problem is fan-out. If you want to fan the query to the top million domains (far too few to match Google's retrieval depth, but enough to demonstrate the issue), you're going to need to implement some sort of multi-level fanout, since just one machine can't send enough HTTP requests -- nor even establish that many connections -- in a reasonable time-frame. There are going to be severe tail latency issues that will prevent you from gathering documents from potentially relevant sites. You will have to make frustrating tradeoffs about when to time-out per domain queries to provide a good user experience. And many more issues besides. All decisions that are unnecessary if you control the index. Also, internet bandwidth isn't cheap, and you're going to need a lot of it just to consume the top ten results from a million sites.
The next technical issue is that the inverted index is only a small part of what goes into information retrieval. Scoring is at least as important. Modern scorers are multi-level, meaning they do one pass over many documents on a simple, correlated representation of the data. Then they do a second pass on a more comprehensive form (i.e. the whole document, and metadata about the document), but over fewer documents. There can even be third and fourth passes. The data for the first pass is often embedded directly into the index, and it would be challenging to come to any kind of agreement among stakeholders about what data should be embedded. This goes double for the second and subsequent passes. Moreover, those second and subsequent passes often use data about the document, rather than data in the document. Data a site owner would be unable to provide or even incentivized to falsify. Not less than these problems is the issue of where to run the scorer code. If you're running it locally, you're operationally 90% of the way to the complexity of an inverted index. Why not go all the way?
Then of course there are the economic issues, which, roughly, are: "Why should I pay all this money to host an index of my site that nobody uses when Google will do it for free and charge me nothing?"
Can you elaborate on what is the "simple correlated representation of the data"? It sounds like you understand this space pretty well might you have any links or literature on how modern crawling architectures and indexing work? Thanks.
[1]: https://nlp.stanford.edu/IR-book/pdf/07system.pdf
[2]: https://pdfs.semanticscholar.org/2795/d9d165607b5ad6d8b97183...
One of the challenges of creating a "web index" is first creating indexes of each website. "Crawling" to discover every page of a website, as well as all links to external sites, is labour-intensive and relatively inefficient. Part of that is because there is no 100% reliable way to know, before we begin accessing a website, each and every URL for each and every page of the site. There are inconsistent efforts such "site index" pages or the "sitemap" protocol (introduced by Google), but we cannot rely on all websites to create a comprehensive list of pages and to share it.
However, I believe there is a way to generate such a list from something that almost all websites do create: logs.
When Google crawls a website, it is often or maybe even always the case that the site generates logs of every HTTP request that googlebot makes.
If a website were to share publicly, in some standardised format, the portion of their log where googlebot has most recently crawled the site, we might see a URL for each and every page of the site that Google has requested.
Automating this procedure of sharing listings of those googlebot HTTP requests, the public could generate a "site index" directly from the source, via the information on googlebot requests in the logs.
Allowing crawls from a "new" bot would not be necessary.
Webmasters know what URLs they offer to Google. Google knows as well. The public, however, does not.
It is a public web. Absent mistakes by webmasters, any pages that Google is allowed to crawl are intended to be public.
Why should the public not have access to a list of all the pages of websites that Google crawls?
I don't know, but there must be reasons I have failed to consider.
What are the reasons the public not know what pages are publicly available via the web, except as made visible (or invisible) through a middleman like Google?
There are none.
Being able to see logs of all the googlebot requests would be one way to see what Google has in their index without actually accessing Google.
Not everyone will do it and those that do may not do it to 100% completeness: people may not keep their http logs in good order, for example.
Not everyone will provide CCBot with the same access that they provide to Googlebot. The question is how many will?
It is sort of a catch-all issue with anything on the web: "Not everyone will do it." I am not sure that anyone aims for 100% participation where the web is concerned.
There is always an uncertain amount of variation involved with particpation in anything across the entire www.
Publicly maintained directory that I believe was at least theoretically independent of the larger web companies. It certainly had it share of drama, but was a decent human vetted index of what was out there....
Will Google be willing to open its indexes? Probably not at their best interest, because it will help its competitors?
https://news.ycombinator.com/item?id=17548623
If you like, I can try implementing that with my next data analysis project. Right now I'm studying the MySpace Dragon Hoard, and I'll soon write a blog post with maps of music genres around the world.
For simple "I know TF-IDF, let's build a toy search engine" it will suffice, but apart from that?
Some users may be happy to wait for hours or days to get high quality answers not available from commercial companies. Can still be faster than emailing a human friend or consultant or tasking an employee or department.
Central search indexes like Google are not going away. There are client-side metasearch interfaces that combine Google results with other sources. Those other sources can be much slower, including human responses. You would still have your synchronous sub-second response from centralized search, but there would be asynchronous results from decentralized search.
This exists today, e.g. when you post a question on HN or a messaging app, asking other humans for answers not available in public indexes. Most of the world's knowledge is not public, it's obscure and may only be of interest to specific niche audiences.
Where is it found ?
I can't find a reference at the moment, but this topic was covered in a professional journal for historians.
I imagine getting status updates with intermediate search results, and I annotate each one with 'warm' or 'cold' and maybe add some more search terms into the hopper to forcibly narrow or expand the search.
Google is a private for-profit company so we cannot realistically expect them to provide something for free to the public without generating profits in return.
The web index is not a locked up proprietary resource by anyone, so people can do the indexing themselves but the real question is how do you fund a service that will keep increasing its workload exponentially and indefinitely? What institution will have the required resources to bare such costs?
I welcome the idea of data being totally free where you make apps to use mirrors instead of APIs
When there was no tabs users could always see the content of the page knowing what they are looking at. But tabs hide other pages so it is important that we know what is in all those other tabs.
There doesn't seem to be a mention of how to alleviate a tragedy of the commons problem (unless I missed it). If common crawl is doing a fine job, who funds them?
A proposal for building an index of the Web that separates the infrastructure part of the search engine - the index - from the services part that will form the basis for myriad search engines and other services utilizing Web data on top of a public infrastructure open to everyone.
He had no idea what I was talking about.
You could implement your own log monitor or use services like crt.sh or certstream to build a candidate list of domains that have registered SSL certs.
Regardless, the notion of a general web index is well nigh moot at this point due to its not having been built into the system from the get-go. Any such attempt at this point will be, by definition, ad hoc and built by some group of individuals, with the vastness of the content, the cost of the project and the intrinsic conflicts that will no doubt arise making independence from finance and legal issues non-trivial, to say the least.
Really, Wikipedia is the most sensible foundation I can imagine, given that Google has become a self-serving for-profit corporate advertising machine.
Really, I am of the perspective that a machine-grokked indexing system will always be less useful in a significant set of edge-cases than a human-curated index due to such factors as language ambiguity and gaming of such algorithms. As well, the sheer size of the internet requires ranking the pages to ensure the most useful links are properly denoted as such.
WP, being likely the most important and useful crowd-sourced and -built human information system, it is up to us to both keep it funded and add the information we deem important.
And yes I know that's pretty much all of Google. It's just that it's hard to get away from the idea that an index of web pages is anything other than the property of the people who created each web page and the links on it.
And it's not such a big leap to argue that data that is generated by my behaviour is actually my data. (if is likely to be personally identifying data - or perhaps a different term like personally deanonimisable)
I do agree with the general direction of GDPR - but I honestly think the digital trail we leave is a different class of problem that needs different classes of legal concepts to work with.
I think digital data is a form of intellectual property that I create just by moving in the digital realm.
And if you have to pay me to use my data to sell me ads, you will likely stop.
As opposed to Apple Open Directory?
An indexing standard seems a critical element.
You have to remember that while the little that the web provides is also its strongest attraction; it allows the web to be accessed and modified by anyone, they're on little bit of the web can be very different from someone elses.
So by adding on a way that the web must be indexed is kind of like moving closer to communism than liberalism. I guess if we start dictating to google where to get their data then we've moved to the full blown hammer and sickle stage :).