But what you hint at might be more correct these days. They are running a reverse wayback machine in that anything not changed in the last year gets removed. If you click the advanced search its "updated within" and the max timeframe is a year.
In fact it seems the date range example doesn't even work: https://developers.google.com/custom-search/docs/structured_...
If I fiddle with it, it returns a result, but I see an hit from just a few days ago at the top...
Sometimes I wish that were true! Try Googling for, say, PostgreSQL documentation and the top result will often be for a 10-year-old version of the software.
Why is it that Google is thinking the older page is more relevant? Does PageRank outrank content (and Google is oblivious to similar pages that have different versions?)
They did that for a long time, but some years ago the index grew so big, they started restricting it. I thing the general timeframe is 10 years or less till the last update.
> If you click the advanced search its "updated within" and the max timeframe is a year.
Because it makes no sense to go further. For older content you can define individual dateranges. And yes, it works fine for me. Tested a search for 2015 just now, first side had entries all from 2015.
> In fact it seems the date range example doesn't even work: https://developers.google.com/custom-search/docs/structured_....
All those examples are not working. Wasn't custom search retired some years ago?
Google has conditioned us into thinking that "an algorithm that automagically separates the wheat from the chaff" is the only way to do things. It worked for them for a while, but the adversarial forces of marketing, spam, malware, etc are very creative and fast-moving, and that's a lot for an algorithm to try to constantly reckon-with, so best case it'll probably stay a stalemate.
But that's not the only way things can be.
Since Twitter is their preferred platform, go put the activity of journalist Twitter accounts into a relational DB and start searching for who always boosts who. You'll find patterns. Of course there's nothing inherently wrong with this, but at the end of the day I don't need to know what a dozen NY Times journos think of a NY Times oped which is clearly written in bad faith, pushing a false narrative about a particular news event.
Non-gameable ultimately means people who influence the results can have no monetary interest in the results.
example:
Twitter thread...
https://twitter.com/jiatolentino/status/1263208982614814722
...vs reality...
https://wearyourvoicemag.com/jia-tolentino-parents-teachers-...
That search engine is like a gold mine! I searched for "black people love us" in DDG and it was the first result, followed by an article written this year explaining how it came about and .. the web felt like such a smaller place back in 2002 and I just don't remember this at all.
Which makes me wonder about how newspapers and free to air tv kept the culture pretty shallow before exploding with the internet and now, possibly contracting again as our filter bubbles shrink? Just an errant though.
But so much novelty and interesting stuff at wiby.me - search for 'trump' and the first result is just surreal.
Facebook also used have RSS feed for public pages and posts. Now they've not only removed that feature but also have heavy restrictions for third apps.
Services connected to the Fediverse, an alternative framework that focusses on connectivity, is slowing growing. It's only a matter of time before they are more successful than the walled gardens.
https://www.hackernoon.com/what-is-wrong-with-the-internet-a...
I had a search today, and 7 of the top 10 results were from today. What I was looking for was NOT news, it was historical. If I wanted news, I would click the news tab. Having 7/10ths of the results come from today makes using google to search all of the web ever near useless, as todays noise is noisier than ever.
I dont even care if they are defaults, but buttons to "exclude big sites" "exclude the news" or "exclude fresh results" would make search so much better.
What would be really cool was if google could show what the results for a search looked like on a given day. Not the current algorithm, not any sites indexed since then, but what it looked like at the time. Going back to use 2010 Google would be a dream.
The other way would be to have the index and algorithm versioned, where you can target any instance of the algorithm against any version of the data.
I am sure it’s technically possible going forward, but it would be interesting if such capabilities could be enabled for historical versions of the index and algorithm. Combined with anonymized historical zeitgeist data, some interesting digital archaeology could be attempted.
All the more reason to run your own crawler! What’s the state of the art for this area right now in self hosted solutions? Can you version your index and algorithm like we’re discussing and do these kinds of search-data time-traveling?
Though document fingerprinting is hard. Especially w/ fungible page elements.
Internet Archive has an angle here.
Are you referring to WARC type tooling or what? I don’t want to put words in your mouth. I’m a complete learner on this topic. I think gwern has written a bit about this broadly? I’m curious to know more about this, if you have time to share more.
We have them, they suck.
> and focused on indexing the long tail of insightful content that is neglected by Google because it lacks SEO
How would you even define that? SEO is changing all the time. And google is fighting it all the time.
And how would you prevent SEO focusing on that new searchengine? If it becomes big enough, people will optimize for it.
Then we might go back to something sort of interesting.
I wonder whether doing a parallel search on google and filter out by their top-results from your own results would be a feasiable solution? Add a filter on the top 500 websites, and whether known ad-sources are used and you might get slowly there.
Maybe instead of a smart searchengine it would be better to focus on a dumb focus which gives access to all the metadata of a page too, and allows people to optimize for themself. Fulltext alone is not the only relevant content for good results. Google knows it and uses them, but has very limited acces to it for the enduser.
Still, I guess that’s only viable because Google rewards lots of links. If you just disable link relevance that part of gaming the system will be gone too.
Then, the more distinct queries a given website ranks for (i.e. the more SEO battles it wins/the more generally optimal it is at “playing the game”), the less prominently any individual results from said website would be ranked for any given query.
So big sites that people link to for thousands of different reasons (Wikipedia, say) wouldn’t disappear from the results entirely; but they would rank below some person’s hand-written HTML website they made 100% just to answer your question, which only gets linked to on click-paths originating on sites that contain your exact search terms.
This would incentivize creating pages that are actually about one particular thing; while actively punishing not just SEO lead-gen bullshit; not just keyword-stuffed landing pages we see in most modern corporate sites; but also content centralization in general (i.e. content platforms like Reddit, Github, Wikipedia, etc.) while leaving unaffected actual hosting by these platforms, of the kind that puts individual sites on their own domains (e.g. Github Pages, WordPress.com, Tumblr, specialty Wikis, etc.)
———
A fun way to think of this is that it’s similar to using a ladder ranking system (usually used for competitive games) to solve the stable-marriage problem on a dating site.
In such a system, you have two considerations:
• you want people to find someone who’s highly compatible with them, i.e. someone who ranks for their query
• you want to optimize for relationship length; and therefore, you want to lower the ranking of matches that, while theoretically compatible, would result in high relationship stress/tension.
Satisfying just the first constraint is pretty simple (and gets you a regular dating site.) To satisfy the second constraint, though, you need some way of computing relationship stress.
One large (and more importantly, “amenable to analysis”) source of relationship stress, comes from matches between highly-sought-after and not-highly-sought-after people, i.e. matches where one partner is “out of the league of” the other partner.
So, going with just that source for now (as fixing just that source of stress would go a long way to making a better dating site), to compute it, you would need some way to 1. globally rank users, and then 2. measure the “distance” between two users in this ranking.
The naive way of globally ranking users is with arbitrary heuristics. (OKCupid actually does this in a weak sense, sharding its users between two buckets/leagues: “very attractive” and “everyone else.”)
But the optimal way of globally ranking users, specifically in the context of a matching problem, is (AFAICT) with IDF(PageRank): a user’s “global rank” can just be the percentage of compatibility-queries that highly rank the given user. This is, strictly speaking, a measure of the user’s “optionality” in the dating pool: the number of potential suitors looking at them, that they can therefore choose between.
If you put the user on a global ladder by this “optionality” ranking; and normalize the returned compatibility-query result ranking by the resulting users’ rankings on this global “optionality” ladder; then you’re basically returning a result set (partially) optimized for stability-of-relationship: compatibility over delta-optionality.
———
All this leads back to a clean metaphor: highly-SEOed websites—or just large knots of Internet centralization—are like famous attractive people. “Everyone” wants to get with them; but that means that they’re much less likely to meet your individual needs, if you were to end up interacting with them. Ideally, you want a page that’s “just for you.” A page with low optionality, that can’t help but serve your particular needs.
Maybe you could try and make a model SEO article for your own search engine, just taking your existing SEO results and figuring out which parameters are contributing the most to their ranking, then filtering out results that contain these parameters that worked well in your model. Rinse and repeat as SEO writers try and step up their arms race, but they should always end up being foiled by your changes to the search engine after optimizing your own perfect SEO model regularly.
Instead every algorithmic content delivery platform, from Google to Twitter to Youtube to Facebook, is constantly chasing just to keep from being underwater against the spammers.
Less than ads and subscriiptions, but enough to fund a lot of crap.
Maybe, maybe not. How useful is Google (or search in general) to you?
For me search is more of a convenience tool than for finding sites that have information. There are questions I need answered but without search would have an easy time figuring out (e.g. "how many cups are in a pint?"). Sometimes I want opinions, but I almost always am going to the same sources. Sometimes I use search because I'm too lazy to click around using a site's own search. The only things that are actually useful for me from search are for specific expert knowledge that I want in a structured manner (e.g. "what do I need to consider when buying a house?"), and those queries are incredibly few
I feel like search is slowly becoming irrelevant
I think I use google in the same way as you.
Most of the time, I could go to the websites directly (MDN, Stackoverflow, HN), but sometimes I'm trying to find something I don't know about, by trying different terms. I usually do this when I want a particular product but I don't know what it's called or what it is "smallest itx case without gpu", "midi router no power supply", "waterproof tarp diy tent setup".
I installed a browser extension called uBlacklist as recommended by someone here a couple of weeks ago, so now 90% of the time Google search is like my Ctrl-P for MDN, Stackoverflow, etc since I've managed to filter out the sites I don't want to see results from.
2. An algorithm can always be gamed.
You're stuck.
The content on the sites I visit is created by humans. Until automation genuinely overtakes us, I'm not ready to accept at face value the scale of the internet has grown so large that humans couldn't tackle the problem.
All I can say is, good luck with your human curation startup.
If there were orders magnitude more pages than humans, I'd agree. But I'd also ask: Who created them all?
It's not easy to quantify the amount of useful content on the internet. The 2bn figure above seems to stem from registered domains, and depending who you ask [1][2] around three quarters of them are "inactive" (e.g. landing page for a parked domain).
At the other end of the spectrum, Google's index surpassed 130 trillion pages four years ago [3]. Point in favour of my opponent!
If everyone connected to the internet indexed one page a day over the course of their lifetime we might just about do it. And anyone creating a new page would need to [arrange to] index it themselves.
[1] https://www.internetlivestats.com/total-number-of-websites/
[2] https://hostingtribunal.com/blog/how-many-websites/#:~:text=....
[3] https://searchengineland.com/googles-search-indexes-hits-130...
Also, any setup that allows everybody to be a moderator will be promptly gamed.
What you want is what yahoo used to do - a hand-curated search engine. It worked when the internet was small, but got buried under the eventual avalanche of web sites.
- Using DNS "zone files", the DNS database for TLDs (which are not available for all, but most) show there's circa 200 million domains registered at any moment
- A large percentage of these are parked, i.e. no unique content.
- Many domains are "tasted", i.e. bought, are alive for a few days then disappear, so potentially you waste time crawling them
- Lots of sites are database driven and can result in millions of pages that can be created in a day
- URL rewriting means you can have an almost infinite number of pages on any one site
- Soft 404s and duplicate content can be hard to spot and can waste resources in gathering/removing them
There's paid for resources like Majestic/Ahrefs/Moz that crawl the web to see who's linking to who and they all contain trillions of URLs.
I think the most detrimental fact is that pages often disappear or change, I don't have a recent number but I'm fairly certain there's a 10-15% chance that any link you see this year, will be gone next year. "Link rot". Hard to build a DMOZ style directory on that scale with that problem.
I don't think it is unmanageable, it just needs to be seen from different perspectives and managed by different groups of people.
You only need to choose an algorithm so that the amount of trash is less than you can handle.
Those article directories were eventually murdered by Google within the blink of an eye maybe a decade ago, and quite frankly on any given topic nowadays it's way easier to find good content rather than SEO filler. Google's algorithms will nowaydays favor fresh (aka published or updated recently), long content over short, "popular" (aka lots of links on it), duplicated content.
Instead, most stacks are designed to enforce a walled garden, from which very little is shareable unless you go through the gateway (the web client app) to some approved destination.
Semantic web and total availability at the personal-computer level, are aspects of OS design I wish had been paid more attention.
Basically what we have now is a very, very expensive system of dumb terminals.
As an example go look at the tags on soundcloud. People tag their songs with whatever they think will get people to take a listen.
There are plenty of trusted sources that would not spam their pages in this way. And if you spam too much, you risk getting dropped from search indexes, directories, links etc. because your source is just not useful.
It's not as if the semweb people weren't aware of the problem, and it's not self-evidently a fatal blow to the idea.
There are so many possible viable methods for ranking search results! Particularly now with higher level textual analysis using AI/ML/[buzzword], and perhaps more importantly, the resurgence of interest in curated content. People are getting better at discerning curated-for-revenue vs. curated-with-love.
Would you speak to why you think this way about PageRank? What are its shortcomings?
To me, who only paid surface-level attention to this, it seemed like Google results were best when PageRank was the dominant metric. As they moved more and more in the direction of prioritising news, commerce and the aspects we call SEO, “number of links pointing TO the resource” became less and less important in the ranking. And as that happened, the quality of results dropped, and the content silo-ing rose up.
PageRank was peer-review. SEO is “who shouts the loudest”.
As far as I can tell, the main reason Google succeeded was that other search engines let advertisers buy placement for keywords (and didn’t label paid links). I heard from an industry insider that was able to strip the paid links that the engine they worked on gave results that were very similar to Google’s.
The second big reason was that pagerank was a useful signal that hadn’t already been gamed to the point of uselessness. I think this let a tiny team blindside an entrenched industry.
That’s not to say there’s no technical insight behind the page rank algorithm, but it was only a useful signal for a few years.
It got to the point Altavista became more or less useless, and when Google showed up on the market they quickly took it over.
Seems the time is ripe for a new revolution. Doesn't have to be a better search engine, could be something completely different.
Simply having half a dozen or more search engines per country/language, with their own indexes and algorithms should help see the web more fully.
ATM in English, Bing and Google have the largest indexes and Mojeek has its own index though smaller. DDG, Ecosia and others are just the Bing index re-ordered.
I enjoyed the OP, though. And I think niche directories/blogrolls would be progress. The current centralised web is a result of everyone dancing to the tune/rules of the large platforms.
The problem with decentralization is that it creates a power vaccuum that is filled by the most interested actors. Even bitcoin, with it's decentralization-by-design is actually centralized to a handful of miners in China.'
If the goal is to rebuild something that is anything else other than profit based, you need to make sure the organization running it is strictly non-profit.
Also I like the idea of making search like Wikipedia where people can edit results. Obviously you’d need super genius level safeguards to protect against scammers but Wikipedia does it ok-ish.