Google's PageRank patent has expired (2019)
patents.google.com
patents.google.com
Google still uses it for the base ranking. But, the results are then run through a variety of add-on ML pipelines, like Vince (authority/brand power), Panda (inbound link quality), Penguin (content quality), and many others that target other attributes (page layout, ad placement, etc). Then there's also more granular weightings for things like "power within a niche", where a new page might do well for plumbing (because of other existing pages on the site), but wouldn't automatically have any authority for medical topics.
Alas, google would shoot themselves in the feet by promoting pages that consume less in their own advertising products.
If the US government split those two business units into different companies, we could have decent searches (google search could still profit by placing ads on its page results)
I've found just using quotes, OR/AND, and some other basic stuff can often get me what I need.
The place where Google does well is that they're a monopoly, so things like image search that require a lot of resources, they are better at, especially paired with their geographic location. (USA! USA! USA!)
Edit: I hit enter and the post went out while drafting.
They do guarantee it. Sometimes Google is buggy and messes up in determining what's page is actually visible on the page compared to just being somewhere in the HTML.
Here's an example someone posted a few months back where google decided the user didn't really want what they said they wanted:
It's made it harder to search for bits of poetry, quotations, or song lyrics, especially.
On Google, I treat my search terms like a Venn diagram if I want good results.
Eg: want an article about the texas blackouts, but not the ones a decade ago?
Type "texas blackout npr 2022" minus quotes.
But if you do that on other engines, it may be MUCH more literal, the literal intersection of those terms, and I need to do the opposite: use as few terms as possible, possibly paired with using the site: operator, intitle operator, or other things.
>It's made it harder to search for bits of poetry, quotations, or song lyrics, especially.
Yeah to be completely clear, my default is DuckDuckGo, then very rarely I fall back to Google, but often if I'm doing that it's because I didn't want to trouble a librarian -- they talk about privacy, but I had a series of unfortunate events when I told one I want to use books as much as possible because I absolutely don't want some of these tech bros to know what I'm looking up.
(That dichotomy of folks who know information science and those who have critical thinking or coding skills needs to end, now. I'm an alumni of one of the highest ranked schools of information science in the world, and I will not be figuratively or literally extorted into a PhD to get roles others get with a bachelors.)
I did learn recently about Kagi on HN and that is marginally better than bing but still doesn't beat Google imo.
I just hope there is more and better competition like this for Google soon. That's the only thing that can put google back in its place and put an end to their evil things.
In my country we have a website that has complete monopoly over sales of used items, services etc. I'm always amazed that Google is able to determine and put it in top 3 results for a wide majority of searches. "laptop mouse" brings up that site, but "kitchen sink" doesn’t, presumably because people don’t buy used sinks.
All the same, the link graph has been truly bastardised because of it being such a prominent facet of ranking for an audience.
So, while the problem changed I am not sure if it got harder or Google’s priorities shifted.
All of Google's fancy ML could've been replaced with a simple report button - enough people report a site and either trigger a manual review (best option) or just ban them (could be used against competitors... but negative SEO was already a thing and is still widely used).
There was loads of SEO back in 2000 even! It brought alta vista to its knees, the number one search engine of the day.
Google got started, grew, because it filtered all that SEO spammy junk.
Do you have specific stats to back this up? Number of SEO pages vs good ones?
Or are you just presuming?
Perhaps in other words, market forces (I'm thinking ROI) and economies of scale will win.
They would experiment in a pretty deep way, varying things like the rate of new links, type of new links, variations of anchor text / bare links, and so on.
It USED to be very effective.
The larger entities don't really have to be that detailed. If you have that brand power, you can just cut partnership/cross-link deals and pay a little attention to things like anchor text in links, contextual text around the link, etc.
Edit: Ah, yeah, agreed. What got lost in all this was good content that had no big brand behind it. The indiscriminate hammer Google used to kill off small-guy SEO spam also pushed a lot of actually good stuff, stuff that never did any SEO at all, off the first page.
What I was meaning was, the players who can manipulate the link graph most cost effectively tend to rank better and it favours those with deeper pockets. Not so bad for competitive niches and high volume terms, but muddies the waters for many other things.
For local search, I almost always use google maps instead. Maps search isn't perfect, but it beats web search hands down.
Some tricks I've used in the past include local church bulletins (most parishes have some sort of a website and a weekly bulletin with advertising on the back), or local sport team sponsors, or local bars.
For towing, you can also try calling a local dealership or auto repair.
This is certainly a problem with the locksmith companies, but I think there's also a Maps problem, too: Google enables this kind of commerce, especially businesses that apparently have a physical location but explicitly stated on the phone that I can't stop by there. It reminded me of all of the thousands of Delaware corps housed in a single building.
If Google made it so that you're only listed on the map if you actually have a physical commerce location, that would help a lot -- at least for those kinds of businesses that need it. Towing companies and the like may be an exception?
> The goals of the advertising business model do not always correspond to providing quality search to users. - Backrub paper
The type of people get into locksmithing are like those who do computer security for fun - they like to figure out how things work and how they can be broken. Which means that every locksmith wants to figure out how to game the system. And the first thing that they all figured out is that people tend to select whatever locksmith is closest. So they all went and pretended to be in a million places around the neighborhood in hopes that THEY would get selected.
That was the case a decade ago. And it was a nightmare. Glancing at a locksmith search now, Google somehow cleaned that up a lot. But it doesn't surprise me that whoever is on top now is someone who figured out how to game the current system.
I had the same thing happen just last week! The locksmith who came for me gave me his personal phone number though and said if I needed help again, I could call him directly and he could give a better rate due to not needing to pay the centralized company a portion. Despite often having social anxiety with strangers, I've still found that just being friendly and treating people well (in this case still giving a tip alongside the cancellation fee after my super finally arrived just moments before the locksmith) ends up getting me better services than any research I've been able to do online. I ended up making a similar connection with the people we hired to install our curtain rods just after we moved in; I was talking to them about how the movers that brought my girlfriends things over to the apartment ended up damaging her desk in transit, and they said that they also do moving as well as installation/assembly, so next time we end up moving, we also will be able to hire people we know and trust (we ended up hiring them again for a few odd jobs since the first time, and we almost always end up getting one of the same three people we've met and made connections with).
"But wait!", you say, "there are some legit reviewers out there." Yes, there sure are, by my starement is accurate, because for every legit review site with aws referrals, there are tens of thousands of ml created spam referral sites.
And so the real review sites are often lost in the mix regardless, which makss arguments to keep those results pointless.
But google leaves them there, and this is the same sort of site which, if it were an email, would immediately end up in a spam folder.
And beyond this, the other part of the problem is their ridiculous aliasing of search terms, which helps spammy sites come back as a response.
You say google lost? It's not losing, if you just don't care.
Frankly, it's just a return on investment thing. As long as only a few people per tens of thousands bolt, why spend the r&d?
This is the root of the problem. If Google targets the low-hanging fruit, the spammers will very quickly find a workaround. Google is trying - and failing - to use more sophisticated signals.
Spam is a virus. Google beneffited from the network effect, eradicating all opposition. But, like a European medieval monarch, it now has a poor immune system because of its lack of genetic diversity. The many, many SEO spammers are constantly experimenting, and they only have to find one flaw in Google's algorithm for it to win and spread quickly.
Google's monoculture cannot save us, however hard they try.
But yeah, a healthy set of search engines will get us better results.
[1]https://fossbytes.com/google-downranks-65000-torrent-sites-i...
[2]https://www.projectveritas.com/news/google-machine-learning-...
The most evil thing Google is doing now is pushing brands to buy ads on their own name so that they appear above the organic #1 search result. It's like the time Facebook decided it wouldn't send messages to followers who like you unless you paid up.
It's not like this is a new thing. The court cases about it are almost 20 years old: https://en.m.wikipedia.org/wiki/Google_v._Louis_Vuitton
You think the results were better in 2006, but the corpus was also different.
That's like an old person complaining that their body felt a lot better back in 2006 when, I assume, they didn't have to use their walker and glasses all the time...
How would you demonstrate that search is objectively worse? And how would you then show that it's a result of Google's algorithms, and not a consequence of the content of the Internet changing significantly?
How would you demonstrate that being 50 years old is worse than being 25 years old?
You ask people that are 50 years old or older because they have been on both sides.
> it's a result of Google's algorithms
Well, it's simply in front of your eyes: this [1] was not possible in 2006.
Anyway, the fact that you cannot easily find on Google why Google search results are worse, proves that Google search results are worse today than in the past.
https://www.theverge.com/tldr/2020/1/23/21078343/google-ad-d...
There's a few ways to do that. The easiest is to point out that almost any lucrative search query has -0- organic results above the page fold on a typical monitor today. It's all ads unless you scroll down.
Then, it's not proof, but how much time do you think Google spends on things that sit below the fold and aren't clicked on much? What would the financial incentive be?
Quality organic search results is the reason they can have ads above the fold. There's tremendous financial incentive for them to care about that.
My bet would be that each algorithm would perform best on the web of its day.
An old person's glasses and walker don't make their body feel worse. They're responses to an underlying change, and in fact make them feel better than they would without them.
Similarly, I'd argue that the ML pipelines and complexities in Google search aren't why search results are worse today. Rather, the web has changed with more SEO spam, walled gardens, content in videos, and search has changed in that you try to find more kinds of information than ever before. It's the underlying changes that make the search seem worse, and all of Google's fancy algorithms are imperfect responses to that. Without them, I'd be surprised if Google's results weren't far worse than in 2006.
The comment I responded to:
> My search results were a lot better in 2006 when, I assume, they didn't have all these ML pipelines...
made it seem like maybe the ML pipelines were somehow causing the decline in quality, rather than simply an imperfect response to changes in the web since then.
Vince (authority/brand power), Panda (inbound link quality), Penguin (content quality)...
This represents a pretty severe narrowing of results on information and opinions - possibly the worst results are on Youtube searches for newsworthy events. It'd be very interesting to see what kind of content a pure PageRank algorithm-based search engine would generate today, and I'd be very interested in using such a search engine. Now, would it be overrun by SEO? I don't know, but it'd be worth finding out.
I kind of wonder if Google Scholar is purely PageRank or citation-count based, it still gives very useful results with relatively simple query strings.
even without Google account?
This is my Google experience from my temporary workplace in Barcelona
https://i.imgur.com/X2Avw2v.png
Then I click on "I agree" and whatever thing I search I receive Spanish results because I am in Spain.
Not my idea of "the best of any search engine", but ok...
Youtube search is awful. I once went through news queries on Google searches of Youtube (site:youtube.com) looking for popular independent media outlets (BreakingPoints for example) by -MSNBC, -FOX, -CNBC etc., and ended up with a string about 25 queries long, and those shows are just banned. They were feeding me unpopular corporate media shows with hardly any views or subscribers at that point, but zero access to independent media. Full-on propaganda manipulation.
It's simply a fact that in 2006 Google search results were better.
Reasons might vary and could well not be Google's fault, but it's not old man yelling at clouds it's provably true.
Of course people that were not using Google before 2006 can't really know, just like people that are not old cannot experience how much better being young is, body functioning wise.
Obsessive programmers like myself really want a search engine that does better and helps us find the weird corners of the web.
Most people do not care. They can type a question into Google and get an answer, and that's all they're looking for.
Relative to what? Google has better performance than any other search engine. It's relatively poor compared to the imaginary ideal search engine that gives you exactly the result you want for any query regardless of whether the information even exists or not.
Google really does seem to be frustratingly bad these days.
I think Google has made one change for the worse, though, which is strongly favoring more recent content. Increasingly, I think that change has been a big contributor to the decay of the web since.
It seems that Google has decided that most people want the most updated information when they look for something, which I don't think is entirely unreasonable.
What I would love, however, is a way to turn that off for particular searches. Researching past events, as a trivial example, benefits far more from exact results rather than most recent tangentially related blogspam.
Since everyone I know in tech laments Google's decline into uselessness, I'm assuming this is not a sustainable strategy.
Everything I've heard from people I know at Google suggests otherwise. Most searches for most people ... work. I too struggle to have google work in specific research cases, and I would like more power-user toggles, but basic searches like "$celeberty_name photos" or "$my_kids_school calendar" or "pizza places near me" just sorta work.
But you can! In search results, click Tools and switch the Any time dropdown to Custom range... and you can specify a date range in the past. (Apparently, the custom option is hidden in the mobile version?!) I'm not sure how precise and dependable it is but it seems at least partially useful when I search for historical events.
Google is trying, they're just losing the arms race of detecting crap content vs generating it.
All Google needs is an obvious, one-click "spam" button for logged-in users. Clicking should add a site to the user's spam blacklist (which they should be able to review).
They know about what users search for and how those searches overlap and separate users into groups (not to mention all the individual details they have access to). When sites are marked as spam by enough different types of users, those sites can then be manually reviewed and their content blacklisted by Google preventing the same or very similar sites with a different URL from appearing.
Unfortunately, Google makes a lot of money from these ad/spam sites, so they have a perverse incentive to keep allowing them.
I think that Google just don't care and are happy to be a search motor for Reddit and Wikipedia and to answer questions like "How old is Lady Gaga?".
I genuinely don't understand.
So the algorithm is still useful not only for search results.
The thesis itself can be found here: https://git.vbrandl.net/vbrandl/masterthesis/raw/branch/mast...
Some groundwork I built upon: "SensorBuster: On Identifying Sensor Nodes in P2P Botnets" https://git.vbrandl.net/vbrandl/masterthesis/raw/branch/mast...
Google still uses PageRank, but at the risk of stating the obvious, the current PageRank is much more sophisticated than the one found in the bibliography.
That completely depends on how you model the queries. It can be a TB sized relation all the way into needing more bits than there are atoms on the universe.
Prior to Bayh-Dole legislation in the 1980s, this was the case for university-held patents: they could be not be exclusively licensed. Repealing that legislation would be a good idea to avoid the rise of monopolisitic behemoths like Google/Alphabet.
Citation needed
A successful Google competitor just needs to be way better than what’s currently available. Nobody seems to have the next big idea for a better search engine yet.
Hand-curated walled garden sub-web that bans SEO spam with an iron hand.
If monetisable, it would turn gaming wikipedia into a whole new level of shitshow of course.
I would speculate that Apple could roll their own search engine using Siri.
We'll see what happens in either WWDC or in a few years, but it seems to be quite early for that and certainly have a long way to go.
"it is executed at query time, not at indexing time, with the associated hit on performance that accompanies query-time processing" - for a search engine that planned to take over the Web that might have been a dealbreaker.
There's also the fact that PageRank was presented as a query independent computation that could be done ahead of time and HITS as a query dependent computation. A resourceful enough person could however modify HITS into topic-wise precomputed HITS scores to be combined at run time based on the query.
Ok so here is how one can game HITS. Create a harvester page that points to lots and lots of popular, high traffic pages on the internet. By virtue of doing this it can accumulate a lot of Hubs score which it can redirect as an Authority score to an intended page.
> Create a harvester page that points to lots and lots of popular, high traffic pages on the internet. By virtue of doing this it can accumulate a lot of Hubs score which it can redirect as an Authority score to an intended page.
I'm not sure how HITS is any more "easy to game" than PageRank? As far as I understand it, the differences are almost entirely limited to performance characteristics, not semantics. The example you give doesn't seem to be specific to HITS (as opposed to PageRank) in any way.
(I'm also not sure how "game theory" is relevant here, unless by "game theory" you just mean "the idea that people will try to game it".)
If you check my original comment, I gave a simple scheme to attack HITS rank. The main drawback is that one can 'harvest' Authority score using 'out-links'. Outlinks are cheap and easy, compared to 'inlinks'. Sybil attack is a little harder for Pagerank.
OK, but how is it harder for PageRank? I can't really see any differences in the semantics of the two algorithms, so I'm not sure what kind of added vulnerability one or the other could have.
> One could pose this as an adversarial game.
Yeah, I appreciate that, that's what I was referring to as "the idea that people will try to game it". It's not really the kind of 'game' that would be considered in game theory, though, because it doesn't have any interesting or emergent properties - the designer's response will just be "oh yeah we should stop people gaming our algorithm".
If you are familiar with the algorithms, which I assume you are, you can work it out.
To make my page score high on the PageRank score I need to acquire links from high PageRank score pages. This is a lot harder because it depends on a) in-links and b) high PageRank pages. With Hits, its easy for one page to harvest a high Hub score. All that is needed is to outlink to known good pages (authority). Providing outlinks is trivial. Once so harvested, one can direct that flow to a designated page to give it a high Authority score.
> It's not really the kind of 'game' that would be considered in game theory
Why not ? Formalize the strategy spaces of both the players and its a very valid game in the Game Theory sense. For the ranker you have to consider some functional space of functions over a graph. For the page player it has a budget of alterations it can make to the graph.
Are you saying that you think HITS doesn't recursively score the quality of references by their own scores? That's not true. It does exactly what PageRank does in that respect: a page's score depends on the score of those which reference it, which in turn depends on... etc.
The 'hub' vs 'authority' distinction is interesting but not really relevant here: we're considering a page's 'authority' score, which depends on the 'hub' score of those who outlink to it, and at that point we're just doing PageRank [again, except performance-wise and arguably freshness-wise].
Like I said: the only non-trivial differences between them are implementation / performance-related, not semantic.
> Why not ? Formalize the strategy spaces of both the players and its a very valid game in the Game Theory sense. For the ranker you have to consider some functional space of functions over a graph.
Yes, again: possible to frame it as a formally valid problem if you really want to; still not an interesting one. We're only talking about this because you want to maintain that your earlier statement was true.
"You have to consider some functional space of functions over a graph" gives no detail (besides that, yes, you can model something–maybe documents, maybe people, who knows?–as a graph) and sounds like something written by a person with a gun to their head.
Or maybe I'm wrong and there's a fascinating problem which you just don't want to divulge to me.
I doubt that reading comprehension is that hard a skill to master. I dont see where I have said anything about recursion or their lack of. If you want to have an imaginary conversation between yourself and what you think I have said, you can continue. I do not need to participate in that. I am sure you alone will suffice.
I think going back to the Pagerank and HITS papers carefully and understanding them will be illuminating. You keep saying they have no semantic difference, which cant be further from the truth. The scores are the eigenvectors of very different matrices, and HITS scores are straight forward to manipulate. BTW the papers cite each other stating in what way the other is different, if their difference was mere implementation detail and no semantic difference, I doubt they would stand as published papers.
The Hub score is not something interesting that I happened to say but a core part of the HITS paper. It is by virtue of the Hubs score that the Authority scores are defined (and vice versa) in the paper. It so happens, its easy to bump up the Hub score of a page by adding few strategic outlinks.
> Yes, again: possible to frame it as a formally valid problem if you really want to; still not an interesting one. We're only talking about this because you want to maintain that your earlier statement was true.
That's your opinion. I am just countering your categorical claim that there is no game theory formulation possible here. If your yardstick for your assertion that no game theoretic formulation is possible is your inability to find one that's interesting to you, there is not much I can do about it. All I can say is that such an yardstick is not very popular or useful.
A game theoretic formulation with a budget constraint adversary is a very natural setting. Research on link spam resistant node ranking algorithms are a thing, as is evaluating how stable are the rankings produced by some of these proposed algorithms to (potentially motivated and adversarial) changes to the links in the graph. SIGIR, WEBKDD proceedings on link analysis and rankings would be a good place to look.
You will need some background in matrix perturbation analysis, especially perturbation analysis of principal eigenvector to understand some of the results. Perturbation analysis of finite Markov chains will also suffice.
https://ai.stanford.edu/~ang/papers/ijcai01-linkanalysis.pdf
is by far one of the easiest papers to read in this area (Andrew Ng has focused on other areas of research since this paper). Note that the stability bounds in that paper can be easily tightened... left as an exercise for you. As you will see in the paper, HITS scores are easier to alter (equivalently stated, they are unstable) compared to Pagerank scores.
> Or maybe I'm wrong and there's a fascinating problem which you just don't want to divulge to me.
This is hardly the forum for extended discussions on a research topic. With the pointers and sketches that I mentioned a competent grad student would be able to fill in/ develop it further.
That's largely what I've been doing with my search engine.
I also think it's worth exploring a simple reputation system. Why not have reputable users evaluate results? (approve/disapprove buttons) I think almost any reputation-free system will eventually fail to bots or cost a huge amount to win the tug of war against SEO. Reputation mostly solves the issue.
Speaking of reputation, I think the great insight would be to apply a ranking algorithms to reputations themselves -- if you can't trust your users, you fall back on the same problem.
To rank reputations, clearly PageRank doesn't work because it values all users equally, which is unfortunately not sybil-resistant. I think one approach is to have an "invite system", where your reputation is associated to who invited you, with earlier users having greater weight somehow (also the administrator can manually assign trustworthy users).
This also suggests a way to formulate distributed trust. You can join a "trust network" by trusting a certain user(s), and then you import users who they trust as well. (I believe this is the rough idea behind Web of Trust, although I believe WoT is not algorithmic -- it should have been!) The problem with this approach for a practical search engine is that you can't aggregate results (you would need to store each user's vote and compute a personal ranking every time) -- so I think in practice a useful compromise is to give the user a choice of a few "trust sets". You trust <Public User A> here? Join this trust set. You trust <Public User B>? Join this other trust set. (Combining a small number of trust sets should be trivial)
As for an algorithm, something like:
-- By trusting other users, up to 100*(1-sqrt(N)/N)% [note] of your trust points will be redistributed (diminishing the impact of your choices).
[note] Obs: Formula arbitrary, chosen to approach 100%
-- The total trust conserves, and phenomena like cyclic trust are not a problem due to conservation.
The signals that matter the most: 1. Anchor text (and all variants of smearing and distinguishing between high quality and low quality anchors) 2. In aggregate, which pages got clicked on on any given query (and all smearing variants -- using ngrams, embeddings, ...). At Neeva, we use it for retrieval and scoring. 3. Query understanding signals mined from the query-click bi-partite graph and the query-query session refinement graph. 4. Page summarization signals built on top of 1 and 2 and body text. 5. To a lesser extent, query-independent page quality signals
Whether you use term-based retrieval or nearest neighbor (embedding) retrieval, a heuristic combination of signals or LambdaMart, whether you calibrate to human eval or clicks, whether your topical relevance function is hand crafted or uses a combination of deep learning and IR signals are all details past that.
tldr; there's a lot of craft in a search ranker, and no one silver bullet. Definitely not just PageRank.
How is this a representative example?
You didn't specificy that the breakthrough needed to be useful to the users or the people wanting to monetize search.
Now that the patent's expired, somebody can create a free open source implementation called PageStink.
Eric Schmidt: We don't want them to catch up. We want to stay ahead. We call for all sorts of techniques to try to make sure that we rebuild a domestic semiconductor and semiconductor manufacturing facility within the United States. This is important, by the way, for our commercial industry as well as for national security for obvious reasons. By the way, chips, I'm not just referring to CPU chips, there's a whole new generation, I'll give you an example, of sensor chips that sense things. It's really important that those be built in America.
https://www.hoover.org/research/pacific-century-eric-schmidt...
What are the chips that "sense things" and that Schmidt wants so much to prevent from being available to China and any other country?
Short answer: electromagnetic sensors used to spy on everybody so Google can show them "relevant ads" and profit. You will find them embedded on your CPU, on the nearest cellphone tower's transmitter and on Starlink satellites.
Slightly longer answer: since the 1980s, Silicon Valley has used semiconductor radars to collect data about what you think (your inner speech) by means of machine learning with data extract from wireless imaging of your face and body. It has proved very convenient for them, as this enables blackmail, extortion, theft, sabotage and murder like nothing else. They can do this because they design the semiconductor used on your phone, computer, TV, car and for your telecom supplier's network equipment, which makes possible to embed silicon trojans everywhere.
Don't underestimate what machine learning can do. e.g. Study shows AI can identify self-reported race from medical images that contain no indications of race detectable by human experts.
https://news.mit.edu/2022/artificial-intelligence-predicts-p...
Also, don't underestimate the number of people Silicon Valley is willing to kill to maintain a monopoly, as you may be the next victim.
Coming from SEO world this looks like a joke.