Google won’t comment on a leak of its search algorithm documentation
theverge.com
theverge.com
https://sparktoro.com/blog/an-anonymous-source-shared-thousa...
It may have been accidental, sure[2]. But is there a basis to claim unauthorized action?
[1]: https://news.ycombinator.com/item?id=40505708
[2]: Reminds me of “Night of the Living Dead”, which accidentally became public domain; source: https://en.m.wikipedia.org/wiki/Night_of_the_Living_Dead
A leak like this (if it’s actually as substantial as people say it is) combined with google’s ham fisted way to insert LLM results into its search could really mean that google search quality will crater in the coming months, potentially opening up space for competitors or reducing usage in a real way.
The major implications are that a lot of what Google was telling SEO practitioners were a lie or cleverly disguised red herrings wrapped up in semantics.
The article that was posted earlier today went through a lot of the stuff they were lying about with examples of Matt Cutts and Gary Ilyes statements that are now directly contradicted with the recent release of this code.
Here is the article: https://ipullrank.com/google-algo-leak
The premise:
What I’ll do here is contextualize some of the most interesting ranking systems and features (at least, those I was able to find in the first few hours of reviewing this massive leak) based on my extensive research and things that Google has told/lied to us about over the years.
“Lied” is harsh, but it’s the only accurate word to use here. While I don’t necessarily fault Google’s public representatives for protecting their proprietary information, I do take issue with their efforts to actively discredit people in the marketing, tech, and journalism worlds who have presented reproducible discoveries.
Thank you for that link, I'll have to check it out!
They are somewhat circumspect about the fact that each example is actually, 8+ years ago Google made a claim about what they did at the time and now there's evidence that they are currently doing doesn't match that claim. That's not a lie.
For example, they link to this tweet from 2016 as such a claim: https://ipullrank.com/wp-content/uploads/2024/05/image21.png
Or a later example where they link to this 2016 talk: https://www.youtube.com/watch?v=iJPu4vHETXw
Later on, they use claims from 2012: https://www.seroundtable.com/google-chrome-search-usage-1561...
Are you asking whether people want discoverability for their product/companies ?
Or if they're willing to give up agency and leave it to a third party for-profit corporation to dictate what is surfaced in people's generic searches ?
We'd have to decide what is quality and what is content first, and I'd expect the heat death of the universe before we ever come to a useful definition that even matches half people's own definitions.
Back in the day (early to mid aughts) as an SEO person you really had to work to get your site noticed. You have to develop inbound links, create really good content, set your site up well for Google to index, used other resources like blogs and social media (when it was still in the early stages) to boost your positioning. In short, it was a full time, month-to-month job. The sites I worked on it was a constant battle to get a decent ranking and to maintain it.
But you are 100% right. Somewhere in the last 8-10 years, everything has gone from search engines really protecting who they put on the front page, to loading the page with tons of ads (half of the results above the fold are now all ads, followed by more ads below the 10th position) and making it insanely easy to manipulate Google and others to get your site on the front page.
Great example: In 2019, I had a medical device company that was a startup. Out of the blue they called me because of my prior relationship with one of the founders parents and the site I built for her. They needed a site designed, built, optimized and 15 pages of content written in less than three months. I gave them some insane price and they didn't even balked and said if I got done sooner, they would send me a bonus.
I grabbed a template somewhere online, revamped the home page and the internal content pages and then proceeded to copy/paste content from other well known and not so well known sites in their industry. I did everything you should never do from an SEO standpoint. I figured it was a long shot, but worth the payoff. I got their site released on time, but the majority of the content was copied from other sites. Very little of it was original.
I figured it would get buried in the first month. Nope. #3 on the first page for various searches I targeted. It was outranking the sites I copied the content from, it was crazy. To this day, the site remains in the top three places on the first page of Google and other search engines for dozens of searches I targeted. It was outranking huge medical supply companies in the same industry - all as a startup.
That experience in 2019 was a huge wakeup call for me. It was blatantly obvious how easy it was to manipulate Google and other search engines. As of today, I have no idea if SEO is really a worthwhile pursuit any more considering how easy it is to do this. If I can do it, then I just assume everybody else is already doing it. If not, then they're missing out.
Smoking gun of what? The article never seems to actually say.
The only accusation I see is that SEO grifters were lied to by Google Search docs, but you look at the article like[1] it's over things like "Google always said Domain Authority isn't a thing, but we found a variable site_authority. We don't know what it's for or what it does, but clearly we were right". Which, even if true, isn't a thing I care about at all...
For a while, I would create pages like “Bank of XYZ Phone Numbers” or “Phone Numbers for $ThatTelecomYouHate” and often rank #1.
Which made sense, because my page was just a very parseable page of their actual phone numbers, but obviously risk for harm there.
Phone numbers on the actual orgs’ websites were hard to find because, well, they wanted you to do anything but call them because calling cost them money. And the big corp websites were always an SEO mess that I’m not surprised a crawler could not comprehend.
Of course the ads were all for competitors, so I enjoyed making money and costing them customers at the same time.
After some years, Google would start returning you 10 different results from the orgs’ websites, playing into their hand of having you do anything but call your bank/telco.
You're right that the real impact is limited though, people woun't be moving to another search engine anyway.
https://github.com/googleapis/elixir-google-api/commit/078b4...
Sure, they can't stop people who received a copy from studying it and using or discussing the information in there. But I think they can almost certainly stop them from (legally) distributing that information to others.
https://arstechnica.com/tech-policy/2021/04/how-the-supreme-...
The README also says:
> Disclaimer
> This is not an officially supported Google product.
which, while it doesn't directly indicate that Elixir isn't used internally to Google, would be a surprising mismatch of support if they did use Elixir. Here are their official clients (including some in languages that I doubt they want to be used internally for service development, like PHP, Node.js, and .NET): https://developers.google.com/api-client-library
Google lied? Of course they did. And Microsoft, consistently (with their pool of NOBUS).
In articles/stories like that my mind wanders to the ST land and the Rules of Acquisition (https://projectsanctuary.com/the_complete_ferengi_rules_of_a...)
125. A lie isn't a lie until someone else knows the truth