Yandex ‘leak’ reveals search ranking factors
searchengineland.com
searchengineland.com
108 MusicQ: The musicality of the request. The results of the sorcerer Anton Konygin.
276 NightQuery: The request is set mainly at night (and accompanying ones for morning, afternoon etc)
316 IsUa: Domain in the .ua zone (ukraine)
747 WordHostWikiSum: The relative popularity of the Word -Host pair, where Word is the word from the Title article on Wikipedia, and the Host is the host that is referred to in this article.
755 NastyContent: Content ugliness factor. (several of the other factors related to port and "adult" content are defined as multiples of this)
892 ShortVideo: A document is a short video (Tiktok, Reels, Shorts)
894 TelegramPost: Document - post in telegram (note I could not find similar for twitter etc, but they do have for tiktok)
1076 More90SecVisitsShare: The share of visits for which the time spent during the day on the host is more than 90 seconds (they track how long users are spending on pages)
1078 RankHackedNovaPhp: Rank of hacked sites
1309 DistanceToAnkara: The distance from the city from where the request was set to Ankara
1603 YellowImgCount: Average yellow images count on host
1614 NewsAgencyRating: Rating of news agency from agencies.json (Yandex.News resource)
I don't know if either of these are relevant here, but it's possible a weight on Yellow Images could be related to NSFW images.
So my suspicion above is probably wrong, and "yellow" images probably refer to a Toloka score/color for whatever domain the images are hosted on.
[0] https://yandex-explorer.herokuapp.com/search?q=yellow&o=all
So say suppose you're in New Orleans and I am in San Jose, now we can compare our distance, because geographical distance usually equates to cultural differences and perceptions of quality.
For a decade and a half, Yandex has had a codename for each big revision of search ranking algorithms and release of new search features. Most were city names, and Magadan was one of them. So “distance” is probably a distance to some rating introduced in that release. Whether it is a positive or negative factor (cutting obvious over-optimized sites), I have no idea.
In other words, clickbait images from clickbait ad networks.
Was the agencies.json published?
I agree though, would be very interesting indeed to see who they promote and/or push down
Sounds fascinating. Anyone have any idea what this is about? Is [0] the Konygin in question?
- These are factors, not the underlying algorithms which use them, which means it's not possible to deduce your own ranking from that data.
- Many of those are so vague as to be useless.
- The most important factors for SEO are those we already know about.
For example, we obviously want the most-viewed and most-bought products to rise to the top. However, we also expect there's an initial honeymoon period for many new products, where people want to see them but they don't have enough history yet to sway the popularity factor. So there's a non-linear term that looks kind of like
( weight_factor / (product_age_days * age_scale_constant) )I remember reading a few years ago that most search engines use some tree based model. If that's the case, that means the idea of monotonic linear weights is not relevant.
So you can run your bulk-recomputing whenever you have spare capacity.
Stale rankings aren't great, but don't hurt that much. As long as your liveness is more frequently updated, so you don't send people to dead sites.
By this point, I can't imagine they haven't automatically balanced {revenueFromQuery} to {costOfQuery}.
No sense delivering hyperoptimized results if you lose money on generating them.
- The X matrix (e.g. the page rank score) for each result.
- The y vector, i.e the score for each result. Although we can observe the relative ranking in each sample (would be interested to hear about techniques to cope with this).
- nobody in the West cares about Yandex SERPs and Yandex isn't Google, so you can't easily transfer the factors
- Your favorite SERPs aren't being overtaken by SEO garbage because they're not big/profitable enough to catch attention. It's really more of a dark forest thing in SEO, and search volume is how you detect prey. SEOs don't care about things they cannot make a lot of money from.
- SEOs aren't that technical. They employ some technical people (me for instance), but they don't listen to technical people, they care about links links links links, headlines, title, description, microdata (for fancier Google SERP display because having a fancy display can make rank 2 work just as well as rank 1) and keywords. Some do keyword-related stuff like WDF*IDF, but it's more of a ritual where they throw text into different tools and wait for all lamps to turn green. They're really not that sophisticated. Source: am working for large affiliates with millions of visits per month.
Everyone in SEO treats Google like a god. If you have a somewhat stable + successful project, you're sacrificing things to Google (let's all adopt AMP, yes, it's so great! let's all do CWV and say we believe in user experience!) and praise Google each day and avoid anything that you've heard might displease Google for fear of Google sending down lightning and turning your projects to ashes. If you don't have one of those golden geese, you must be more nimble and make sure that the lightning only strikes where you've been yesterday and doesn't catch up with you.
The blackhat part is much more stressful (I've had clients get super depressed when their old tricks no longer worked and they burned site after site after site and nothing worked for months), and everyone I know that did that has transitioned to whitehat as soon as an opportunity presented itself and gotten rid of all the blackhat stuff to not have their main sites get caught in a penalty.
I have no idea how ML generated content will change all that, but Expertise, Authority and Trust (=EAT) is what everybody has been worried about for the past year, and whether Google will believe them that they're experts and should be trusted.
what are some suggestions you make that your clients don't follow?
They vastly over-estimate Google's abilities to identify patterns ("we use similar wording on other sites, that's a pattern" for wording that everyone uses, like talking about the beach distance on a hotel review, or just using bootstrap) and are like "I don't understand, but I trust you so I'll consider this okay ... for now". If some ranking drops, they're always eager to roll back any changes that were done, even if they didn't affect the frontend ("who knows how Google works").
At the same time, while they want to avoid patterns to connect their sites, they run all their sites in the same search console and analytics account.
Most of them pick up very few technical things and transition from Product Management into SEO by learning how to run reports and what blogs to read for instructions. Introducing the DOM inspector is something I'm often blowing minds with, same with viewing request in the network console ("you mean I don't need the redirect tester 3000 any more?")
I've only once worked for clients that had someone doing SEO who was deep into technical SEO and had all kinds of automation set up and roughly understood how HTTP works, the differences between server and client etc.
Much of it is quasi-religious, I've often jokingly suggested trying to sacrifice a chicken. I'm not sure everybody laughed and never gave it some serious thoughts.
I get involved when the SEOs have decided what to do, but I don't keep up with trends and developments.
Of course, the "how exactly is it implemented", "what weight does this factor have", "can I exploit how it works" etc is another story, but if you started building a search engine and wrote down your ideas about ranking and approaches to deal with some manipulation attempt etc, I'm sure you'd come up with a lot of the same things.
Google got a lot of flak over the fake news stuff so they've gotten even more conservative with what they rank in recent years. I think that's a big reason people have noticed results getting worse, there is a lot of really great niche stuff on small blogs but Google would rather rank more generic content from an established website. A few major sites like CNET have already gotten caught churning out AI produced stuff and Google still ranks it due to site authority
Here's the shortest explanation of why you'll never get anything better.
Apparently. Had heard of SEO, but not SERP
The latter is interesting as a factor because it suggests that visitors are visiting for the value rather than because they've been targeted by ads, implying higher value.
Not surprised that some websites have an artificial preference though. It feels like this is basically the selling point (other than privacy) of Kagi, DDG with its bang commands, and other alternative search engines – artificial inflation of rank.
Some pages, that I won't name, but one of them starts with a pin- and ends with -terest consistently get high google results while offering garbage content.
I don't think that if a user stays a short time necessarily means the content is garbage. It might also mean that the user found directly what he was looking for, such as the answer to the question.
It might even mean the page is very good as it is clear and provides the answer and does not have other click-bait type content.
Landing on a bad page, might require you to read and browse around and maybe after a while figure out it's garbage or you get distracted by some "interesting", but unrelated articles. But if the search engine would rank it high because of this then that would be wrong.
But looking at my kids, looking directly at the info box and considering the answer as given by God, I suppose I am an exception.
But Google definitely tracks such short clicks and considers them a bad signal, very important in ranking. Same for Yandex, search for "dwell time" (visit duration) in its code/factors.
On the other hand, maybe some people actually like the pinterest pages in results, and i'm the weird one... who knows.
It isn't just you! I suspect that for certain people it's easy for them to get distracted there though.
I've seen some data on this, and that data was inconclusive. Sometimes it seems to work, sometimes it seems to be counter-productive, sometimes nothing happens. It's hard to truly test things in isolation, and maybe their methods were shit, but I don't think it really does anything, or maybe Google has solid ways to detect anomalies and ignore them.
In fact, a good way to downrank a competitors pages is to do a search for them, click the result, then click back 5s later. Repeat 1000x from different google accounts. A week later, the competitor will go off the first page.
> and doing it with an acceptable latency
Lots of interesting optimizations possible here, but the big obvious one is multiple level models: score documents with a cheap model (FastRank in yandex lingvo) first using a subset of the fastest available features, then rescore top docs with your best slow expensive model. Perhaps rescore multiple times at different points in the stack with models of varying complexity, at each index shard and after aggregating the results from subset/all shards. Also sort documents in each index shard by some other ML model with query-independent features to push all the junk to the end of the index where you'd likely skip it when running out of time budget to process a query.
> Also, what happened to Google page rank, is is still relevant today?
Vanilla 1990s' pagerank obviously not, but the idea of such graph-based calculations is still very useful yes.
what did we learn about the flaws?
PageRank is factor 0.
Without going into specifics, I've seen ranking treated as a multi-stage algorithmic problem. Initially, you rank results with a lower-quality but low-latency ranker at the first stage (inverted index, tf-idf, knn, etc), and subsequent stages rerank the top-K results with higher-latency ML models outputting a relevance score.
I believe progress is currently being made to combine everything into one giant neural model that just ranks everything from the get-go rather than pass in multiple stages.
Maintenance sure is hell, lots of sweat and tears. Just glancing over search/formula/webcommon/select_ranking_models.cpp makes me cringe, they must have many dozens if not hundreds of different models in prod by now. Each of them needing maintenance and lots of training data. Work on new ranking factors I suspect must be also highly frustrating: throwing stuff at wall^W catboost black box and seeing if it sticks, and if it doesn't you'd have little idea why and control over it. Imho google's approach (white-boxish interpretable top level ranking formulas) is far superior and maintainable at scale.
> It is equal to one if the site has a Ukrainian geoist (i.e. 1 - Ukrainian site)
Also keep in mind Yandex was different company a decade ago. Back then Ilya Segalovich been CTO there (he died of cancer in 2013) and he supported opposition and even participated in street protests in Moscow.