Google Has Officially Killed Cache Links
gizmodo.com
gizmodo.com
These days Google will offer a result where the little blurb (which is actually a truncated mini cache) shows a part of the information I'm looking for, but the page itself does not. Removing cache means the data is just within reach, but you can't get to it anymore.
Internet Archive is a hassle and not as reliable as an actual copy of the page where the blurb was extracted from.
I'm not sure how that serves Google's interests, except perhaps that it keeps them out of legal hot water vs the news industry?
Shoot, even a meta description qualifies for that - thankfully Google uses them less and less.
Isn't this the web we want? One where big corporations don't steal our websites? Right?
I'm tired of saying this, yelling at clouds
It is definitely an area where there is no single "people of HN" - opinion varies widely and is often more nuanced than a binary for/against⁰ matter. From an end user PoV it is (was) a very useful feature, one that kept me using Google by default¹, and I think that many like me used it as a backup when content was down at source.
The key problem with giving access to cached copies like this is when it effectively becomes the default view, holding users in the search providers garden instead of the content provider being particularly acknowledged never mind visited and the search service making money from that through adverts and related stalking.
I have sympathy for honest sites when their content is used this way, though those that give search engines full text but paywall most of it when I look who complain about the search engine showing fuller text, can do one. Also those who do the "turn off your stalker blocker or well not show you anything" thing.
----
[0] Or ternary for/indifferent/against one.
[1] I'm now finally moving away, currently experimenting with Kagi, as a number of little things that kept me there are no longer true and more and more irritations² keep appearing.
[2] Like most of the first screen full of a result being adverts and an AI summary that I don't want, just give me the relevant links please…
It makes some sense, too because the edges are blurry. If a user from France receives a french version on the same URL where a US-user would receive an english version, is that already different content? What if (as it usually happens), one language gets prioritized and the other only receives updates once in a while?
And while Google recommends to treat them like you'd treat any other user when it comes to e.g. geo-targeting, in reality that's not possible if you do anything that requires compliance and isn't available in California. They do Smartphone and Desktop-crawling, but they don't do any state- or even country-level crawling. Which is understandable as well, few sites really need to or want to do that, and it would require _a lot_ more crawling (e.g. in the US you'd need to hit each URL once per state), and there's no protocol to indicate it (and there probably won't be one because it's too rare).
The recommended (or rather correct) way to do this is to have multiple language-scoped URLs, be it a path fragment or entirely different (sub)domains. Then you cross-link each other with <link> tags with rel="alternate" and hreflang (for SEO purposes) and give the user some affordance to switch between them (only if they want to do so).
https://developers.google.com/search/docs/specialty/internat...
Public URLs should never show different content depending on anything else than the URL and current server state. If you really need to do this, 302 Redirect into a different URL.
But really, don't do that.
If the URL is language-qualified but it doesn't match whatever language/region you guessed for the user (which might very well be wrong and/or conflicting, e.g. my language and IP's country don't match, people travel, etc.) just let the user know they can switch URLs manually if they want to do so.
You're just going to annoy me if you redirect me away to a language I don't want just because you tried being too smart.
> This header serves as a hint when the server cannot determine the target content language otherwise (for example, use a specific URL that depends on an explicit user decision). The server should never override an explicit user language choice. The content of Accept-Language is often out of a user's control (when traveling, for instance). A user may also want to visit a page in a language different from the user interface language.
So basically: don't try to be too smart. I'm more often than not bitten by this as someone whose browser is configured in English but often would like to visit their native language. My government's websites do this and it's infuriating, often showing me broken English webpages.
The only acceptable use would be if you have a canonical language-less URL that you might want to redirect to the language-scoped URL (e.g. visiting www.example.com and redirecting to example.com/en or example.com/fr) while still allowing the user to manually choose what language to land in.
If I arrive through Google with English search terms, believe it or not, I don't want to visit your French page unless I explicitly choose to do so. Same when I send some English webpage to my French colleague. This often happens with documentation sites and it's terrible UX.
To answer your comment: yes you should return the same content (resource) from that URL (note the R in URL). If you want/can, you can attend to the Accept header to return it in other representation, but the content should be the same.
So /posts should return the same list of posts whether in HTML, JSON or XML representation.
But in practice content negotiation isn't used that often and people just scope APIs in their own subpath (e.g. /posts and /api/posts) since it doesn't matter that much for SEO (since Google mostly cares about crawling HTML, JSON is not going to be counted as duplicate content).
IOW pragmatism.
https://developers.google.com/search/docs/specialty/internat...
> If your site has locale-adaptive pages (that is, your site returns different content based on the perceived country or preferred language of the visitor), Google might not crawl, index, or rank all your content for different locales. This is because the default IP addresses of the Googlebot crawler appear to be based in the USA. In addition, the crawler sends HTTP requests without setting Accept-Language in the request header.
> Important: We recommend using separate locale URL configurations and annotating them with rel="alternate" hreflang annotations.
As a real-world example: you're providing some service that is regulated differently in multiple US-states. Set up /ca/, /ny/ etc and let them be indexed and you'll have plenty of duplicate content and all sorts of trouble that comes with it. Instead you'll geofence like everyone else (including Google's SERPs) and a single URL now has content that depends on the perceived IP location because both SEO and legal will be happy with that solution, and neither will be entirely happy with the state-based urls.
So what do you propose that such a site shows on their root URL? It's possible to pick a default language (eg. English), but that's not a very good experience when the browser has already told you that they prefer a different language, right? It's possible to show a language picker, but that's not a very good experience for all users, then, as their browser has already told you which language they prefer.
E.g. images deleted from Reddit can sometimes be seen in Google image results, but when following the links they might lead to deleted posts.
The simplest answer could be that making the cache accessible costs them money and now they're tightening their purse strings. But maybe it's something else...
For sites that manipulate search rankings by showing a non-paywalled article to Google's search bot, while serving paywalled articles to regular users, the cache acts as a paywall bypass. Perhaps Google was taking heat for this, and/or they're pre-emptively reducing their legal liabilities?
Now IA gets to take that heat instead...
The cache is arguably a strategic resource for google now.
The paradox of the internet archive's existence is if all that data were easily searchable and integrated (i.e. if people really used it) they would not exist by way of no more money for bandwidth and by way of lawsuit hell. So they exist to share archived data, but if they share archived data they would not exist.
and so it is a wonderful magical resource, absolutely, but your "power user" level as well as "free time level" has to be such that you build your own internet archive search engine and google cache plugin alternative... and not share it with anyone for the above existential reasons
This has always been one of Google's best features. Really sad they killed it.
Coral Cache maybe? The caches you listed were my manual order to check when a link was Slashdotted.
Google's cache, at least in the early days, was super useful in the cache link for a search result highlighted your search terms. It was often more helpful to hit the cached link than the actual link since 1) it was more likely to be available and 2) had your search terms readily apparent.
It is not as complete as Google's, but it is usually good enough.
For quite a time they stopped being a simple obvious link but where available in a drop-list of options for results for which a cashed copy was available.
For most users the internet has 5, maybe 10 web sites. I can use Wikipedia search or LLMs when I have questions.
Maybe just use Wikipedia search only then!
I would welcome some rules and regulations about this kinds of stuff. Can you imagine that google wakes up one day and decides to kill gmail? It would cause so many problems. It’s already impossible that even as a paid gmail user you can’t get proper support in case something goes awry. Sure you can argue they can decide what ever they want with their business. But if you have this many users, I do think at some point that comes with some obligations of support, quality and continued service.
https://blog.archive.org/2024/09/11/new-feature-alert-access...
Looks like cached pages just got more useful, not less.
For now, you can still view Google’s cache by typing “cache:” before the URL, but that’s on its way out too.Man, wish the Internet Archive hadn't staked it all tilting at copyright windmills...
The cache is often still accessible through a "cache:url" search. There's been no official announcement, but it does seem like that could go away at some point too. That is even more likely now that Google has partnered with the Internet Archive.
What I'd really like to see, and maybe one good possible outcome of the mostly bogus antitrust suits is to have a continuously updated, independent, crawl resource like Common Crawl.
Lots of discussion then:
https://news.ycombinator.com/item?id=39198329
More recently:
New Feature Alert: Access Archived Webpages Directly Through Google Search
So searching them in Google was exactly how students found the answers, I assume, but we wouldn't have had the smoking gun without a cached, paywall-bypass, dated copy. $Employer was definitely unwilling to subscribe to services like that!
(However, the #1 most popular cheat site, by far, was GitHub itself. No paywalls there!)
Too lazy to find a link, but this is now public and live, although pretty well hidden. Three dots menu for a search result -> More about this page.
"The Wayback Machine has not archived that URL."
A large part of the usefulness of the cache links came from the inherent freshness and completeness of the Google indexing.
Was invisible on the search UI for some time now, but the service itself is still accessible.
The reality is that people who create these filter policies often do so with very little thought, and sans complaints, they don’t know what their impact is.
God help you if you need something that's not tcp on port 443. Yes, I'm still a little bit bitter, but I have spent a lot of time explaining the difference between TCP and UDP to IT guys who have little interest in actually understanding it, and ultimately won't understand it and will just deny the request. Sometimes after conferring with another IT person who informs them that UDP is insecure and/or unsafe, just like anything, not on Port 443.
https://google.com/search?q=cache%3Ahttps%3A%2F%2Fnews.ycomb...
And some people considered that a "victory" for IA.
They'll just foot the bill while Google reap the rewards
VERY rare these days a google search result actually contains what was searched for - anything with a page number in the url and cache was guaranteed to be the only way to access it.
Combine that with the already absolute epic collapse of their search result quality and ms copilot locally caching everything people do on windows, and this may well be recorded in history as the peak of google before its decline.
very sad day.
But here is an instance where all the opprobrium is justified:
> So, it was decided to retire it.
“It was decided”? Not you decided or Google decided, but it was decided? Come on.
You wouldn't want to be a thief... right?