Update those periodically (hours / days / weeks). The adverts don't change particularly quickly.
On the flip side, some people can't change their IP addresses easily, and getting IP banned (even if rare because of the reasons you stated) is actually a major hassle when it actually happens for those people. :/
If only they all did this. So many sites I get to and they're a blank page or an absolute disaster....
However, I don't see cache links on Google :(
Edit: Oops, I'm wrong. The article does say that the Google bot only sees the first paragraph or so.
"The reason: Google search results are based on an algorithm that scans the internet for free content. After the Journal’s free articles went behind a paywall, Google’s bot only saw the first few paragraphs and started ranking them lower, limiting the Journal’s viewership."
So maybe this is why there's no Google cache.
Also, if Google can only index the first few paragraphs, the results are much less comprehensive.
Of course, the problem with nuclear is collateral damage. Drop the bomb and ads don't work, but neither does a lot of other stuff. E.g., the site shows a blank screen, images are invisible or blurry, drop-down menus don't drop. And, of course, the deal-breaker: videos don't play.
The remedy for killing JavaScript is more JavaScript (and CSS). But supplied inside a Chrome extension targeted at the offending site. An injected stylesheet makes `<body>` visible again, hides assorted useless junk, and styles injected UI elements. Your content scripts load the missing images, drop the menus down, and play the unplayable videos in button-activated pop-over windows displayed at superior resolution.
Of course, the problem is, there are a lot of sites out there, and they change unpredictably, requiring your extension library to change in response. That argues for crowd-sourcing the extension library, but the crowd needs to be proficient in HTML, JavaScript, and CSS and know the ins and outs of browser extensions and care and have time.
You can completely change how a site presents. E.g., change a slide-show in a static slide window that barely moves due to the background ad-tech load changes into a set of `divs` that roll upwards as your finger swipes.
It's a hobby at best. Disabling ad-tech components by origin is the practical option.
I used to play around with filtering sites to make them less antisocial, but find that slog less entertaining these days. So now when confronted with a site that's useless without JS, eh, there's almost always another site out there that doesn't mind the terms I demand for my attention.
Sometimes I do a Google Image search because I found something interesting, but don't know what it is, so I'm hoping of going to a page that describes what I was looking at. Pinterest shows up as a result, but with no backstory nor does it lead me back to a source, so it's worthless as a search result. It'd the ExpertsExchange of image searches.
Endy's advice of using `-inurl:pinterest` seems invaluable and I'll be adding that for all image searches in the future.
Pinterest needs more competition.
In the context of Google being the alternative, this is sad and funny at the same time.
Hacker News in the abstract: The web needs more decentralization.
Hacker News IRL: Let Google own every vertical.
That said, there aren't many players in a position to provide actual competition to Pinterest like Google can. The obvious concern is where you draw the line at anti-competitive if Google tried to do the equivalent of what they did with Yelp ratings.
Would also love to have one that rewrites URL's in search results to avoid the frequent 5 second pause when Google's redirector gets it's head stuck up its ass or whatever the problem is.
Note that it doesn't block some spam domains that sneakily use certain special characters i their domain names. Unfortunately Google hasn't fixed this issue for forever.
And for Firefox: https://addons.mozilla.org/en-US/firefox/addon/personal-bloc...
http://jesuschristsiliconvalley-blog.tumblr.com/post/4896203...
I want a world that supports business models other than web advertising. It negatively impacts journalistic integrity and freedom while further exacerbating the race to the bottom search engines and other content aggregators create.
WSJ is responding poorly to a bad situation. I suspect it'll cost them.
http://brave.com/ is one approach.
Eg, have newspapers _ever_ had integrity, and if they _did_, what's different?
They did somewhat before the massive corporate consolidation starting in, IIRC, the late 1970s when newsrooms started getting axes and the major dailies progressively became skins over wire services and lightly-rewritten press releases.
The internet often gets the blame, but it providing actual competition was decades after the terminal quality and subscribership decline of American newspapers began.
It actually was the internet competition (both wire services being directly available to readers and the loss of advertising) that actually got some of them talking about building up newsrooms, rebooting investigative journalism, and relying more on subscription income (paid subscription was never paying the bills before, it was pursued as a key metric advertisers used in determining how much it was worth to advertise in a paper.)
Multiple revenue streams (sales, classifieds) would make them less beholden to advertisers.
Of course they would still have to write material that sold!
I don't like to use John Oliver as a source, but there's some decent content in this:
If Steve Jobs was still alive, I'd bet Apple would be working on a competing search engine with some of these features.
krschultz's comment is really relevant as well.[1] In a complete system, search should actually know about what you subscribe to already, and not penalise those results for you.
I don't want search services to know what I subscribe to. That's private information.
Steve Jobs reinvented PCs, reinvented mobile. I think "the next Steve Jobs" could do the same thing for search. I'm less and less certain of Google's monopoly on that space going forward. It's still built around Web 1.0 tech, has hacks into Web 2.0, but there's a Web 3.0 it's not ready for.
Free zero-revenue startup idea: there'll be an IP-over-ham-radio or something to preserve "internet classic". (Largest use will be bitcoin-for-pornography).
There's got to be a startup idea in there somewhere.
put this in perspective: there are millions of sites out there. theres just 1 major search engine. Yeah I think i know who has leverage here.
It specifically mentions cloaking.
> Webspam pages try to get better placement in Google's search results by using various tricks such as hidden text, doorway pages, cloaking, or sneaky redirects. These techniques attempt to compromise the quality of our results and degrade the search experience for everyone.
Facebook doesn't even need Google, their users just visit the site directly.
I guess Mark zuckerberg doesn't lose sleep over this.
- hidden text,
- doorway pages,
- cloaking, or
- sneaky redirects
They just show a big popover nagging you to log-in. But you can click this away.
If certain Facebook content pages rank low, or do not rank at all, it is because Facebook actively blocks Googlebot from accessing the content, not because Facebook is trying to deceive Google (or the user).
Though Facebook does not need Google, it could get quite a lot more visitors if it lowered the wall of its garden a bit. As is, Facebook is an inaccessible social echo chamber, and I don't lose any sleep over this.
For the record, if anybody needs a draggable WSJ paywall bypass bookmarklet, I put one up here:
WSJ subscribers, online and print combined, are in the low single digit millions. Google search monthly unique users are about three orders of magnitude greater. WSJ online subscribers are close enough to 0% of Google's users as to make no difference.
So, that "some" is essentially all.
If we take just the population of the US in 2017 (326.5M), and assume every American searches via Google for their news, we're looking at give or take 1% of the US with the WSJ subscriber estimate you provided (~3M).
We can refine these numbers further...
19.4% of the US population is less than or equal to 14 years of age (18 would be better, but couldn't find) - so that gives 263M potential American news readers
The Pew Research Center http://www.journalism.org/2016/07/07/pathways-to-news/ states that about 38% of adults in 2016 often get news online (~99M)
That leads to at least 3% of US Google news searchers are potential WSJ subscribers.
So that's not a tiny number (yes the number could be adjusted for worldwide English speaking news googlers - but I think I've made my point).
How many of these subscribers are wealthy and coveted by advertisers?
What if more news publishers follow WSJ and you happen to be a subscriber of that content?
With the amount of information that Google has on its users, I don't see why it can't adjust search results based on whether or not you subscribe - and bring value to whichever side of the paywall you reside on.
There's an interesting difference between overlay paywalls like Wired uses, and content-not-loaded walls like WSJ uses. In the Wired case, they text is sent to you but they try to stop you from reading it. In the WSJ case, they don't even send you the text of the page you supposedly clicked on.
Since we're in the second case, this isn't even a decision by Google. The WSJ actually isn't sending you the data in the search result snippet, so the crawler rightly says "wow, nothing useful here". The complex, ideal solution might let me tell Google "search as though I'm a WSJ member", but short of that they're accurately assessing what content is actually available.
I could imagine Google adding some sort of "content is locked behind a paywall" indicator on search results, but if I'm searching for something on the web, a link to blocked content is not very helpful most of the time.
At least being able to know it exists can help consumers decide if they should pay for an article/subscription.
That's called advertising, but the WSJ will have to buy ads like everyone else.
Not sure if they track this but whatever the Googlebot's view of the site's content, if it's constantly bouncing users back to the search page it should get hit with a hard penalty.
For example, Google scholar will search papers that are behind a paywall. By blocking them, I wouldn't even know what to purchase.
This is google foisting their business model down our throats.