Google doesn't recognise or penalise stolen content
pi-datametrics.com
pi-datametrics.com
What if I published a book, it was copy-pasted in blogs, and then later I put it somewhere crawlable by Google? You certainly can't just say "first time we saw it, that's the proper owner". It would either require a massive amount of manual QA to get right (and even then, there are going to be interminable copyright battles), or have a super high error rate.
I think Google's best value is letting proper content owners easily find violators via normal searches, and let them deal with them via takedown notices or the court system -- which is where it should be done, not in a pseudo-court run by a Google who does not want what responsibility.
I suspect that Google simply doesn't care. They get Ad revenue regardless and in their laissez-faire editorial position it doesn't matter. What are you going to do, use another search engine?
In the mean time it isn't even Google's content so its not a hosting issue, they are just the "neutral" third party providing their 10 blue links (oh and supplying the advertising engine those sites are using)
My guess, having been at both Google and Blekko, is that "whitelisted" search will be the next wave in the industry. For those old enough to remember Yahoo!'s original "directory" model, once Yahoo!'s contract with Microsoft is up one could hope they rebuild their search team and technology into something with a strong editorial bias for "quality" content.
Three things which would be a major issue for Yahoo! on that sort of search would perhaps be first that they themselves rely so heavily on content syndication to power their various verticals, second they keep losing search market share (especially as more search happens on mobile devices and Google has mobile locked down with their Android contracts), and they also screwed up their old directory before they moved it to Yahoo! small business as part of the Alibaba share spinco.
I also don't see how Yahoo would effectively differentiate their search engine enough to be able to (profitably) buy share at prices set by Google, particularly if they over-promote their internal results & rely on a smaller search index.
I expect it does include fees paid to Apple so that Apple would send search traffic to Google, and fees paid to browser vendors.
Our experience as a search results provider was that there was demand for a more 'functional' search capability (not casual searching) many of the techniques we used have been adopted by Microsoft in their Bing engine which has improved both their recall and quality with respect to Google results on highly contested searches.
I certainly agree that Yahoo! has made a number of missteps with their search technology. I talked with them once (post Marissa's arrival) and in many ways they were confused as ever about how search engines generate value for the parent company, but such things are rarely permanent.
One interesting bit from the most recent IAC investor conference call is on it they mentioned that their search deal with Google was renewed for another 4 years & that the rev share on mobile was lower than it was in the past. An analyst asking a question mentioned both Google and Yahoo! were lowering revenue share on mobile.
> ... strong editorial bias ...
Interesting. Care to elaborate?
When you use words like 'whitelisted' and 'editorial', I imagine humans adding something to a database one by one. But the volume of useful pages (and the number of site) is really large now, so I guess that's not what you mean.
One thing I like about search today is that it's almost comprehensive. If I know something exists on the (open) web, I can usually find it with a few searches, even if it's very recent or obscure. I don't want to go back to the days when I browsed gopher directories, or even to the days when finding good quality content meant a hierarchical journey from a directory to a site, to a site map, to an individual page.
Duplicate removal is essential for making a web search engine that works. For instance, together with a CS research group, I built a search engine for a major university library that had more than 80 web sites. We found huge amounts of duplicate content produced by various mechanisms (for instance, multiple people posted the same stuff to the web.) If your ranking is content-based, all of the duplicate documents are going to rank the same and form a "plug" that excludes other documents.
It has long (post 2006) been a common story that "I wrote a blog post but somebody else ranks for it." For instance, I made a blog post that got a huge amount of traffic in the day, but right now you search for it and you find a presentation from some fresher at Oracle that is based on those ideas.
There are many factors that make this hard to control and these include: (1) for one "real" origin there are probably ten or a hundred fakes, so if you are picking at random you strike out -- you have to not only outrank one fake you have to outrank all the fakes, (2) freshness... copies are fresher than the original, also they can be updated years later, (3) also the bad guys think a lot more seriously about indexation, Page Rank, and other variables they control than do most content creators.
The behavior does seem weird in any case, like there is a certain slot for a given piece of content, and Google is swapping different domains in and out to fill that slot. It seems like Google is actually trying to identify the original content, failing, and then actually inadvertently penalizing the original producer.
Also, the combination of the pagerank algorithm and normal user behavior typically helps Google to understand who was first and who deserves to rank higher. That is, most people don't plagiarize content, they quote it and then cite the source, which (thanks to pagerank) tends to rank the original better than sites which have plagiarized it.
Which is how Google makes blogspam such a good business to be in, even if your content is inferior to the post you used for "research".
But Google certainly isn't intending to make blogspam a good business to be in, and I'd argue that they aren't; over the past four years Demand Media's stockprice has fallen from $400/share to $4, and the general marketplace for commoditized SEO services has shrunk by a similar degree over the same period. 19 out of every 20 SEOs who were active five years ago have thrown in the towel... just check alexa graphs for the top SEO forums.
The SERPs are clean these days. Google has done an amazing job every year for at least thirteen years now of improving them constantly. The new wave of spam is social. In practice this means Buzzfeed writers stealing user-produced content from AskReddit threads and it ending up polluting my Facebook feed to the point that I can't even find any good counterfeit Raybans.
Here's an alternate take on that http://www.johnon.com/1075/bullish-on-seo-rankbrain-vs-seobr...
"This is because SEOs follow and influence the intent of searchers in the marketplace, while Google’s algorithm (and AI) merely monetizes it."
Where does the extra monetization on page 1 results come from? Unless he's implying that Google provides bad search results so that people will click the ads instead.....
And then there is the knowledge graph & other flavors of scrape-n-displace, which is largely content recycled from elsewhere, given prominent positioning not based on merit or editorial quality, but based on who the publisher (or recycler) is.
Another parallel trend would be the confirmation bias / brand bias factors promoting older and staler sites. Or simplified "take" articles in the mainstream media rather than the original source articles on niche hobbyist blogs and forums or such.
And in taking broad sets of new niche intents and trying to guide those streams of users back down well worn paths. For example, sometimes when you want to find a particular news story about a broad & well-known web platform like Apple, Amazon, Facebook, or Google it can be hard to find sites other than the official site. And on some other longtail queries Google rewrites what is being searched for in a way that brings up some results that don't match the true searcher intent. Probably the best example I can come up with on this front is say you wanted a pair of shoes of a specific brand, size, width, and model number. If they are not the most recent and most heavily marketed versions it can be tough. Auto-generated internal search pages on trusted brand sites rank well, while a small retailer carrying that specific shoe might be penalized by Panda.
That explains why I've noticed some older sites which are still around, and have plenty of detailed technical information, seem to have disappeared from the search results. Somewhat sad that the "newer is better" mentality appears to have taken over completely... if I really wanted the newest things I'd look at Google News.
[1] https://news.ycombinator.com/item?id=10103545 [2] https://pubsubhubbub.appspot.com/
If Google wants to be the best search engine possible, returning the original result for an article relevant to the user's query is a better result than returning some second-hand copy littered with low-quality ad junk. And if that's not Google job, then let me know whose job it is and I'll start using them instead.
Saying that Google is "selling" stolen content isn't that clear, though. Yes, they're selling ads on search results, but wouldn't they get the same ad revenue regardless of where those links pointed?
It's easier to make the case with AdSense, where Google literally profits directly from stolen content.
Huh? Now I'm worried. Are you telling me that the Rolex watch I paid $30 for, that I bought from a street vendor near Times Square, might be fake? Oh no, the horror! /sarcasm
I don't think that Rolex is too worried about this. Nobody would mistake a $30 watch for a real Rolex. And, give it credit, my fake Rolex worked for a year or so. It probably just needs a new battery.
Besides, you can't sue a street vendor. They're what's known as "judgement proof".[1]
And as to police action against them, the de Blasio administration seems to have adopted a laissez-faire attitude about all this stuff. If they're willing to allow squeegee men to operate with impunity, they certainly won't care about novelty watches being peddled.
I think plagiarized is more accurate.
Plagiarism comes from a word meaning "kidnapping" though, so the tone of both words is pretty similar.
e.g. The Soviet spies stole the plans for the hydrogen bomb.
(Well, unless you're talking about getting the courts to tell the original owner it's yours instead. Which has been at least attempted a few times.)
See definition 1d.
'gunna'. Is it still wrong?
--original comment--
Common usage does eventually lead to validity, though. Language is not static and evolves through usage. Besides, how are you going to measure validity? Is the OED the sole arbiter of what's "correct"? Common usage and mutual understanding is a great way to determine what's "correct" in a language.
I won't debate the legal usage of the word, as there was no mention of legal interpretation by the courts in the previous comments. However you are correct from a legal perspective where words have specific meanings.
For anyone interested in copyright and legal issues, I'd recommend checking out techdirt.com. They have a great starter section at https://www.techdirt.com/blog/?tag=techdirt+feature, and they cover legal, copyright, patent, surveillance and all sorts of related topics. High quality journalism.
You are very wrong. For many years, nearly 100% of Google's revenue was from AdSense.
What if someone spends days writing an article and posts it on his blog. Then, someone else copies and pastes it onto BuzzFeed, which becomes the top search result for that topic.
BuzzFeed is making money that the same blogger would have made from his own content. Now, also assume Google serves ads to BuzzFeed, but it does not serve ads to the blogger. Google has a financial interest in ignoring the provenance of the content in this case.
Is all of that ethically acceptable?
That's incredibly untrue. A substantial portion of Google's revenue has always been and continues to be from first-party AdWords ads.
The fact that you're using BuzzFeed as an example, a firm which emphatically does not use display ads, shows how little you know about this.
Google has always gotten a very large majority of its advertising revenue from ads on its own sites, not on third party sites.
Consider a novelist who works for 10 years on her novel. A hacker steals the document from her computer and publishes it online under his own name. He makes $100M.
Is it wrong for the novelist to feel like someone stole from her? What word would you use instead?
Saying that someone is "stealing" when they infringe copyright is like saying someone is "killing you" when they present convincing arguments against your cause. It isn't literally stealing or killing, it's an exaggeration made for emphasis.
The reason there is so much contention is that a) the same language has been extremely common among hysterical content industry lobbyists who insist that it is literally stealing, and b) stealing and copyright infringement are both unlawful (and therefore more easily confused) even though there remains a meaningful distinction between stealing and copying.
But that distinction is very important in practice because we can't treat stealing and infringement the same. If you don't like someone's speech you can't be allowed to steal any of their webservers but you have to be allowed to copy some of their work in order to effectively criticize them.
Infringing. (duh)
Different circumstances, different terminology. The correct terminology (see US Title 17 or CDPA 1988) is "infringing". Anyone who insists on using the word "stolen" is signalling their ignorance of the first, most basic fact of copyright law.
1.) If the item is being given away for free, there can still be infringement.
2.) If a person would never purchase an item at the available price (due to the law of supply and demand for example), that person might still infringe. No revenue was lost or gained since the transaction would never have completed at the existing price.
In either of those cases, no revenue was "stolen", but infringement still occurred. These are some of the many reasons that stealing isn't a good way to describe copyright infringement.
If I point a gun to your face and take your money, that's a law being broken, but taking that money is not stealing, it's robbery.
If I threaten to expose some dirty secrets and demand money or things from you, that's blackmail but not stealing.
If I take your textual content and re-publish it under my own name, then that's an infringing use but again, not stealing.
Since internet isn't free, you're still paying for content. At what point are you paying "enough" that the information isn't public anymore?
Are you saying no one should monetize their content using ads unless they're willing to allow anyone else to do that as well?
https://googleblog.blogspot.com/2011/02/microsofts-bing-uses...
http://searchengineland.com/google-bing-is-cheating-copying-...