More content by people, for people in Search
blog.google
blog.google
In fact it seems like Google worked better when wikipedia was almost always the first result on various topics
What I mean by rambling is if you type something like "types of oranges" and suddenly land on a page where some dude is clearly just filling up paragraphs for google like "so you want to learn about oranges? an orange is a great type of fruit. here's a bunch of text about oranges being described in historic literature"
Anyway, from what I've heard, that's the reason recipe sites have started posting 4 pages of drivel about the dish before getting to the actual recipe.
Or so I've heard, I don't really have any good sources for this, could just be hearsay
Or, Google Analytics. Or Google ads. Or, if you return to Google and try other links for the same query. Or if you return to Google and refine the query (e.g., a new search where the term is within a predefined threshold of word vector similarity)
There are plenty of ways to approximate time spend on link.
E.g. if a link to your page is shown to the user in search and they don't click or they return to search page too quickly, Google sees this as a signal that the result was not helpful.
If it is a hidden algorithmically social function then of course it will be gamed.
But IRL there are certain people whose advice and recommendations you value and those whose you ignore. And in other cases you can easily ask the source of other information to find out if it is high value or not. MLM is the gamification of the IRL social structure and it's fairly easy to opt-out.
That's what search needs, a way to see the path information took to be presented to you and a way to filter it.
Unfortunately right now, so much of the best information is in Facebook groups, post and comments. The interface there is absolutely horrible though and not designed to provide you information, but to maximize the amount of ads that come across your screen.
The same is true for video information. It's not easily searchable or digestible. The web peaked when information was predominantly text form and not fragmented into walled gardens.
Appending site:news.ycombinator.com instead of Reddit?
Can you create a fake account with a verified credit card number, verified phone number, passport, drivers license, account history consistent with human usage, Google One subscription, etc.? You probably can, but doing it at-scale is going to be quite costly.
Quite cheaply
Google is already scraping bits of website content and showing it to the user as a "featured snippet". Nobody is going to write short pages if Google can already rob you of a click so easily.
The update literally says they are working to get rid of the type of useless content you are talking about.
Enormous swathes of the platform are filled with gossip channels (politics, crypto, stonks, celebrities, music) hosted by clout chasers desperate to be seen as authorities on that topic.
Of course there can be good content inside of these categories, but you can generally tell which channels are "optimized for engagement", i.e. run like a business, and those are more amateurish and perhaps more authentic.
reality 1: The www exists. Google indexes it, analyzes it and delivers it to users. Users like certain things, like original content.
reality 2: Google's ranking policies/algorithms influence the web. The "original content" that exists in a world without Google is different to the content that exists in a world where Google ranks such web pages more highly.
Google refuse to see or present themselves in the role that they actually occupy. They avoid thinking of the search algorithm as encouraging or discouraging anything. It's just analyzing.
On youtube, This mindset is even more loopy, because on youtube they actually own the platform. The recommendation engine or whatnot implements what is clearly a new policy, Youtube pretends that there was no policy to change in the first place.
Can you clarify what you mean by this?
As an outside observer, it seems that Google recognizes that their search engine does have a "Heisenberg" effect akin to "measurements of certain systems cannot be made without affecting the system"
E.g. the Google search ranking causes the rise of content farms. Google then fights back with a new revision to the algorithm to downrank them. That constant arms race between various blackhat SEO and Google's new algorithms was commented on many times by Matt Cutts (Google former head of search quality): https://hn.algolia.com/?q=matt+cutts
Another example is the RapGenius punishment by Google to discourage content that tries to game the algorithm: https://www.google.com/search?q=rapgenius+penalty+google+ran...
Google often makes manual human intervention e.g. M Cutts team telling HN they're looking into RapGenius SEO hack: https://news.ycombinator.com/item?id=6956658
Are those are not examples of Google understanding its effect on internet content that tries to game their algorithm?
I think the issue is that Google's algorithm isn't perfect -- and therefore it appears like they don't discourage bad content.
In a sense, I am overstating. Having an anti-spam team is obviously a recognition that SEO/Spam exists and that Google is trying to discourage or mitigate.
But Google is well beyond just attracting spam. Success in Google search rankings is, for many sites, the better part of online success. If Google ranks needlessly wordy recipes more highly than concise & useful recipes... then wordy recipes get written. It's not about ranking recipes anymore... it's about the effects google has on recipes. The recipes get written in a wordy fashion, because of the ranking system... not just ranked because of the writing style. Google is dictating the nature of online recipes with their rankings.
A few years back, Google made changes to youtube's recommendation engine in a way that really "encourages" frequent, regular videos. Viral hits became much less common. The result has been to send many professional youtubers into a frantic grind. There's no point in taking time with a video, or trying weird ideas. It won't go viral anyway. A week off or a couple of flops is harshly punished. There are tens of thousands of these youtubers, basically small businesses.
There's never any reckoning with 2nd order effects. It's always communicated as it is here. I believe this is how Google execs actually think about the issues. As a quasi-spam problem to be mitigated with the next “helpful content update.”
Can you articulate what concrete actions Google could take that would address these 2nd order effects?
>, Google made changes to youtube's recommendation engine in a way that really "encourages" frequent, regular videos. [...] A week off or a couple of flops is harshly punished.
I've seen this repeated many times (especially from Youtubers making videos about "burnout") but the analysis about cause & effect seems incomplete. As one counterpoint, Ben Krasnow "Applied Science" channel slowed down from weekly uploads to a video every few months and yet his views and subscribers went up not down: https://www.youtube.com/c/AppliedScience/videos
Another yt channel that had a year between uploads and the views went up : https://www.youtube.com/channel/UCX7katl3DVmch4D7LSvqbVQ/vid...
My pet theory on the contradictory anecdotes: the Youtubers making videos on a topic with lots of competition from other Youtubers uploading every week -- such as fast fashion clothes shopping -- are the ones that seem to suffer if they slow down. E.g.: https://www.youtube.com/results?search_query=zara+haul
However, if you're making videos in niche topics with originality (maybe "weird ideas" as you put it), you won't be penalized by infrequent uploads. Many examples including Technology Connections, Applied Science, etc
So a different conclusion can be reached... if one makes a high quality videos, one can even upload just once a year and the Youtube algorithm won't penalize you.
If one youtuber produced a 32-minute video per month and got 500,000 views while another produced eight 4-minute videos per month and averaged 200,000 views per video, who do you suppose gets the most ad money? I'd wager the latter.
Black box or not, google are creating these systems, analyzing, optimising and implementing them. It's indeed hard or impossible to trace back individual examples and their whys. But the macro effects are visible and choices are made.
It's just easier to "blame the computer," internally or externally. It's just like bank employees merge always blame "regulations." In fact, whatever they are blaming is a policy created to implement that bank's compliance framework which exists to satisfy the regulator's goals and the bank's goals." Most of the time, the underlying regulation is distantly removed from whatever piece of bureaucracy is annoying the customer or employee. But, if it can be blamed on the regulator, it will be.
It's just easier, if at all possible, to blame an immovable and mysterious force rather than the more likely culprit: humans doing a bad job.
> Are you writing to a particular word count because you've heard or read that Google has a preferred word count? (No, we don't).
https://developers.google.com/search/blog/2022/08/helpful-co...
a) dont summarize other sites
b) dont write overly long articles just to make google happy
It's a shame a) is never helping with b)
Happiness meta-rule: Throw electronics away, use analog equivalents instead.
In an ideal world, such an endeavor would be supported by governments all around the world (or a single "Earth government"), with yearly updates to keep up with the state of the art.
I've started just closing the tab. I don't give a shit that a certain country planted oranges and it gave them naval superiority, I want to know what types of oranges there are!
Google prolly cannot earn money simply by featuring high quality stuff like Wikipedia or nerd websites. Usually SEO farms spend money on Google to earn money. So that's why search really cannot improve in my opinion due to this conflict of interest.
Back when "SEO" was new, I would read the Matt Cutts blog. He was head of "anti-spam." I remember thinking back then that anti-spam was an ignorant frame.
Once Google gained importance, websites started trying to improve their rankings. That might mean migrating from Flash to HTML. It might mean meta-tags, content, keyword stuffing, link collecting... paying a consultant.
Google's early view/advice seemed to be: "Just ignore rankings. You do your thing, makes your site as useful you can. We'll do our thing: judging your site algorithmically and deciding how to rank it." Websites intentionally trying to improve rankings was ipso facto spam. The very idea of SEO was spam.
Naturally, cracks appeared. Flash and embedded images was an early one. Google's initial position was: "HTML is better for users." They thought websites should use clean HTML regardless of rankings. No contradictions need be confronted
Once a two sided conversation starts, that crack becomes a wedge. It becomes clear that website owners don't care about the usability of plain html text. Usability doesn't matter until you have users... and users come from Google. The "language" of SEO continued developing in this disingenuous way. Google pretended to be giving tips about accessibility or content, that tangentially also improve rankings. Websites pretended that their keyword stuffing was about usability or whatnot.
Google have been carrying this culture of euphemisation for almost 20 years now. Elephants stampede all over the meeting room. Everyone tries hard to pretend they don't hear the deafening trumpeting.
Youtube is an even more extreme example. Google's algorithms, policies and processes throw an entire media industry around, while Google pretend these are just minor side effects, or that it isn't happening. The gaslighting is awful.
They can look at how long people watch; but that ends up rewarding rambling pointless blather when a concise two minute video would be perfect.
They can look at likes or subscribes, but that causes many wasted human lifetimes per day, of people saying and listening to pleading to click said buttons.
Perhaps, but only because Google is being small minded. Define the "problem" in a certain way, and come up with certain solutions. A broader minded frame would define the problem(s) more broadly. Presenting users with UI. Creating a good, or at least respectful, commercial incentives framework for creators. The litmus for this is the content. What content gets created. Mitigating or avoiding bad consequences, like spiraling towards pointless blather.
Yes, this would require subjective decision making, but at least it doesn't require dim witted, corporate delusion. Call the dog by its name.
Maybe a search engine that could prioritize getting to the point would be good.
I was looking up some niche stuff in JavaScript space, and the article started out with a sentence:
"An increasing number of developers are looking to get started with web development. And because of this, they are searching to get started with <a fairly complex topic to grasp for a beginner>."
Sooo.... developers are looking to get started with development...
And like you say, the same goes for all the "What is a <topic>" - feels like a fundamental error on Google's part for parsing language intent.
Hard disagree. If the best content for the problem is a video, then sure, bubble that up to the top of the search results. But to limit yourself to only one medium is ridiculous.
People have very different preferences in the format of content they consume.
I've seen it more and more during the last few years, from ads around the city I live in to hipsterish magazines. There's something about it that has started bugging me the wrong way, I find it kind of infantilising but at the same time trying to "sell" me something as an adult (an insurance product, a "life is good" vibe because goofy drawings, that sort of thing). Or maybe I'm seeing too much into it.
Later edit: I'm talking about visuals like this one [1], which, looking at it again, I find quite similar in style to the one Google is using (hence my question, there must be a trend or something).
[1] https://www.reginamaria.ro/sites/default/files/inline-images...
https://en.wikipedia.org/wiki/Corporate_Memphis
https://knowyourmeme.com/memes/subcultures/corporate-art-sty...
http://clipart-library.com/data_images/177477.png
Guys in baggy suits and ties running through fields. You probably remember it from things like manual covers of scanners and stuff.
C'mon Google put Imagen to work, I'm so over this art style.
I think Google is part of a cartel; their search funnels people to sites that run ads that often are part of the Google network. Yes, they make a ton from search advertising but also a ton from the greater network.
This is really the Mob boss cracking down on underlings who have gone over the line and are collecting too much protection money or are cutting the "goods".
I find myself searching <product name> + "review reddit" to find real honest reviews. My anecdotal experience is that everything else is written by content creators who are given the product for free to use for a few days/weeks/months w/o going in depth.
Another one is <product name> + "forum" or <product name> + <some trusted forum>.
I train myself to use DDG, but, to my own dissapointment, I end up with "!g <query> site:something-trustworthy.com" way too often.
For regular Google searches, I've had the feeling of being "scammed" for years, though. The "fishing devil" on the no-results page also feels like buttering up the user. I always subconsciously interpret this as a lightly humiliating, but cowardly way to say "we decide what you see, bro". But, YMMV.
Never wanted to bash anybody, but this has got to be the most negative post I've ever written on HN. Please, please, please, Google, give me the search results of around 2008-2010. Pages, pages, pages of results for almost any query. I worked as a journalist then, and googling was actually useful, even educational, in real life.
https://chrome.google.com/webstore/detail/search-the-current...
Your example is a really good one, because it illustrates the power of OR remarkably well. The key is using it to combine several different operators, not just duplicating a single one, e.g. "query site:a.com | site:b.com". That has been my main way of digging deeper over the years, along with "#" to specify a date range.
Interestingly, in case of "query site:firstsite.ee|secondsite.ee" the "|" doesn't seem to work, I get zero results. Why is that? To narrow down the results based on both the domain name and top level domain, I have to add the inurl: operator. E.g. "query inurl:firstsite|secondsite site:.ee"
Strangely, "piping" is the essential thing I do on the Unix command line every day. I haven't put (that) much thought into using Google's "|" operator in a similar vein, that is, to combine several (3-4) different operators. Possibly partly because one used to get pages and pages of results even with very simple, single-operator queries.
Time to refresh my memory about Google's search operators, I guess. Thanks again for sharing this example.
Marketers may be dumb, but they're not stupid - all it takes is a few people to notice Google starting to add "reddit" to their autocomplete, start looking around, figure out what's going on, and then suddenly the whole "SEO"/marketing spam sphere is aware of it and will start making sockpuppet accounts.
Might want to check out https://plagiashield.com/
I use it occasionally but don't have an active subscription. But it's decent at finding stolen content (except for false positives on privacy policy & TOS)
By my understanding, this should hopefully cause Pinterest pages to disappear from image searches too, but by my experience, they're a big site so Google will have made an exception for them.
And it should likely deal with those StackOverflow 'clone' sites which sometimes rank better than StackOverflow itself.
These days there is no shortage of tools that monitor SERP fluctuation, and if this update brings about _significant_ impact it will be talked about everywhere.
And I fully expect spammers to go ballistic once their content gets obliterated.
Did they go ballistic when Google introduced the Panda update in 2010, killing many of the the web-scraped, machine-generated blogspam that was poisoning search results? No, they moved on to fresh territory like Youtube, gaming the recommendation engines. And they eventually came back to the web to poison results with hordes of awful 'best bicycle for fitness in 2022' type afilliate marketing sites.
Health and medical queries are the worst. The top results literally have the same generic content.
Maybe a no-filler health info website could be viable if it spread by word of mouth and not via Google. After a while it should become well-ranked on Google anyway, like Wikipedia or StackOverflow. Getting it to that point is a chicken-and-egg problem though.
It seems there's a new one I haven't seen before called appdividend blog, from India too. And another one "DelftStack", allegedly from Netherlands (tweets from 2019 and weird lack of info). programiz, from Nepal (you have to dig to find it).
What should these search pages give, in order: Official documentation, then Stack Exchange and similar, then real people's blogs, and at the very bottom on page 25 these spammy copy-n-paste sites/blogs.
Google is full of developers, so they have to know about this but for some reason the company does nothing about it. Almost like the company is fully driven by ad revenue and doesn't want users to find actual content.
I suspect they have an inventory problem, not a ranking problem. Why would any real humans publish publicly on the internet at this point?
If you’re posting for fun you have to deal with moderation and stolen content.
If you’re posting for pay Google just takes your content and shoves it onto the SERP, skipping your ads.
Bill Gates had it right. If you want a sustainable ecosystem, you need to make sure the other players are making more money than you in aggregate. As far as I can tell, that’s only true on the internet for eCommerce and so that’s all that’s left.
Unfortunately, what is not mentioned in the article is that sole existence of these content-farms is enabled by Adsense. So basically they are in the chicken and egg problem, and the only way to really solve it is to acknowledge that ad-based business models lead to detoriation of content on the web, get rid of Adsense (probably miniscule revenue contribution compared to Adwords) and then figure out how they resolve their own business model incentives and outcomes for users. Tough one.
You basically have to act as a mind-reading translator to understand what they actually did.
I know Google isn't going to reveal the exact secrets behind what changes they're making to do this to avoid black hat SEO, but haven't they been trying to do the above all along? What stops the SEO community from observing how rankings change, guessing what the new metrics are, and then optimising for these metrics as usual until the search spam is back where it started?
What about letting users give a vote/rating or leaving a comment on if a page was helpful? A big reason imo that Googling for "reddit <search term>" works is that spammy stuff on social sites (including Hacker News) gets punished quickly by spam filters, down-votes and negative comments - regular Google search lacks that because people can't down-vote or comment on a page. It feels inevitable to me they'll have to start doing something like this.
Negative seo is a thing and any commercial site that is remotely successful is targeted with it daily.
Difficult for such a system not to be gamed, e.g. via botnets.
Also voting is an extremely skewed signal, what people say they like and what they actually like are very different and with search you can actually measure that by seeing which results users actually go to and don’t bounce from.
I'm not saying it's as simple as adding an up/down vote widget next to each Google search result, but what can Google do to compete with social sites here where social sites have a lot of rich signals about what users prefer? There's no way to weigh or moderate each up/down vote based on how reliable the source is? What stops Hacker News, Reddit and Slashdot being overrun with spam from bot votes?
A better comparison is Amazon search results and reviews, which are known to suffer from being gamed.
This translates to "more YT videos made by our content-creators", right? Searching for "macbook pro m2 reviews" gives me 5 YT videos at the top, all repeating the same stuff over and over again and (no offense) made by people who aren't really experts. Or are they?
(/s, if it wasn't blatantly obvious.)
Nowadays I just append "site:reddit" to most of my queries (although I'm sure there's more and more paid/fake posts there as well)
Especially for the shopping category Google dropped the ball. When I am, wallet in hand, ready to buy a widget, why doesn't it do the best to help me find it? Isn't that the essence of search? Or just to search again and show more ads?
Such a failure of experience to try to use Google. Behaves like the doctor who secretly wants people to be sick so he can have more business.
But the writing is on the wall. Dialogue based search agents are coming, snippet based search is on the way out. The language models become better and better, they can even do sub-searches. The trend towards natural language based search is being accelerated by the mobile phones who lack proper keyboards and large screens.
The only way to get half-decent reviews at the moment is to append "reddit" to your search query, otherwise you get pages full of comparatives that haven't tested any of the items they recommend.
> "Currently, the predominant business model for commercial search engines is advertising. The goals of the advertising business model do not always correspond to providing quality search to users. For example, in our prototype search engine one of the top results for cellular phone is "The Effect of Cellular Phone Use Upon Driver Attention", a study which explains in great detail the distractions and risk associated with conversing on a cell phone while driving. This search result came up first because of its high importance as judged by the PageRank algorithm, an approximation of citation importance on the web [Page, 98].
> It is clear that a search engine which was taking money for showing cellular phone ads would have difficulty justifying the page that our system returned to its paying advertisers. For this type of reason and historical experience with other media [Bagdikian 83], we expect that advertising funded search engines will be inherently biased towards the advertisers and away from the needs of the consumers."
For meal recipes I would just like to see the recipe and not a blog post with dozens of ads.
For programming questions I don't want to see the sites that just scrape and repackage stackoverflow questions; for a company run by software developers they have to know about this problem lol.
Why not just buy a cookbook? It doesn't seem like the convenience of accessing recipes at the touch of a button is worth the tradeoff of ads.
Check it: https://search.marginalia.nu/search?query=scallops&profile=f...
[1] https://git.marginalia.nu/marginalia/marginalia.nu/src/branc...
This fact alone tells me they're not willing to take search result quality seriously. I'm sure medical professionals feel the same way about relevant content that is scraped and republished.
While it’s still early, I will say that I’m seeing better results from a variety of search results, whereas I had 2-3 consistent search performers that drove most of the traffic prior to the recent change.
Just the perspective from someone who runs a long-running tech-meets-history newsletter.
Sites containing "fonts and scripts" are not baggage. They have just been updated for modern times. Sites have to measure traffic somehow, and it's not going to be via a "Guest Counter" widget like they had back in the Geocities days.
You might be looking for an alt-web like that hosted on the Gemini network. It's all text and HTML-only as I understand.
https://searchmysite.net/ is a curated search for exactly what you describe though. It's been on HN a couple times.
This is why I like HN and artisan forums so much, or even Reddit.
Google could do just one thing - prioritize giving out the discussions of real people. But they just don’t wanna do it apparently
This phrasing sets off a little alarm in the back of my brain. Overall, I agree with the intent of the update and I think Google is maybe internally realizing that their questionable search quality combined with legitimate and privacy-focused competitors (DDG, Neeva [1]) could turn into a serious and existential threat to them.
That said, this is not quite the issue a lot of people have with Google search results. For me, when I search a movie for reviews, it’s not that I want new reviews to show up from smaller sites. It’s that I don’t want the whole screen filled with Google’s own little special sideshow for movie reviews. I can easily see this becoming even more frustrating actually if I search for “x reviews” and rather than get the usual Rotten Tomatoes, IMDB, newspaper reviews, I get RandomGuy.com. If I want RandomGuy.com, that is actually what a social media-esque site like Rotten Tomatoes or Letterboxd is for, not a search engine.
Hopefully this will be what we want it to be and not, for example, the TikTok-ization of Google search results, constantly and algorithmically showing you something “new.”
[1]: <rant> That said, the game is still very much for the taking and no one is still able to easily answer one of my most common queries for work: “python regex.” Seriously, it feels like a battle to get a simple example showing how to use groups/named groups to extract specific information. I don’t want a reference guide to common RegEx wildcards, I already know about \d and +! I don’t want a toy example showing how to use RegEx for seeing if “rain” is in “The cat is in the rain” (that’s using a hammer to kill a fly - you would just use “if rain in string” for example). I want a helpful and succinct guide on the specifics of Python’s RegEx so I can stop fiddling with it ASAP! </rant>
You mean content farms? Oh yea, more of that please.
Google has a giant hard problem to solve here. I hope they are successful.
If they aren't successful, it could be a door opening for someone else to take a stab at search.
Maybe somebody with a more analytical eye could elaborate on this. Thanks.
If google search really wants to improve results all they actually have to do is let people see the results instead of hiding them. 400 results is not enough.
I'm hopeful for this feature, but I'm not sure how scalable it can prove to be, to help manually rank pages for myriad of topics.
Sometimes adding "reddit" helps.
This is something youtube should be updating their algorithm for
O, wait. That's how the SEO hacking industry was born.
That’s not gonna change.
Especially if you’re searching for local services above the fold will be ads then maps listings.
I get it. But that blog post is some 1984 shit. Hahaha.
No one wants the back story for the "the best brownies", I want to know the ingredients, cooking temps (Celsius and Fahrenheit) and time.
I guarantee you that none of them are aiming to provide just a grocery list or instructions, they definitely enjoy writing their story as well. They’re building their own cookbook essentially, not a bunch of lists.
To be clear, I don't really care if the person wants to put their life story on their recipe site. I am annoyed that they are forced to do so in a really mechanical and disingenuous way just to make a Google algorithm happy - and as a result I lean heavily on the recipe sites that have alternate marketing channels/revenue streams (i.e. Serious Eats, ChefSteps, etc)
Im honesty asking since. In my mind (i could very well be wrong). I. would think there are more ppl interested in 'how to make Great brownies' , then there are ppl searching for 'the backstory of the great brownie recipe' ?
Yet the results i get from my searches are more inline with if i searched for backstory of recipes.
Hell maybe im just searching with the wrong keywords ??
I don't know the ins and outs but my understanding is that if you don't follow the Google format then you recipe drops way down in Google search results.
TBH I'm not entirely clear why the issue is so important to the people who publish these things, but clearly it is!
Otherwise, it is almost useless for most of the world.
1) Im mad cause now i need to google a convert query and author assumes the world is only full of U.S.A ppl
2) Mad cause I should know approximate conversion value, since I *think/claim" to be not a stupid person.
Both of the are usually false, in my case at least
and they're increasing surface area for the results to become more garbage.
what are trusted experts ? does google wanna be a publisher now ?
But until people pay for that search capability directly, the economic incentives will not align.
And there is the ever present temptation of accepting money from advertisers. This is potentially lucrative. Will consumers ever pay enough directly to offset seller interests to sway search results? I doubt it under the current regulatory framework.
We need changes.
Some people (policy people and economists mostly) know about Ronald Coase. One of his key points is we can rebalance who has the upper hand at the beginning of a market (initial conditions) and still have economic efficiency.
Coase's classic analysis compares two hypothetical worlds. In one world, smoking is legal and non-smokers have to compensate smokers in order to have a smoke free experience. In the other world the opposite is true. Both worlds can be completely economically efficient (defined as the equilibrium where there are no additional exchanges that make all parties better off). The difference between the worlds is purely distributional -- who gets to have more money.
Distributional questions have a considerable ethical component. Like most people, I prefer the second over the first.
Regulation could also similarly apply to search engines. We could tilt the balance in favor of consumers instead of sellers (in the case of public companies, this means investors).
What am I proposing? Simply put: more discussion about such options.
I will not try specify the best legislation for all situations in this already long comment. This does not mean no significant improvements are possible nor feasible. I simply do not want to get mired in debate over only one policy option.
---
About me: I listen to good libertarian style arguments as long as they are realistic about the standard economic model's limitations and failure modes: externalities, market power, imperfect information, nonrational consumers, and so on. / I lean liberal, so I care about transition costs too. Rapid change is hard because people have limited geographic mobility, especially in the short run. I also care about observable, measurable opportunities more than only hypothetical talk about options.
Hah. Just another day, someone sent me this link[1] which took images and text from my page[2] for the Google query "linux check disk space". Here is the thing my page doesn't even appear on Google. The ripped page points to the source[2] and has a link back to the original images. These scammers know how to game Google with their AI-driven sites, and they adopt it faster than Google rolling out new changes. I hope Google fix this issue.
[1] Spam/scam page - http://blog.imm.cnr.it/content/linux-check-disk-space-comman...
[2] Original my page published on 2016-01-23 - https://www.cyberciti.biz/faq/linux-check-disk-space-command...
[3] Google search query - https://www.google.com/search?q=linux+check+disk+space
* Use Copyscape and perform automated plagiarized content detection.
* Automate the sending of DMCA notices to Google and the website host (using captcha solving services, if necessary).
"Plagiarism" is not a legal thing: it's not a crime codified in law and doesn't appear in laws [1]. Instead, usually it's a violation of some internal code, e.g. a school's or newspaper's policy of integrity. As you say indeed, usually these policies only consider something plagiarism if there is a lack of (clear) attribution.
However copyright is encoded in law and copyright infringement is a crime. The linked page clearly violates the copyright of OP.
I've been burned in the past by relying on extensions like these, which only work until Google change their HTML and then the extension author is (understandably) burned out and doesn't update the extension, and eventually I just give up and uninstall it... I'd be a lot happier if this was a core Google Search feature. But I understand that Google don't make money by blocking their (paying) customers (spammers who run ads) from their search results.
!py urlunsplit
redirects you to:
Before using this, I was finding the first page of my search results were overwhelmed with StackOverflow-scrapers, which returned SO answers reformatted into some sort of garbage and unreadable 'blog-post'.
Nowadays, after blocking twenty/thirty/fifty(?) of these sites, I get to where I want as fast as I did 10 years ago.
If anyone want to share other such lists, however subjective, I would be grateful
edit: found another one, https://github.com/arosh/ublacklist-github-translation
I mean, I am 40 years old programmer/entrepreneur interested in sports and traditional board and card games. I like nice restaurants and hotels and even sometimes write Google reviews on them. I have my Gmail account since the beginning and I am paying YouTube and Google services customer. I am using Google phones since the Nexus. Google knows who I am, what I like, where I visit. It's pretty good at estimating revenue of my company. What about using the info for something useful for once and just show me pages people like me liked and don't show me pages people like me think are spam/scam? Please?
These are not scammers, it's probably an amateur researcher which published that information in a blog which is usually read by no-one. Probably it wasn't even meant to be public. But because of the prestige of the domain, it becomes first in relevant searches.
If you track down the author and send him a quick mail, I'm 100% sure they'll help.
I've worked at CNR.
There are plenty of scam pages, but specifically this one doesn't look like it
More like someone who didn’t care much about copyright or license decided to back up information they found useful in their personal blog.
It's good they linked to the original post and that the link is the first thing we see, but nothing says that it's the source, and no paragraph explains that the content comes from elsewhere.
I also don't see why the content should be copy-pasted instead of just a link.
I see no malicious intents, it's probably done in good faith, but meh.
To the author: did you try to reach out? I'm sure something can be done about it if it bothers you. I expect the author of this copy to be receptive.
Quite possibly an easy way to preserve the content in case the original goes offline and to share it with colleagues. Perhaps they wanted to link to it from some long-lived or even printed material. Not saying it’s the best way to do so, but it’s plausible and not malicious.
This is a public blog, not some internal website.
Anyway, my previous comment probably sounds harsh because it is the way I wrote it (because I previously worked in a research lab, so I kinda feel disappointed), but I still consider this a minor fuck up and it happens to everyone, for sure.
Ah, but this is the new method of SEO trickery and credential scamming. Publishing a 'guest post' on a high-ranking blog subdomain of a trusted instititution. There was a story of someone doing this on Harvard University's blogs, which I can't find right now.
But I found something even better. An actual UpWork posting promising to publish your crap on Chapman.edu's university blog:
https://www.upwork.com/services/product/5-high-da-dofollow-g...
It has a garbage/content ratio of 32 (the browser downloads 32 bytes for every byte of content) while the original page has a 400 ratio (the browser downloads 400 bytes for every byte of content). It's borderline denial of service attack against the visitor.
The original also has a slow aggressive cookie box that is unnecessary as visitors can be spied on using server logs, no need to have cookies for the spy infrastructure. The bootleg has no cookie box (though does set a couple of gratuitous cookies, and does load one gratuitous googleanalytics.js, so maybe it does the right thing in a non-compliant way).
Maybe all these factor make the original page look "less good" to google quality ranking? (which would be ironic given Google's general leadership of the Orgy of Waste school of web design.) Also if google starts ranking such pages up they may expose themselves to class action lawsuits as users could ask for a refund for the power, hardware and telecom bills incurred in being lead to load them?
Sorry, but this is completely insane. Search engines should most certainly and definitely not ever be liable for linking to original sources for such a stupid reason. This American pro-litigious attitude has to stop.
In the physical world one can't dump a rusty aircraft carrier on someone else's lawn and get away with it. Should be the same in tech.
Google's search results are most definitely not your lawn, nor are heavy websites very much alike to rusty aircraft carriers. Or however this analogy is supposed to work anyway.
What is this referring to?
Sure, goatse people if they refuse to observe your requests if you feel like it.
#2 cyberciti.biz
#3 cnr.it[0] https://opensource.com/article/18/7/how-check-free-disk-spac...
Content farms run AdSense.
Google algo updates have revenue as one of the core metrics.
While this is almost trivially true based on what we know of (the world, big tech companies, Google in particular, etc), do we have direct supporting evidence for this?
Well you see already in ancient rome they made bread this way...............
[2 pages of story time]
AD
AD
AD
Ingredients separated by more Ads
What you may witness is less amazon affiliate posts and more made for AdSense posts - such as those recipes that begin with a random love story, for example.
In the interim they tried to push AMP, but had to admit finally that it's more to keep people inside their ecosystem than make the web faster/better.
Why was the company blatantly denying obvious issues for so many years? What has changed? Why should I trust their judgment all of a sudden?
It's kind of like someone who is a pathologic liar having been caught in various lies for many years, swearing to tell you the truth, but not admitting to the past lies.