Why Is the Web So Monotonous? Google
reasonablypolymorphic.com
reasonablypolymorphic.com
From the article:
> Lets look at some examples. One of my favorite places in the world is Koh Lanta, Thailand. When traveling, I’m always on the lookout for places that give off the Koh Lanta vibe. What does that mean? Hard to say, exactly, but having tourist amenities without being touristy. Charming, slow, cheap. I don’t know exactly; if I did, it’d be easier to find. Anyway, forgetting that Google is bad at long tails, I search for what is the koh lanta of croatia? and get:
This is a near impossible query for human beings, let alone for computer given the state of the AI at this point in time.
What we get instead are incorrect results which are being presented as reasonable answers to the search query in the name of advertising revenue.
When I discovered Google (I think it was running at beta.stanford.edu at the time) it was just short of a miracle.
There is no blank space to fill with something useful.
Instead everything is filled with SEO driven, ad riddled, bullshit.
And creating something useful is more cost and effort than creating more SEO driven crap.
And even if you do create something useful, you’d then need to focus on SEO in order to compete.
There’s no way this problem isn’t solvable. But google seems to lack any motivation to solve it.
Their apparently paying a lot less attention to certain things (like link text—"what might someone use to describe a link to this resource, but which might not appear on the resource itself?" used to be a very fruitful way to search), and freely substituting words or dropping terms that are merely not common on hits (but not totally absent from 100% of results) has made this kind of thing impossible.
That being said, I remember my cousin had a blog where he was writing utter nonsense about trending topics and he was actually able to pull enough money from Ad Sense to afford a pub crawl once a month.
Now of course all of this is streamlined - one button website generators with AI content and posts for given topic and self optimising for engagement and advert clicks etc.
But I see Google has thrown in the towel and no longer cares for search results.
It sounds like this person wants to have a philosophical discussion with a friend about this type of subjective question rather than simply receive a factual search engine response.
> Places where things move at a less hurried pace, where Croatian life can be savoured, where you get a flavour of what the Dalmatians call fjaka – the art of doing nothing. These islands and mainland destinations are what you want in a post-lockdown escape: peace, beauty and the chance to discover why Croatia is such an enticing country.
Now I know I want to search for fjaka -- if I want to search further because this article might just be what OP wanted.
The success for search always has been finding the keywords to search for. You can see here it took me just a few minutes.
Once I lived in the city, it was too big and noisy
So I moved to the country to stop and smell the rosies
All my city friends joined me and put up nice new housies
Now it's too big and noisy, think I'll move to the country
The reason even a fairly advance NLP system is going to struggle to find what the author considers to be the "Ko Lanta of Croatia" is that they can't read the author's mind about qualities associated with Ko Lanta, so they're going to highlight islands, not "small, quiet resort villages with $10-$20 rooms and hammocks" (Not really Croatia's "vibe" though I suspect the OP would probably like most Croatian islands anyway...). The same is going to be true even if you put in strings that return comparisons humans directly make themselves like "Venice of Asia".
Well, I was about to post a similar reply. It would have taken some messing around with the search string, but I'm confident I would be able to find something like this on Google 18 years ago, and I'm certain that if it exists, I can't find it now (unless it's on some big name site).
I don't think Google of 15 years ago will do a good job on the current age of internet. The size of web has increased a lot like 100x where most of the increment comes from unstructured formats like video or image and its signal to noise ratio deteriorated considerably. SEO has become much more sophisticated ever. A large fraction of useful information is now locked in the unindexed walled garden.
In fact it could be less than 1x in many cases. People that would have blogged about their Croatia trip might now instead just have a 10 word tweet, an instagram story (of similar low quality and probably not publicly indexable), or a youtube video (so non text content)
And then you learn that on a good day you could see Croatia from Venice, or something like that.
(I just completely made that up, don't quote me on that.)
It was less likely to ignore the absence of a particular keyword but certainly wasn't better at NLP
Fed a bunch of countries coastline into it, and popped out ranked beachfronts to investigate more.
1. I once went to Ao Nang, Thailand a few years after the tsunami and found a really cool chill cheap beach town and loved it. Went back a few years later and right on the main street there was a McDonalds. Across from that there was a Burger king, ~100M up the road was another McDonalds. I will never return.
Where did you get the data from? Do you still have the code somewhere?
It worked for a select group of toy problem domains and the answer for most high-concept searches would be noise or nothing.
The author never manages to convince me that the information they're looking for is on the Internet for Google to index in the first place. The author then asserts "don't just hit me with garbage," which is a fine assertion, except what they're interpreting as 'garbage' is the information other people making the search could actually use. Google A/B tested the hell out of "say nothing" vs. "guess something close to what the requester might care about" and the latter won out every time.
If this data is there now, it is most likely in a Facebook post which Google isn't seeing.
Google could find them on Archive but we both know that's not the business.
Ugh, this reminds me of the people that move from high tax places to low tax places but want all the benefits of high tax places at low cost.
Now, back to the main point, if this was searchable it wouldn't exist.
The number one complaint people have today is popular places are too busy, too many tourists come and prices go up and it's over ran. So what do you think is going to happen to some place that 300+ million wealthy people around the world can search in a few seconds and find out it's cheap an empty. The answer is "become busy and expensive".
The only reason places like this still exist is Google can't find them.
Here's a tip for the blog author. Learn a foreign language. Then search google in that foreign language. You might be surprised. Be the change you want to see.
Youtube is worse, I have two YT accounts in English where I only watch videos in English, and Youtube keeps recommending me videos in French that are not relevant at all to my interests (even though most recommendations in English are perfectly relevant). It's as if their AI just saw my IP and decided to recommend to me the most popular French videos of the moment.
What I needed was to use a different search engine. DuckDuckGo gives me a toggle to include or not French results.
I passed back through a few months later and barely recognised the place. Wooden chalets has given way to hotel rooms and beachside stalls to night clubs.
The whole point of even having tourist amenities is that there's a community that that's attempting to draw tourist money to their location, and probably create a thriving economy.
An unknown tourist attraction with amenities is a failed location where the amenities are unmaintained/dilapidated, the local economy has nosedived, the young have moved out and the population is collapsing.
In the context of the discussion, I don't understand your question. The money comes from the hypothetical tourists who're hypothetically making all the usual spots too busy. What point are you making?
> there's a community that that's attempting to draw tourist money to their location, and probably create a thriving economy
Usually that's not how it works (at least from first-hand experience). It's about a local business/mafia trying to draw tourist money while exploiting cheap labor from the local community. Let's not even get started on the scheming to privatize public beaches, environmental pollution (noise, lights, trash) and other shenanigans the tourist industry is creating to the detriment of local communities.
Like, Koh Lanta is surely not just a vacation destination, it's also a place where people live their lives, go to work and school, are homeless, deal with illness. The expectation that Google will assume that by "koh lanta of croatia" you're really asking "vacation destinations in Croatia that are similar to Koh Lanta, Thailand" really rubs me the wrong way
> Google has gotten exceedingly good at organizing everyday life. It reliably gets me news, recipes...
What I mainly hear on this site and elsewhere is people complaining about how the news is partisan rubbish, fake, and serves the "elites". And recipes. Seriously? I had a half dozen people berate me earlier for suggesting that cooking your own meals should be an enjoyable part of life.
Nobody wants news and recipes.
Here's the problem.. "The Internet" was military project that got loose. "The Web" was a solution for research scientists to exchange papers. That's all. Driven by a massively profitable industry, a solution looking for a problem expanded in expectations, and took over the consciousness of generations of an entire society.
"Organising everyday life" is an inadequate appraisal. Most businesses had to be dragged screaming and kicking online. We still possess a vague, poorly thought-through ideal that somehow a connected, "technological" society is a better one, and a thirst for convenience and speed. But let's be honest, we don't really know what the net is for. It has no telos or guiding design principle other than what we overlay on it.
Military tools in the hands of "the people" tend towards insurgency and revolution. Wasn't The Internet's "killer app" the Arab Spring? The brightest moment of the Web was in it's formative years, as an explosion of ordinary speech and new political power. Much since then has been an reaction to try dampening it, domesticating netizens, or even putting the genie back in the bottle.
Maybe a reason to be optimistic, excited even, is that we're still in the infancy of the network. It's still pregnant with unimaginable possibility.
What seems like an endless army of Google proponents (perhaps they are simply "status quo" proponents) discourage anyone from even attempting to think about how the web could be organised without the need for a Google.
In a library, I can do searches restricted to certain subject areas. I might only be searching certain databases that pertain to certain subjects.
Whether it could ever be feasible to search the www by "subject area" is left as a question for the reader. In any event, if it were I would search for "koh lanta" in a relevant subject area, e.g., "travel reviews". One can search for pages that contain "koh lanta" on www sites using domain names registered from the .hr registry, e.g.,
"koh lanta" site:.hr
What if all travel sites registered names from some registry, .travelreviews, and one could search pages that contain "koh lanta" on www sites using registered .travelreviews domains.Even if I could just obtain a list of travel review websites, I could index those sites and then search that index for reviews that contain the terms (a) "croatia", or names of various destinations in Croatia, and (b) "koh lanta". If I was really patient I could use a script that simply performed the same search on each site, either using the site's "search" option or a search engine.
I am still not sure this search even makes sense. It seems to rely on an assumption that some traveller will compare Koh Lanta with some location in Croatia. What could make more sense is to define what are the specific charactersitics of Koh Lanta that one wants to find in a Croatian destination, then search for those. Then I might search not only travel reviews but sites that describe characteristics of Croatian travel destinations.
In any event, it is this sort of catgorisation of sites that is generally missing from the www. I believe it is feasible but "tech" companies like Google and its followers are not interested in promoting such facilities, preferring instead to pursue data collection about www users and programmatic online advertising. If Google is the "front page" of the www, then it can portray the www in the way it sees most beneficial to Google. For example, an endless sea of disorganised information that is impossible to utilise without Google's assistance. I certainly do not need to search billions of web pages to find the "koh lanta of croatia". I only need to search sites with travel reviews. But I do not get to limit a search like that. An advertising company gets to decide what is "relevant" to the terms I input. _Popularity_ (potential value for advertising to a wide audience) dictates what I can and cannot see of the www through Google.
About 10 years ago I started a proof of concept that aimed to catgorise a web (cf. "the" web) via non-ICANN issued domain names issued by a non-commercial domain name registry not governed by ICANN. The idea is that the FQDN can contain 1. a subject matter description (subdomain) 2. a trademark (domain) and 3. a Nice trademark class (TLD). Searches can then be done largely based on FQDN instead of heuristics such as "page rank", popularity or other metrics designed for advertising purposes. This new web has value to me because it is constructed from legitimate organisations that have invested in trademarks, i.e., they have a registered address and they pay their legal bills. Registering a mark suggests they have an existing or future brand to protect, a business that needs a mark to protect consumers from becoming confused about the origin of a product or service. (How many of those Chinese sellers on Amazon today have trademarks.) Ideally, this system filters out those who are not legitimate businesses. The garbage at the top of Google would not be possible. To game this system would require obtaining trademarks, not simply gaming ICANN's system of registries and registrars. (And we know ICANN itself is not a trusted steward of DNS, but a means for a select few to make huge profits from it.)
This prototype web was not meant as a replacement for "the" web, but as an experiment to separate out the commercial entities on the www from the non-commercial ones. Google and its ilk want a www where there is no distinction between commercial and non-commercial www use. All www use is surveilled for commercial purposes.
This mixing of commercial with non-commercial, to me poses one of the biggest threats to the www we started using in the early 1990s being wiped out. In the early days, there was this idea that TLDs would represent different categories of websites. For example, ".org" would be non-commercial websites, ".edu" would be educational institutions, and so on. Things have changed. The root.zone has exploded in size with countless "gTLDs", most of them are purely commercial. History is showing that it is infeasible to demand that TLDs enforce some sort of rules over the contents of websites that use them. One needs "Google" to help figure out what websites are worth visiting.
The prototype web in theory allows one to find legitimate brands on the www by searching registered marks or subject matter descriptions within Nice trademark classes that are contained in FQDNs, not the contents of web pages. It is the antithesis of what one sees on Amazon today, what with a gazillion Chinese knock-offs, including Amazon's own "brands". There is no "SEO" on this new web because the contents of web pages, e.g., "backlinks", are not the basis for search.
But do not worry, the prototype web does not exist and would never work. It is a terrible idea. Google will never die. See you in the Metaverse.
Well exactly. If anyone ever mentioned Koh Lanta in their review of Croatian destinations, we'd hope that that it would be at the top of the results.
Or that there would be no results.
Getting results that don't even mention Croatia is not the intended outcome.
Forgot about 4. an ISO-3166 country code. Thus, the FQDN becomes 1. subject matter description (subdomain), 2. trademark (subdomain) 3. Nice TM class (domain) and 4. ISO-3166 country code (TLD).
Also which country's trademark office would be the gate keeper? There's one in every country. I assume you mean the USA trademark database in this case, but that only covers US trademarks.
This is irrelevant to the purpose of the "prototype web". It is a different problem. The purpose of the prototype web is not "to make things much more difficult" for "spammers". Its purpose is to _separate_ the commercial web from the non-commercial web and to _categorise_ the commercial web in a way that makes web search easier.
Given that trademarks are allegedly "easy and cheap to get", akrymski could share with us how many he has registered. Surely it would be equal to the number of domain names he has registered since the UDRP favours trademarks registrants and trademark holders have the additional option of using the ACPA. It would make sense to have an "easy and cheap" trademark for each domain name, registered in every class in every territory.^1 How much would that cost.
1. Because domain names have no such limitations.
"Also which country's trademark office would be the gate keeper? There's one in every country. I assume you mean the USA trademark database in this case, but that only covers US trademarks."
It depends on the ISO-3166-1 country code TLD. If the country code is US then the USPTO would be the applicable office.
I really love the vibe of Koh Lanta in Thailand. It has tourist amenities without really being touristy. I find it charming, slow and cheap. Is there a place in Croatia like it?
Completion:
There is no exact place in Croatia that has the same vibe as Koh Lanta in Thailand. However, there are places in Croatia that have a similar charm and slower pace. These places include Hvar Island, Korcula Island, and Istria. While these places do have some tourist amenities, they are not as touristy as Koh Lanta and still offer a more authentic Croatian experience.
Another:
There is no exact equivalent to Koh Lanta in Croatia, but the small town of Nin on the Dalmatian coast comes close. Nin is a charming, slow-paced town with beautiful beaches and a relaxed atmosphere. It is also quite inexpensive by Croatian standards.
It will require multiple dozen auxillary Google searches and reading through countless articles and message boards.
GPT-3 is a great research tool as a brainstorming engine, but sadly cannot currently be formally trusted to output truthful information, even with conditioning.
It takes someone with prior knowledge or subsequent research in order to make use of the result. I'd say this isn't what the author was intending.
Humans, unless they're disturbed or malicious, tend not to hallucinate things out of whole cloth.
Something like coming to accept some data-set local optima as a global truth and deriving beliefs without underlying basis.
The whole point of these models is to make up something that seems "realistic", but not regurgitate the original contents. If that's not hallucination I don't know what is.
Me: What is your name? Model: I'm Mark A. <continue> Dreyer, and I'm with the <continue> Federal Communications Commission (FCC)
Obviously, it's not actually a guy named Mark A. Dreyer and, as far as I can tell, there is no Mark. A. Dreyer in a prominent role at the FCC. That, as far as I'm concerned, is essentially a hallucination.
If you think about how they work, this isn't totally surprising. The model has been programmed to take the input, turn it into a series of tokens, and then produce another series of tokens that are especially likely to follow the input ones. That's it.
To the extent that this corresponds to reality, it's only because reality is overrepresented in the training data, and so token sequences corresponding to reality are often more likely under its model. This also explains how it goes off the rails: Mark is a very common name (in the top 25 for men for decades) and once you generate a stereotypical middle-aged white guy name, it's not a huge leap to him being a bureaucrat, etc. I forgot to delete the huggingface boilerplate, which talks about advances in technology (etc), which probably sent it down the FCC route too.
Prompt: "Is Fortigate open source?"
Article: "Based on the parameters we compared, it seems that pfSense is the more favorable solution because it is open source and also offers great features."
Google: "It is open source!"
query: what is the koh lanta of croatia?
response:
Koh Lanta is an island in Thailand that's known for its beautiful beaches and clear waters. It's a popular destination for tourists, and it's also become a popular spot for digital nomads in recent years.
Croatia doesn't have an island that's quite like Koh Lanta, but there are a few options that come close. The island of Hvar is a good option, as it has beautiful beaches and a laid-back vibe. The island of Korcula is another option, and it's known for its medieval villages and stunning scenery.
I wonder how long it will take before we have a GPT-3 bot commenting here on HN
Needless to say, all of these people get told to fuck off and get a job.
https://www.theverge.com/2020/8/16/21371049/gpt3-hacker-news...
I could write pages of this stuff, but let me write more positive things. The best way to enjoy the islands in Croatia is actually by boat, rented or otherwise (they are not expensive). There are beautiful quiet coves where you can spend your days and nights, only going for resupply in towns. You anchor in a quiet cove, and in the evening you take your tender boat to the town if you want to walk around it. Next day you sail to another island and repeat. That experience is actually awesome.
If a boat is out of the question, being in some apartment is an option, I would just advise not to be in the town center or near it, or anywhere where some very local beach is. Mostly there will be a caffee with blaring music and nobody likes to sleep with earplugs. Also this is and advice for any island, and true, a lot of those islands are quite different than each other.
This shows, if you had any doubt about it, that it did not actually understand the question.
Anecdote: When I told a random person I got there by plane he automatically assumed I owned a plane."
All the results it gave are flooded with tourists, and some of them are among the most expensive you can get in Croatia. :shrug:
Unsure of what to think of the results, as it seems to have very little relevance to the input description other than being in Croatia and being tourist spots.
Tried and failed.
How would that prevent spam?
As is every query that we need computers for (e.g., a SELECT query that goes over millions of records).
On the other hand, expecting a search engine to find pages which at least mention both Croatia and Koh Lanta is a very, very, very low bar.
And it's going to be better than what Google is giving today.
The real answer to this question is that walled-garden social media took over everything. The much-pined-for "Old Google" worked because people used to actually create content on their own sites, and not just post it on walled garden monolithic social sites like Facebook or Twitter.
I've had very little problems cutting through the SEO spam running an independent search engine, but then I don't shape traffic like they do.
I love GraphQL, and I think it has a number of features that make it better and more usable for most APIs than a RESTful interface.
The thing I despise about GraphQL is they put "QL" in the name, so legions of software developers think that it is somehow comparable to SQL, or somehow is a generic query language for data.
GraphQL has absolutely nothing to do with SQL. Comparing "GraphQL vs SQL" is like comparing "HTML vs Java".
In what sense is GraphQL not a generic query language for data?
Those are all the right analogies for comparing GraphQL. GraphQL doesn't say anything about how the underlying data is queried - at the end of the day it's just a bunch of resolver functions that you can implement however you want.
I'm sure it's not Google's intention to ruin the web, it's just something they happily accept as a by-product of their optimization strategy towards maximum value extraction via ads. Like some chicken farmers happily accepting creating chicken hell because it lowers cost. I don't think they set out to torture chickens and go "hey, I'll start a chicken farm as a cover story".
Just as in nature, monoculture rewards specialization, whereas ecological diversity rewards generalization. Not that specialization doesn't exist in diverse ecosystems, it's just not as devastating.
But they didn’t want to do that anymore. Why? Because of ads money.
To drive the point home, it is not that difficult for Google to have a junk score and simply phase out the junkiest of junks. But they didn’t even want to do that.
Google released several updates (such as "Panda") that greatly lowers the rank of domains that aren't linked to by major domains they deem to be trusted, such as those with edu or gov TLDs or whatever whitelist of domains they decided to add. There have also been updates that specifically lowers the rank of websites that are running forum or blog software. This was probably done to fight spammers but it has effectively killed off the long tail of topics.
This is admittedly speculation since the inner workings of Google search are not released to the public. But these findings have been corroborated by others on the internet.
I used to have a blog and several forums that used to rank highly for some niche keywords and over the years the search rankings started to drop off the front page while a single reddit comment with no content mentioning the same keywords would be #1.
I'd say that Google is very much to blame for this situation as they took the easy way out to fight spammers.
No, no, Eric Schmidt openly said that "Brands is how you sort out the 'mess'" (in search). And he started prioritizing big brands in search results and disappearing everyone else.
Panda even wiped out individual webmasters or small software houses - they were adding small back links to the websites that they built for their clients, per google's OWN recommendations, for years.
Then google turned around and suddenly penalized all of those legitimate links while 'sorting out the spam'. millions of small businesses, developers, software houses found themselves with zero traffic within a day. their hard earned traffic coming from legitimate business clients. whose sites were also affected similarly.
This caused the rise in internet marketplaces. from amazon to elance, upwork. because the small businesses and individuals were now totally invisible in search results and instead 'brands' dominated. a side effect was killing of independent content in search results and forcing everyone to have to post in social networks for visibility, forcing content creators out of their own blogs into social networks. Which exacerbated the damages that algorithms because now everyone had to obey individual corporations' unaccountable algorithms.
So yes, google has created this mess to a very large degree. Their nonchalant, uncaring attitude towards their users, customers that plagues all of their products crippled search as well. By saying "Brands is how you sort out the mess" and enforcing it top-down, one single ceo single handedly decided the fate of millions of small businesses, blogs, professionals. In a totally unaccountable way with no input from anyone affected.
Speaking to a group of magazine executives at the Google headquarters, Google CEO Eric Schmidt said yesterday that the Internet is becoming a breeding ground for false information, reports Ad age. But trusted brands help weed through the disinformation:
"Brands are the solution, not the problem... Brands are how you sort out the cesspool."
https://battellemedia.com/archives/2008/10/cesspool_brands_e...
"Brand affinity is clearly hard wired….It is so fundamental to human existence that it’s not going away. It must have a genetic component."
Well, I built sort of exactly this: https://search.marginalia.nu/
It's not great, but it sure has its moments.
Should be noted, regarding Koh Lanta, that travel is one of the most aggressively SEO-spammed topics, along with pharma and online casinos. It's extremely difficult to cut through the noise and reach any sort of signal.
Point is, I don't think SEO at the time was any more or less "aggressive", per se. Just different. And clearly, the game has since escalated altogether.
What's the elevator pitch? What do you actually do that is special?
From the about section:
> This is an independent DIY search engine that focuses on non-commercial content, and attempts to show you sites you perhaps weren't aware of in favor of the sort of sites you probably already knew existed.
> The software for this search engine is all custom-built, and all crawling and indexing is done in-house. The project is open source. Feel free to poke about in the source code or contribute to the development!
But that doesn't really give me much info.
For one I have a sort of budget for how much javascript I will tolerate. Some is fine, like your standard wordpress config probably will fly, but not much more. I also do a personalized pagerank biased toward the blogosphere. The likelihood your website shows up in the results is directly determined by whether real humans link to the website.
It's not a Google replacement by any measure, it's meant as a complement, but if you are looking for more in-depth content on a topic than you can find elsewhere, it's often a good starting point.
What about sites built with React, etc? Like would you be filtering out cool things like Ableton's "Learning Synths"[1] or do you mostly mean third party scripts?
Do babies go out with the bathwater? Almost certainly. But when you're running a small scale search engine, you're never going to index everything anyway, so that's entirely fine.
(also low value, perhaps :)
FWIW I think it is cool and much needed.
But this was indeed a great time. Remember being on a social network ca 1999 and people used it as a diary where others could comment, give advice and meet others with similar interests. There was a timeline and posts from people you "followed". I think it didn't have likes or ratings though. Most people were super friendly and we were doing meet ups, so very much everyone knew each other. At peak time there was about 20k people and at that point owners couldn't cope with it. With that amount of people you'll find some nasty ones that will ruin it for everyone else and that's what happened. Owners couldn't afford to host it, so they were doing crowdfunding to cover the server costs and maintenance. Some bitter people didn't like that and started making unfounded accusations of theft or that they will report everything to the tax man and the police. So owners eventually closed it.
About ~5k USD total investment including a small UPS, operational costs are $40/mo. Doesn't actually require all that much maintenance, at least not compared to the work I've put into building a search engine (including crawler, index) from scratch.
I was going to suggest that perhaps adding info about whether the site had ads, etc. might help the usability. Then I tried marginalia ... :) Perhaps we need to invest in similar detection for other antipatterns (listicles, blogspam etc.) - this is mostly what I was talking about. Less the understanding for indexing, and more the classification for user safety and authenticity of results.
But again, this is very cool.
I don't know how useful these "hobo signs" actually are, but they did feel like a natural response to a lot of web sites justifying their use of trackers and cookies and affiliate links and whatnot with "everyone is doing this!"
The big problem is that Google is the major way people discover goods and services. If you have a site that sells garden hoses, the best way to get eyeballs is to write blog posts that a potential user of garden hoses might search for. "How to install a garden hose." "What is the best garden hose."
This kind of content marketing can drive huge traffic and thus huge sales.
But only results on the first page of Google matter, and the farther down the first page you are, the less traffic you get.
Now, the 15 companies that produce garden hoses are fighting against each other to get the 10 slots in Google's first page for any given keyword.
With AI-generated articles, it's become a race to the bottom. Google rewards AI-generated, keyword-optimized blog spam. If you want to be one of the 15 companies that makes it into those 10 slots, you better believe you need to write AI-generated blogspam too. And when everyone has to do that in order to compete, all that's left is the AI-generated blogspam that has infested the modern web.
You almost have no choice as a business. Either write AI-generated blogspam that Google loves, or your competition will, and bury you in the search results.
> Anyway, forgetting that Google is bad at long tails, I search for `what is the koh lanta of croatia?`
Well, in at least one significant case Google searches are more helpful than they used to be, in that they now take into account natural language queries like "what is..." and try pretty hard to DTRT. Contrast that to "the internet of yore" that the author yearns for, where search engines - including Google - treated all words as keywords, and tended to weight them in the order they were given, so that the query would mostly look for pages containing the word "what", and "is" and "the", and then hope to find a subset containing "koh" and "lanta", and finally, almost incidentally, rank those by whether they might seem relevant to "croatia".
Incidentally, those old-style keyword searches may still provide better results in some cases. Try searching `croatia place charming slow cheap`, and see how that compares.
TRTD works fine though.
In the past, Mr Joe Blogs was into a passion, he would write about it on his own shoddy blog with gif links to other blogs part of a blog ring of nerdy stuff he cares about.
Nowadays, Mr Joe Blogs writes on Medium, shares on Facebook and potentially has an outdated instagram account with 4 pictures of his passion. All buried down by algorithm to never be found, because his passion "isn't sexy".
The closed web and their algorithms are just as much to blame.
From what I've observed, the opposite tends to be true more often than not. Plenty of people have been able to create communities around niche interests because platforms bring something independent blogs never really did - an audience. Joe Blogs probably started a subreddit and facebook group for his passion and gets more people reading his content in a week than likely happened to stumble across his blog in a year.
It's a good theory, but becomes circular when you consider that one of the primary ways to determine search intent is . . . seeing what does well on Google.
So if you search for, say, "customer feedback", you see a lot of general guides about strategy, offering general definitions. That's likely because the first couple pages that ranked did it, and then more players saw that and said "oh, search intent for 'customer feedback' is [whatever they are doing]."
When you see all the listicles for certain queries, that's what you're seeing. Yes, it probably reflects some desire from people to see a list of options. But it's also just SEO people saying "look, the search intent is X." It leads to everyone copying each other and reduces the incentive to try and be innovative/think outside the box with content.
Search is ridiculously hard. Everyone underestimates the forces you are up against with SEO/Spam industry. Microsoft has invested into Bing for 13 years and has no market share and the #1 search term is "google."
I recently decided to create my own landing page, stored locally on disk, with a funny gif and favicon.ico, and a shortlist of my most used work and personal links (including this site!).
No ads, no tracking. Shockingly - I had to install browser extensions to override the new tab page in both Edge and Chrome. This used to be a built in setting! Shady guys. Very shady.
Although... Not as shady as the fact that Chrome by default sends your entire browsing history to Google for analysis.
Citation needed. Are you saying that, as I click around the site here, Google knows what comments threads I'm interested in?
Edit: I think they're talking about the History sync, which does indeed by default send Google every link you visit, but only if you explicitly enable it. We should get access to this data and leak porn preferences of every senator. That ought to get the laws changed quickly.
Settings > Sync and Google Services > Other Google Services > Make Searches and Browsing Better.
Underneath the setting it explains, "Send URLs of pages you visit to Google".
I think Today is the day I try moving back to Firefox again. I hope it's gotten snappier.
I would 1000000% never enable this setting. And I suspect that Goog is pulling a Zuck with "accidentally" re-enabling it and hoping nobody notices.
I have Termux installed and there are a few options for running a simple webserver (e.g., via Python). I've yet to look into hosting my own homepage and directory there, though ... it would be nice and useful to do this.
Their discussion forum is https://www.resource-zone.com/forum/
If anyone is curious, this is at least part due to copyright law. A recipe is a set of facts, which is no copyrighted. But recipes with substantial literary expression (making it unique) can be copyrighted. Content creators don’t want their stuff copied so viola…that’s how you get a long winded story before the goods.
https://copyrightalliance.org/are-recipes-cookbooks-protecte....
Also, getting people to pay you money is now a proven concept on the Internet. Whereas, in the early days, not only where there few options for taking payments or for monetizing content with ads, but it also wouldn’t have been worth the effort because that just wasn’t a big enough target audience.
The current document-centric approach has an intensely graphical (attention grabbing) aspect which very much encourages the least common denominator types of results that the author abhors.
Those who want to be more thoughtful about these things need better toolsets to allow them to focus the way that they go about their online lives (including general search).
Google is in indeed more than search. They are also the machine learning framework, Tensorflow. Linux on the Web is in as good a position as anything else to start putting the JavaScript implementation (Tensorflow.js) to very good use.
We don't need new search engines as much as we need new search engine interfaces.
Another explanation is that the type of company that buys thousands of links is also the type that pays rock bottom for their writers. I don't think you need to necessarily write trite crap in order to do well in SEO.
> Why don’t other search engines compete on search results? It can’t be hard to do better than Google for the long tail.
I think that's the "citation needed" that the author's missing. It's not like Google is unaware of this problem, and a huge piece of their research spend goes into improving search.
It's entirely possible nobody's doing better than Google because nobody knows how. The entire paradigm Google's leading in might be a saddle-point with significant activation energy needed to escape it while still having something as usually-useful as Google.
They mirrored the sentiment of the author that search results have gotten significantly worse and do not provide much auxiliary insights. The noise/hit ratio is really high.
Interestingly, Google Search works really well for Stack Overflow.
Maybe it's because people who use these kinds of websites are generally less tolerant of BS websites than the average person, so Google tends to derank the crappy ones just because we always click away instantly.
These two things have co-evolved in response to each other, and what this article laments is the result of exactly that process.
But the interaction between SEO and Google search is way more complex than the article implies. In particular, the influence goes both ways, and it's not a clear one-way cause and effect relationship. Also playing into this are a lot of other factors, like the increasing commercialization of the web, how the online advertising business works, Googles dominant position in both web search and web advertising, etc...
I really don't think the solution to this can be as simple as "change how Google rewards keywords". It's not that simple.
Google "reminds me of Koh Lanta" turns up Boracay, Philippines in the top three links
https://www.google.com/search?q=%22reminds+me+of+Koh-Lanta%2...
No you ever used this phrase about anything in Croatia
https://www.google.com/search?q=%22reminds+me+of+Koh-Lanta%2...
Translating "reminds me of Koh Lanta" to Croatian doesn't help
https://www.google.com/search?q=%22podsje%C4%87a+me+na+Koh-L...
I wish Google would just penalize Quora. I can't ask any question on Google without seeing a bunch of Quora fake experts on the first page providing vacant answers to any question.
If you compare 'new' vs what goes on the wall, you'll see a non-trivial amount of Big Tech bias and very good posts (often criticisms) effectively getting censored. Is the paycheck that good?
Yea, it's more monotonous, but it's far less hate filled than most forums and keeps on target far more often.
I browse new occasionally and I encounter a lot of lowbrow substack spam to be honest. Rants against big tech make you a contrarian, problem is you have to be contrarian and right to be interesting.
I don't think HN has a big tech bias, most criticism of big tech is just plain awful. If anything in today's discourse there is a stupendous anti big-anything bias (for context, I'm not paid by big tech, I work for a small German company on the other side of the pond and find much of SV culture annoying personally)
Any other website, where people just share information freely, where you might get different opinions, experiences, hell, _anything that is different from just trying to sell you something_ (even if it is ads), will be very far in the back.
In identifying a 95/5 reward for supporting "mass-appeal queries", is much of the answer right there?
And will the average general-purpose alternative search engine escape similar incentives?
It's amazing how much stuff/information there used to be, instead of whatever it is we have today on the WWW.
https://semiosis.github.io/about/
https://semiosis.github.io/posts/imaginary-internet-survival...
Better start encoding your personal mythology into code right now.
Google is good at many things, but their leadership is questionable now, unlike in the past.
I would pay a subscription fee for a service that provided partity with Google and customer service that answers the phone.
No idea how good it is, I remember it was posted here a while ago
I sort of agree, I also sometimes use Yandex and am usually positively surprised.
I can immediately see two reasons:
A) SEO efforts are targeted Google/Bing, not other search engines that don't focus on the target market/language/region. Of course that argument can be twisted around: these search engines don't focus on Western languages, so could in theory be more susceptible to SEO trickery. It's interesting that they aren't.
B) Ad money. Google is strongly incentivized to prefer search results that show Google ads. For foreign search engines there is no money to be made, so they can disregard that signal.
A bazaar of boutique search indexes provided by People seems like it would be nice to have. Better than Google constantly trying to sell me shit.
[1] https://paste.sr.ht/~sircmpwn/048293268d4ed4254659c3cd6abe67...
can't say whether this is true for sites but they definitely encourage this behavior in ads -- the ads dashboard will shame you ('poor quality ad! we may not show this' kind of prompts) if you don't obey silly rules about 'how many keywords' appear in your title.
I'm not an expert in this area but that seems very optimistic. Creating an algorithm that gives better quality results than google for certain kinds of queries probably isn't too difficult. But scaling it up to indexing the entire web, continuously, would require a LOT of resources.
The prior barrier if entry for quality content was that you had to be tech savvy. Now anyone and their mom can publish shit content, play the SEO game, and spam dunk the Internet.
Maybe the solution to this is really some kind fo web3 variant which Wil put up some hurdles again.
No they aren't.
In the old days, you might search something simple and it would have the answer in the listing. So you wouldn’t have to click. They are now gone too.
The patent expired a while ago now.
"They paved paradise to put up a parking lot"
Possible catalysts include: regulation of search or advertising negatively affecting google, an unknown search competitor entering the market, another of the behemoths competing (amazon with search, bing getting their act together, etc.), antitrust suits, insert black swan event here.
I think the pendulum is also swinging towards smaller, local communities. Maybe that's just my bubble. Eventually everyone will get sick of it though, and something will change. I don't think we can predict what will replace it, but I hope it's not another big, centralized source. We've seen the negatives of that.
I don't know how to build it, but to me the ideal Google search killer would be:
1) Decentralized somehow. Hosting search indexes collectively in order to reduce the need for a single entity to host the data. Crypto has this possibility maybe. But I don't know how well it could be implemented, and whether you could have search still be "free" like Google is. How do you solve for that problem?
2) Filters out the SEO spam. Yes, this is a huge problem. My idea to fix it is manual curation. Not scalable. How do you solve that? I think something like StumbleUpon. A curated list of sites that is searchable in its own index. Perhaps there could be a trusted network of curated indexes. Members only to host the curated indexes to keep out the black hats, but free to search for the public at large.
I also think we need to educate everyday users to host their own sites on their own hardware. Making it super simple to spin up a web host in a VM that is running on a laptop in the bedroom. Or off a raspberry pi 0. Or whatever is cheap and available. Turn it off when you need to and the website goes offline. Yes, but that is okay in the interim to get Joe Regular hosting a website independently. Sure it's not okay for a business, but for your hobby site about bonsai trees, why not? Of course this is against most ISP's ToS so how do you solve for that? How do you give Joe Regular the ability to deploy a simple, secure website easily? How do you incentivize them to get out of the walled gardens? Give them a garden of their own.
That was what made the early web (as I knew it in the 90s and 00s) awesome. You were figuring it out on your own. Google/Bing/DuckDuckGo can stick around and be the digital yellow pages. They are good at that.
We need to build a new web for ourselves.
For 1) we have an open app platform that allows everybody to collaborate on that first entry of the web
For 2) We let people decide the sources they want to see and if they don't like a source or app, they can downvote it and see it less in their ranker.
We will try to decentralize more over time though it's a big effort to make it fast enough still.
Exactly right. Big opportunity.
Optimistic. Which are those?
So the top hundred sites link to each other and everything else gets ignored.
This the equivalent of social media's echo chambers.
In the next 5 years, "Knowledge Networks" are going to become curated high-density graphs of domain-specific knowledge. Some will be paywalled, some will have ads, but you'll know how to find the information you're looking for.
Psychologically, at the core of the problem, this is what happens:
Cancel culture fear => Peer pressure hypersensibility => Monotonicity