Search engines and SEO spam
twitter.com
twitter.com
Google no longer producing high quality search results in significant categories - https://news.ycombinator.com/item?id=29772136 - Jan 2022 (1167 comments, spread over multiple pages - note the "X more comments" links at the bottom)
I guess what I'm saying is that if you want better reviews, you probably want to start writing reviews and figuring out how to sell them for money. Many have tried, few have succeeded. But there probably isn't some Javascript that will fix this problem.
That's not strictly true, given that reviewers are often sent pre-release versions of things in order to do that work before release day.
Sites that receive free review samples and are supported by affiliate links are kind of the exact opposite model.
Only purchasing review units at retail would remove this conflict.
Industry has already done this with the "food pyramid" - influencing, capturing governments to make the food pyramid more based on economic reasons and much less on science - with the government putting it out and distributing it into schooling of different levels, giving it an unearned or undeserved authority which then people blindly trust/follow - not understanding that or when systems and their output or oversight have been captured; why the pandemic bringing the classroom home via Zoom, so parents could see/hear the learning material has outraged many parents - an example I've heard, where white children are being taught to feel guilty about their 'white privilege', or parents being upset their children are being taught at a very young age that they can decide what gender they are; I'm not stating what I believe here, just giving examples I've heard of.
This capturing of the government is why I think ultimately the government should be developing and maintaining such platforms, as per law, and requiring individuals and organizations to in real-time add and update their data (simply example being restaurants, their menu's ingredients, their open hours) - in part to de-risk the government having an unnatural power as "the single arbiter" of truth, perhaps instead to de-risk capture that the government funds multiple independent organizations at the federal level - that States can decide which ones they follow, if necessary, part of why States exist - to de-risk the potential capture of the Federal umbrella; however the system is in an imbalanced, broken, captured state - with the duopoly evolving to be more extreme lead or formed by the establishment, with a broken voting system in arguably most countries of the world, and mainstream media being captured by for-profit industrial complexes that fund MSM through ad revenues - which further develop or mould our culture and narratives/talking points and beliefs, whether truthful or not; without fixing these the other platforms/systems excelling won't be possible.
What's happened since then is that almost all the normal "people linking to things they like" has gone behind walled gardens (chiefly Facebook), and vast majority of what remains on the open web are SEO spammers.
Fewer unique blog domains due to “blogging” sites that aggregate users? Sounds plausible. Fewer people blogging overall? I’m not convinced yet.
Ten years ago, the majority(!) had at least something up and running, where they would post essays, thoughts, whatever came to mind.
Nowadays? All gone. All! When asked why, the answer almost always is along a mix of ever-increasing negative feedback and harassment from randos, and aggressive automated spamming of their forums. Loss of the pseudo-anonymity plays a large role as well. Many have deleted years' worth of work, simply because they are afraid of someone trolling through their posts to find something to harass them with.
I was never a blogger myself, but I am sad about the change. There was a lot of good stuff out there for a while, and sometimes it just plain made me happy to read someone joyfully nerding out on a favorite subject of theirs.
Not to zero. You can still find things tucked away in a post on reddit or the like. Almost never, as far as I have experienced, on Facebook or its ilk, as the affordances are different. I genuinely think there has been a loss.
Now?
Nope. Putting anything out there is basically just doing the rest of the world's Open Source Intel for them. Maybe it isn't the Net that changed. It's just there's way more sharks out there that can't just leave well enough alone.
Most of these many more people are mobile users, where creating long-style text content can be quite bothersome.
What ain't bothersome, with a smartphone, is taking pictures and videos to slap filters over them, alas that's why we are where we are with TicToc, Instagram and Twitter dominating large parts of the web.
It's even noticeable in a lot of online discussions with text outside of these communities; The average length of forum posts feels like it's gotten way shorter over the decades. People have less attention to read anything that looks longer than a few sentences, often declaring it a "wall of text" based on quantity of text alone.
Imho it's a big part of what drives misinformation; Doing any kind of online research on a small phone screen is extremely bothersome compared to the workspace an actual computer/laptop, particularly with multi-monitor, gives.
There's also the difference in attention; When I sit down at my laptop/desktop, I actively decide to spend and focus my attention on that task and device.
While smartphone usage is mostly dominated by short bursts of "can't do anything else right now", I don't chose to take out my phone and surf the web, it's something I do when I'm stuck in some place with nothing else to do and no access to an actual computer.
But for the majority of web-users [0], that smartphone access to the web is all they know, which then ends up heavily shaping the ways they consume and contribute to it.
[0] https://techjury.net/blog/what-percentage-of-internet-traffi...
For my part, I'm glad these fora aren't indexed well; I don't want my search results dominated by single-sentence posts and photos. In particular, I don't have accounts on any of these services.
I'd be happy if search engines would decline to index sites behind paywalls. Links to Medium, Substack and Washpo are very common, and if the first thing I see is a popup demand for payment, that browser-tab gets closed.
Don’t know how hard it would be to know which is which. Maybe non-commercial : don’t run ads, don’t sell a product or service and provide information only.
Younger people TikTok, they Instagram, they chat in private conversations with eachother, they occasionally post short messages in walled gardens like Facebook, they YouTube, they listen to music, they watch Netflix & Co. That's what they do. They do not persistently write LiveJournals, Tumblrs, blogs. That pre video/audio-focused era is over and it's not coming back (even if there's occasionally a bubbling up of hipster fakery centered around how cool it is to write text).
I'm somewhat skeptical, it seems a little too poetic to blame Google's ultimate downfall on a decision that was notably hated at the time. But it's plausible. If you want it to be a conspiracy theory, you can posit that killing off independent blogs was the intent, to convince bloggers to migrate to Google Plus.
So everyone worried about SEO became afraid to link to anything except:
1) Their own website 2) High reputation sites like NYTimes, etc.
It's sad. Makes it harder to navigate the web.
People's best chance is stopping using Google and pushing for it to be broken-up.
The incentives to game the algo remain. People adapt to the environment.
I'm pretty sure google could do strictly better (i.e.: better in all reasonable accounts) than they do now if they focused on the users' experience instead of revenue for a couple terms.
Perhaps it could work if the algorithm changed its algorithm all the time.
-- Goodhart's Law.
Google's algorithms didn't create this situation; people chasing high Google rankings did. Had Google used completely different algorithms yet became equally dominant, people still would have poured their hearts and souls into getting higher rankings.
Basically, an application of the tragedy of the commons. Or: "why we can't have nice things".
Then Google came along and we all found it a lot more convenient than the bad search engines we were used to. And of course, we all know where that led. In some sense, Google built an 8-lane superhighway and bypassed all the small towns.
We all traded away paradise in exchange for convenience. Now we have neither.
It's just fragmented - i.e., catering to a specific group. Because if it isn't, it's awesome for 5 minutes and then monetization rot sets in.
[1] None of these work for everyone; conversely, all of these are seen as great things by some and have people who prefer that one thing over others for its quality.
But lowkey Google incentivized such behaviour by not being open and transparent on how exactly their algorithms work.
It's also dangerous to ask for the exact criteria because they are ever changing. Google et al don't want to be prescriptive about what a good site is, they want to recognize what a good one is. You make a good one, they'll figure out how to recognize it.
They can't sit down and publish "The Definitive Guide to a Good Website". That's just not their role and it will be out of date before it's published.
You're technically right. You'd be more right if you said people chased the highest spots on search engines for the widest breadth of queries.
If there were implicit alphabetical ordering of search results I guarantee you'd end up a bias toward A's, Z's or otherwise in people trying to get top spots.
Some would even say it killed the web by centralizing all the content in the hands of a few [0].
Which is the direct consequence of everybody optimizing to better show up on Google/Facebook/Amazon/Microsoft and ultimately even migrating all their hosting to these companies.
[0] https://staltz.com/the-web-began-dying-in-2014-heres-how.htm...
Knowing the actual rules might give them a fighting chance, since the bad guys already know these rules anyway.
Also if you are going to make a purchase somewhere, any website would try to get a cut of the money you spend by actually sending referral links to the product. So small websites that do not allow this service will not get linked so much.
On a metalevel it is thus that links or connections between items are information. Information is money. And as soon as that became evident links and connections also became more scarce.
That is really sad. Metric folks inventing metrics for the sake of metrics, which dubiously correlates to profitability of the company.
The algorithm made linking valueable. So instead of writing hypertext people tried to create isolated sites and boost their rank. Remember rel="nofollow"?
Eventually bigger sites took over the small ones.
But Google doesn’t run on PageRank anymore. PageRank is merely one of hundreds of signals they use to sort results.
It’s not a problem of doing the review, it’s that there’s not much of a market for written reviews, most people would rather watch a video instead.
There’s top channels like LTT and if what you are looking for is out of their niche, you look for the biggest channels in that niche and go mostly by association (who they have made collaborations with,..).
EDIT: of course the big win of video reviews is that you can see the thing working.
I'd say it's more that YouTube offers a clearer path to content monetization than text does. YT is a much more lucrative platform for the same level of effort as SEO for their text blogs.
Everyone just makes blogspam because its far less work than actually buying products and developing expertise and testing them and writing out a whole thorough review. Google's algorithms just can't tell a quality review from a surface-level, uneducated take.
1. You need to have enthusiastic reviewers (people who care enough about a product category to review them semi-throughly.)
2. Proper reviews can take time and may need domain knowledge.
3. Competition. When there were one or two people doing reviews on some category of products, maybe the economics worked out. Once you have hundreds or thousands competing with you, the time demand may be overwhelming and not worth it.
4. If you are a trusted reviewer or site, you will get economic pressure to review a particular thing or brand you may not like very much but the money may be good. So you will begin to experience conflicts of interest.
5. If reviews are just a hobby and not a way to make money, eventually you will slow down or move on, opening a hole that gets filled up by spammers.
7. Some things are timeless (a pipewrench, let's say) and some are seasonal (consumer electronics, toys, etc). The former deserves a through review but the latter doesn't deserve as much but it may get the bulk of interest due to seasonal demand). Does it really matter if the latter's latest iteration has 2% increased battery life to discuss?
I'm sure there is a lot I didn't think of. But it's a doomed category, unless people are willing to pay for professional reviews (Consumer Reports types and other independents).
In my experience, the best reviewers are hobbyists. The thing is, it's not reviews that are their hobby. Rather, they review the products go along with their hobby.
So, for my hobbies (espresso and aquariums), there are tons of easily accessible reviews on all kinds of aquarium gear and coffee machines, grinders, etc. On the other hand, nobody does plumbing or HVAC as a hobby (that I know of) so it's very difficult to find high quality reviews of water softeners or furnaces. It takes a very special rare sort of person who would install these things just to review them. The closest thing I could find was this video [1] on a DIY water filtration system by an RV/off the grid type hobbyist (from what I can tell).
Review sites suffer from a singular problem. They are overwhelmingly SEO spam content farms. People go find some product niche and pay some Fivvver/whatever people to write literally fake reviews of products. Because they're pulling all the SEO tricks and are in a niche category they shoot to the top of search results for that niche.
Their reviews sound realistic and viable but they're pure fantasy. The writers never touch the products being reviewed. Many times they'll pull details from Amazon listings (including factual errors) and even other "review" sites.
Once they get established in their niche they'll accept paid placement from product manufacturers without marking it as such. A single scammer might own dozens of these sites, even supposedly competing ones.
We regularly see partnership opportunities with customers interested in our API [0]. I presume Bing see the same, though their terms are more fixed and require you sharing more data. Definitely big opportunities in other categories, which are often squandered through a naive, if understandable route, of choosing a Scrape and index route.
One possible solution to this could be:
- Let the community vote on the most trusted sources
- Include results from enthusiasts that have little incentive to write biased reviews (Reddit, HN, expert forums)
- Look at the ownership of the site and how transparent they are about it
- Regularly reassess these criteria
This wouldn't scale for a generic search engine, but I'm working on a service that does this for many product verticals/niches.
Not sure if this is a good model for a search engine, but it does work to a small degree in those forums.
Internet points are a terrible reason to write anything. They're completely meaningless. We should all judge comments on their own merit and not because the author has a lot of karma. Apart from mine, obvs.
Instead of creating income opportunity and crowdsourcing people in foreign countries for common (more positive) good, companies rarely create opportunities for the people who would normally turn into spammers and scammers, and that's what creates an endless army of people that constantly destroy online communities like Soundcloud, FaceBook, Twitter, and TikTok with stolen content, trend scams, fake news, and spam messages.
Google search has been invalidating and subverting their most accurate search results based on abstract SEO rules for quite some time now. It was likely done so that they could implant paid ads first into content, because that makes them the most profit. Doing that has destroyed their reliability and reputation as a search service leader, and they're never going to admit it, but payola is the undertone that is ruining their search results... There is a certain type of corruption that occurs when a company turns away from upholding customer service and value towards a monopolistic "profit-first economic stranglehold" business model... That strategy never ultimately works out well for both companies AND users in the long run. The next leader will likely be a search that avoids the same pitfalls until they themselves become a profit-driven monopoly.
There is no algorithm that will usefully and fairly counter spam based on desperation, companies need to realize that creating opportunity for people to operate equally on their platforms is the best move, otherwise, spam will drive any community of rule abiding users away or into madness.
Not quite right because cybercrime aka hacking, cracking, spamming etc. originated in US not in East Europe, Russia and third world countries which are dominating hacking and spamming scene today. Main motivation of cybercriminals is quick money and ease of getting away with it since you are not physically committing a crime but digitally/electronically.
Hacking heavily originated in the US because the US practically built the entire modern tech universe from the ground up. The US was far out in front when it came to utilizing the Internet and the Web, so of course unethical people in the US pioneered various types of online crime, the US was the early adopter.
If you're an elite engineer in the US, you can make millions of dollars doing legal work for big tech. It helps in a big way to drain the labor pool as it pertains to criminal activity online. You generally can't do that today in the countries that dominate SEO spam, online scams, etc. In those countries elite engineers suffer terrible wages doing legal work compared to what they should be able to earn for their abilities; commonly they can earn a lot more doing illegal work instead, it's a very potent lure.
You're an elite engineer in Russia, top ~1%-3% globally. What do you do? Earn several thousand dollars per month doing legit software development in Russia (with either zero or little consequential equity compensation); flee Russia for a more affluent market; or do illegal work where the rewards can be dramatically greater. It would be difficult to resist if you were unable or unwilling to leave Russia.
Become software entrepreneur?
And many international software companies have software development teams and presence in Russia.
Exactly. Hacking for hire, making cheats, botnets, SEO farms, selling exploits and hacked social media accounts; practically anything you can think of that US software engineers can't be bothered with, as they already earn a healthy salary. That is entrepreneurism.
How many? 20? 30? 50? IMHO the cases are rare (and get widely publicized whenever that happens, creating a disproportional visibility), you get a couple captures per year but the number is just a tiny fraction of the actual participants, more like an exception than the rule.
SEO is perfectly legal. So is spamming, regrettably.
I think Goo no longer cares about the quality of search results; they have other business priorities, so SEO works. Spam is another thing again; we still, after 30 years, don't have an agreed definition of spam. We still don't have a flawless spam filter - far from it. So it astonishes me how much email spam I get promoting SEO services for my non-existent website.
SEO is legal but spam is not at least not in US and many other jurisdictions.
>I think Goo no longer cares about the quality of search results; they have other business priorities, so SEO works.
Google cares about spam but there is so much data and information on the web that it is impossible to figure out what is spam and what is not. Another big problem is fake data and information that is also very hard to figure out. Generally Google prefers popularity over quality because it is easier to detect what is popular than what is of good quality.
>Spam is another thing again; we still, after 30 years, don't have an agreed definition of spam. We still don't have a flawless spam filter - far from it.
Definition of spam is unsolicited message. So if I get pharma emails in my email inbox that I didn't ask for it is considered spam. SMS messages are another example for example if I get SMS message promoting free coupons but I didn't ask for it then it is spam.
Considering how much spam Google saw in the last 20 years both on the web and on the Gmail they should have some decent machine learning/AI algorithms which could flag spam pretty efficiently.
That may be your definition. As I observed, not everyone agrees with it. Most definitions of spam include the word "bulk", for example.
Russia also has its own SaaS enterprise sector with companies like SKB Kontur or Diasoft.
Just like the US has to this day warez and cracking groups, where it's for the longest time mostly been about scene prestige, and not making the big bucks.
If you can't get a work visa to a first world country, you do have less options than someone already living there; and the salaries offered by first-world "international software companies" in their remote subsidiaries tend to be 'according to local market rates' (the same "several thousand dollars per month" mentioned by the parent poster is a decent rate) and thus not as competitive with "black entrepreneurship" which pays according to global standards.
I wonder how a search engine could distinguish between "honest & professional" and "fake & amateur" headset reviews without having a head and two ears?
Modern Google actually makes the content problem worse. When our notional TV blogger is starting out in our world he publishes two or three essays, nobody reads them, he stops putting in do much effort, posts occasionally, dwindles off. In a world with a perfect search engine his early essays get some attention to encourage him to post more, a feedback loop starts, and before you know it he's a full time TV reviewer.
I remember doing the exact same thing in the past and being overwhelmed with information. The detail and data in reviews would take a long time to collate and make sense of. But this time even the big name sites seem to be much shallower. Less models reviewed, less testing and benchmarking, more regurgitated press releases and other news.
Last time it took me a while to sort out all of the information, this time all my time was spent trying to find any that wasn't 100% fluff.
I think in practice this is actually largely untrue -- with technology products, video games, movies, and just about anything I can think of, most well known reviewers are given early access to the product so that reviews can come out on or before day 1 of general availability. That said this does create a dirth of 100% trustworthy reviews on day 1 since companies are naturally disincentivized from giving early access to reviewers who they know are going to write a negative review.
But on top of that you do have the problem of whether or not someone is really qualified to write a review. So Joe User thinks product X is good. What is their metric for good? It reminds me of an LTT review of the Amazon TV from a few months ago. They gave it an awful rating but noted that the reviews on the product page were generally very positive. And their reasoning was that the people buying these TVs and reviewing them didn't have a good comparison point for what a good TV actually is. They are probably comparing it to a much older and less advanced product not to a contemporary one.
So then you think the answer must be get reviews from industry related media. But then you fall into the classic problems of unethical journalists or simply ones that are out of touch.
Professional critics usually try to distinguish themselves by producing well written in-depth reviews but not from their own perspective but that of a hypothetical everyman who, ideally, is similar enough to a critical mass of their audience.
So it always interests me when people complain about popular gaming review sites being out of touch because almost always it's the reader that's out of touch but doesn't realize their bubble. It's not an absolute rule but I'm in enough niche hobbies to realize that my desires for products are way out of whack.
It’s a bit funny because this we sort of done by Jason Calcanis’ Mahalo back in the day - but maybe he was just ahead of the SEO curve.
"Obvious, unactionable, or wrong" is a term of art created specifically to talk about Calacanis' advice.
"Mahalo", a Hawaiian term loosely translated as "who is this freak and why does he think our language is there to promote his shitty startup?" was the answer to the single most terrible thing VCs could identify at the time: that someone would create Wikipedia without a plan to monetize every aspect of it. And yet Calacanis did a worse job than the "original fool", Jimmy Wales, whose similar attempt at least didn't end up as a parked domain of SEO spammers.
Instead of having 1 search engine which returns the same results for everyone (depending on interests, etc, like Google does), we could have trust networks. E.g. you trust a few people, those people trust other people. From this network you could build something like PageRank, which computes some kind of transitive closure of trust for one given person. This will then determine the search ordering for that person.
> site:reddit.com <QUERY>
If I want to know about headphones, or TVs, I'll find better answers in the sidebar than anywhere else on the Internet. But, I will not be able to find those quality answers until a reasonable period of time has passed where products can be tried and reviewed by real people.
The issue is the immediacy. We want answers now, but we won't ever have them until later. This requires a cultural shift in consumerism that companies will not like: "Wait and see."
The same problem happens for people pre-ordering video games that end up releasing in a state of complete trash. You can only get reliable information once the early adopters have tested it. If you are an early adopter, you are shit out of luck, but you are doing the rest of us a service that we appreciate greatly.
They're in Montreal. They buy their products retail. They get their funding through subscriptions and affiliate links. Overall ratings are formulaic based on measurements applicable to the category. The formulas are available on the site. So are all the test designs.
They're surprisingly thorough too--when I was shopping for some over-ear headphones I really appreciated that they measured the clamping force (when you wear glasses, something that clamps onto the arms too hard gets really uncomfortable pretty quickly) and breathability (temperature differential between your ears and ambient when wearing the headphones). These are pretty important for all-day comfort but don't really factor in most of the time.
Their methodology is all available:
Breathability: https://www.rtings.com/headphones/tests/design/breathability Clamping Force: https://www.rtings.com/headphones/tests/design/comfort#compa...
Back in the olden days, there were lots of organizations that collated high quality content from the best writers. They nurtured expert writers and paid them well. They fact-checked the content and employed diligent editors and proofreaders so it was accurate and well-written. Over the years, they'd build a reputation for reliability and trustworthiness that kept people coming back for more. If you wanted to learn about fitness, or cars, or cooking, or science, you'd find a reputable author and publisher and buy their magazines or books.
But then, in the early 2000s, the geniuses from SV "disrupted" the publishing industry and its financial model. They brought us a much better way to find content, the search engine. Because they were so much better than the old-fashioned publishers, search engines gobbled up the advertising money and became the dominant gateway to content. Publishers had to abandon expensive high-quality writing because rankings and eyeballs now mattered more than quality and trustworthiness. Instead of investing in writers, they invested in marketers and SEO specialists.
The result: worthless content, writers banging out garbage for peanuts, and useless search engines.
Two decades later, looking at the barren wasteland they had created, the SV geniuses thought: I know what we need, more search engines, but smaller ones that collate high-quality content from the best writers. There must be money in that, right?
This implies that big budget advertisers (the CPGs, like Coke and P&G), are buying Google/FB because they have better conversion metrics. That isn't true today; only SMBs and gaming companies care about conversion metrics. There are interns in LA/NY probably collectively spending millions on FB for P&G and only reporting the number of likes back to their bosses. Google and FB has never meaningfully delivered on conversions past anything like app downloads.
Tech undermined publisher's revenue because the internet cratered distribution costs. Advertising revenues for big media crashed because the eyeballs moved away, not because it was any less efficient.
Did any of these heavily buy newspaper ads before 2000? Definitely TV, possibly magazine, but newspaper? I just don't remember seeing ads for Tide in newspapers.
I stand by my point that tech just exposed a bad business model. Newspapers were only viable because they were the ~only game in town.
Now what actually killed newspapers wasn't search, it was Craigslist killing the classifieds...at least according to a study Google paid for.
It went from "search engines and the web will usher in a new era of wisdom and democracy" to "useful content is dying at the hands of monetization schemes, and also the internet will be the death of liberal democracy, woe unto us all" in about 15 years.
So from an individual developer perspective, it's very hard to get people to change their habits. And Google/duck/Bing is the one stop shop.
It's still out there, but I haven't worked on it much lately. I always think that if I had some good advertisers, a better UI, and a salary coming in, maybe it could take over some of Google's usage!
(2) Do you have a bang on DuckDuckGo? I'm pretty aggressive with bangs, and I suspect a lot of DDG users end up being aggressive with them as well.
Link please
I like and use Duck Duck Go‘s !bangs [1] all the time, maybe try to add your site with a rememberable name.. may I suggest !garlic ?
[1] - https://duckduckgo.com/bang [2] - https://duckduckgo.com/newbang
One piece of feedback: It seems to somewhat liberally fuzz the term, doesn't tell me about it and not toggle it in the UI. I searched for "Natto" and got some great results at the top (Spicy Kimchi Natto!). However, a few results down the recipes start including "Nattu" which seems to be a Indian chicken dish and then quickly have nothing to do with the term at all and give me stuff like Tafelspitz.
Tweaking search can be a whole full time thing. I'll write it down and maybe look into if I can tweak the search so if there are exact matches to not show alternatives.
1 thing that would be helpful is to add in the ingredient amounts.
Never quite got around to it unfortunately, as I wanted more users before new features.
Of course, the massive volume of content creates a fundamental problem, but user curation & categorization on sites like Youtube would be possible, were Google to provide the software support so people could do that. Whether this and similar decisions are deliberate or accidental is likely one of those things that we will never know.
I've noticed this, and it's frustrating. I have assumed it's intentional. I am left to guess as to what a change in this behavior would accomplish.
The invariant has always been: find people who make falsifiable predictions and improve. Back then the pool was small and you had no choice. Now, fortunately we have a choice.
No, what happened is that the publishing industry lost their monopoly. They could no longer extract monopoly rents from advertisers.
One problem with search engines is affiliate marketing. If Google de-indexed the junk affiliate sites, the web would be much less polluted with affiliate spam.
Publishing was not a monopoly. Google/FB are a duopoly. If publishers capture rent from advertisers, they plow it back into content, aka the thing consumers actually want. If Google/FB capture rent, they don’t provide a living wage to content creators and plow the money into buying other startups and whatever “metaverse” is.
Popular HN post in 2022: “Search engines and SEO spam”.
Inevitable popular HN post 2025: “How to avoid getting flagged when content marketing“.
Often times I find myself searching for "best ($product|$thing_to_do)" which I think many other people do as well because we all want the best. Other times I'm looking for a music or a book recommendation with some depth. This of course nearly always leads to SEOd trash. There is no relevance nor is there trust. So, I like others to use keywords like "reddit" or "forum" to get to real humans who I trust and intentions are not to sell via affiliate links.
These issues often lead to the need in finding trust in real human-centered recommendations that stem from real human interests and needs. I've never found an algorithmic solution to this problem. This is why I think college radio stations or those south-of-the-dial end up being so, so much better. And why beer recommendations from your local brew-shop owner are better than anything you can find on the net.
I think building search vertical that are hand-curated would be very interesting to see. But I also think we need to build more communities which allow recommendations to be shared without an incentive to get hits via search and aren't paid for by large corporations and where community impact/quality _is_ incentivized. I do worry that those days may be gone and there are just not may be enough folks (not in tech) willing to spend so much time online and contributing to niche communities. A lot of folks spend much of their time in walled-gardens like Facebook, Instagram or Twitter, so it'll be challenging to be sure.
The larger the gap between paid results vs organic results, the more users click the paid results.
Not sure how to solve this problem.
So this would end up with displeased users and bounce backs.
I do too. I'm wondering why didn't Google invest some effort into "best X" searches? I bet they could extract such information from the web and correlate various sources. They already answer all sorts of semantic knowledge questions.
That was my inspiration behind a side project I made a few years ago — a decentralized, hand curated "search engine" [0]. Never got beyond the side project stage. But I see promise in this in the future. Eventually we'll figure out that moderated crowd-sourced curation is better than the best machine learning. The filtering capabilities have to be pretty sophisticated to make it work, though.
And therein lies the problem. Reddit makes very little money. Forums probably make negative money nowadays. Google has decided to demonetize the organic internet and subsidizes SEO crap and AMP or whatever dumb thing their signals consider valuable. We get what we incentivize, and right now the incentives in almost all of tech are pretty atrocious.
>You might need to do a lot of manual spam fighting initially. That could be both the thing-that-doesn't-scale, and the thing that differentiates you by being alien to Google's DNA. (They must hate manual interventions; so inelegant).
Is he describing...Yahoo circa 1994? A manually curated directory service.
The gamification of the system then would have to come through onboarding fake users, pretending/mimicking real user behaviour to send that signal into the system; not sure if SU ever ran into that problem or was actively paying attention to trying to identify and removing fake or suspicious signals from their output?
I feel a much better system is easily within reach, it's simply getting the right structure to it, the right foundation, and then it will quickly take off due to the quality difference. I've already figured out a design pattern that Twitter and Facebook has indoctrinated us with, making us think it is normal - and keeping us blind to an actual normal way or organizing or communicating, but that isn't conducive to control or ad revenues - and so extending my future plans to include a better search-directory system would fit snugly into my efforts.
For health information and recipes in particular there are only a handful of really high quality sites that have quality content for 95% of the information most people need. I bet if you wanted to increase the coverage to 99%, that list would expand to less than a thousand sites. At those numbers manually curating the information would be easily achievable.
How to get people to use your top notch Google replacement instead of Google, however. That's the hard problem.
Siebel and PG see blood in the water no doubt, they see G's market share and want to fund companies to take some of this.
https://cse.google.com/cse?cx=dc408db269da4e769 (try searching for something you want a review of)
Make a search, whitelist the domains. Every time you run into a good review site, add it to the searchable list.
I'm not saying that curated search results for particular verticals is a terrible idea (though I'm sure like anything the devil is in the details), but on the whole Google search is very, very good considering the constant assault they are under from spammers (which most other search engines are not, at least directly).
There is no scenario - none - where thousands of engineers at Google working on search wake up in the morning and say "we sure have made it good enough wr2 SPAM. I think I'll have another Danish."
I think it’s more likely that because they are just building hundreds of tiny tweak experiments and it’s someone else who desides what to build and if it even worked. Search quality is such a meta-problem that it goes beyond any real hope of simply working on it in anything beyond piecemeal trial and error fashion on their dataset.
I don't claim that Google has any of these and certainly have no insight into their search group. But I've personally been at powerful companies with best-of-the-best talent that were blind to the decay in their own living room, so I would caution against immediate dismissal of PG's take.
The hard part in all of this isn't finding and stopping spam - it's defining what spam is. Are all the pie recipes where there's a 2000 word essay about their grandma at the top 'spam'? They still have the recipe, and Google Home devices pick up the recipe instructions just fine so people end up not reading it, but many people would still consider that spam since it adds such an obstacle to getting the information you want. Same for cnet articles like "Best smart home devices to buy in 2022" - it's a reputable brand with a list of smart home devices, but it's hardly a review and exists to funnel people to their Amazon affiliate link.
What's the point of writing succinct, to-the-point mini articles about problems and solutions if nobody finds them on Google?
I've done blogging for the last 10+ years, and many of those I spent as a freelancer working with startups/brands/editorials. Everyone is after "word count" and I absolutely hate it.
Whenever I work on articles for my own blog, I just don't consider word-count at all. I think if your content is great and informative, then readership will be natural.
Later on, I sold it because I needed the money. Not so much that I didn't want to keep working on it. Unfortunately, the new owners didn't have any idea how to maintain a "healthy" content blog, and it has plummeted down to around 30,000 monthly visitors. All the content they're publishing now is some thin headline-clickbait bullshit.
I even gave them free advice on how to fix it, but I think that for a lot of people, they just don't care and will mindlessly pump out as many pieces of content as possible. And such blogs can be identified from a mile away.
And therein lies the problem with Google SEO at the moment. Even myself, someone who has done SEO work for more than a decade, I can see that results are getting worse. In some niches, the same crappy articles that dominated 6-7 years ago are still dominant today.
I guess we're stuck in time, or so Google thinks.
Incidentally, prioritizing long content seems odd to me, in my experience the best pages are short and get right to the point, at least in the context of something like a recipe or other "how to" resources.
Google's algos, while advanced, still rely a ton on text to actually tell what the page is about. They need it.
If they just relied on other factors (title, links, website, etc.) they would end up with worse results for users. Im sure they've tested it.
Google's core algo in a lot of ways is much simpler than people think (in other ways of course it's very complex).
(The fact that many of these books were rushed out with very poor quality control also didn't help.)
I think this isn't entirely related, but that's perhaps the beginning of a bias you might end up having that everyone experiences technology in the same way as it marches on. I've yet to encounter a Google Home in the wild, I imagine far more people are consuming recipes on phones, tablets and PCs.
Hmm why stop there let's actually make the users do the curating and even the content creation by rewarding them with social validation. Let’s have hard working moderators who work on the community full time.
Then we could just build a search engine over it. We could call it Reddit. Or HackerNews.
Maybe the users aren't all as good as professionals at curating the information. Let's hire professionally trained curators pay them well and we could call them newspapers. Then we can come in disrupt them and replace them with an algorithmic marketplace that eventually becomes infested with click bait.
This is one area where Google could use personalised results to provide a better experience for the user. Let me decide what spam is for me. Let me mark results as good or bad, so that the algorithm knows what kind of pages should be prioritised or filtered out the next time. Google SearchWiki was a step towards this but they killed it off.
We have seen what this leads to inside the social networks as well as YouTube, and at a macro scale I think we might want to have a shared concept of what constitutes a good search result for a given query.
At micro scale, it can seem more optimized to get exactly the type of result you want, but if we take an absurd example like an Apple Pie recipe shouldn't we all have shared understanding of what types of ingredients would make for an Apple Pie?
The shared understanding, I believe, is core to communication. If all of us have our own specific ideas of Apple Pie, then who is actually right on what an Apple Pie really is? What happens when your search results insist that an Apple Pie doesn't actually have apples in it, but instead pears?
At some point, Google must have moved away from using site-level reputation in search rankings, as I almost never see recipes from reputable sources like King Arthur Baking, Serious Eats, or Food52 in the first page of results.
Not only that, Paul and Michael have seen plenty of startups, and at least in recent memory, the number of vertical search and consumer startups that Y Combinator has funded hasn't been that high
As a consumer startup, I know this issue firsthand. Paul and Michael assume that if you build a better product, they will come! That's simply not true these days.
Instead, you need to:
- Build a better product
- Option 1: Figure out a channel with enough growth on an existing platform. This likely means you're doing SEO for your new search engine
- Option 2: Get your customer lifetime value high enough so you can pay for ads. This is tough, since it's a bit of a chicken and the egg problem since most search engines are monetized with ads
As the founder of Wanderlog (YC W19; https://wanderlog.com), a consumer vacation planning app [1], I definitely remember the idealistic days when I thought the best consumer product on its own would win! But growth doesn't just come, and the same can be said of vertical-specific search engines.
[1] Try searching "[your city] itinerary" on Google vs. Bing: it's much more likely you'll find a small blog rather than Lonely Planet or the local travel bureau as the top result
nonoonononooonono. No. Don't monetize anything for the first 10 years. That's the only way it can work. Then you can go monetize it and buy an island and not give a shit if you destroy what you created.
Oh but don't worry. You'll have investors.
Bing: https://i.judge.sh/ShareX/2022/01/www.bing.com_search_q%3Dat...
Google: https://i.judge.sh/ShareX/2022/01/www.google.com_search_q%3D...
Interestingly Google didn't have a top-result ad and the google.com/travel carousel is 4th from the bottom.
For the actual results, both thefearlessforeigner.com and paigemindsthegap.com seem to be actual travel blogs (the pictures didn't appear in a reverse image search, so they are probably organic), but they're clearly geared towards being a 'faq' for visiting the city and have affiliate links where appropriate. Bing went straight for discoveratlanta.com, and frommers.com is well-thought-out but not a personal travel blog.
The best part of it was (going to a foreign country) being able to find / identify all the attractions relative to each other, so I could go to cluster A on Monday, cluster B, on Tuesday, etc.
The hardest part of it (and why I needed to create a separate google sheets anyways) was--once I figured out opening hours of different locations, hard-to-book activities with limited reservations--the ease of moving things around more fluidly e.g. cluster B on Monday, cluster A on Tuesday, etc. and having a more information-dense view so I could see larger portions of the itinerary at once.
It would be cool to have an "input everything" --> "input time restrictions / unmovable things" --> output planned activity cluster type workflow.
It's like the bad-old-days of ExpertsExchange, which somehow was never delisted by Google for its shady SEO tactics.
As a user, your best personal and ethical move is to install an ad-blocker, to make ad-based business models less viable, which will help promote business models that don't abuse the customer.
This seems a bit overly cynical. Some search engines only served ads, but they're long gone. The survivors are those who dedicated themselves to finding links which were responsive to people's search intent. They seem to have gotten into ads because it was the best business model in this market.
Some even made slideshows of SO screen captures and put that on Youtube, with a fake video or spoken intro to make believe an actual content will be discussed... A number of shameless people would go any length to grab bits of money anywhere and anyhow, and I've hit those links a couple of times.
This is really what made me suspect that Google was teetering on the edge of the MBA death spiral: these problems run for years when they'd be easy to block, which suggests to me that whatever metric gets you a bonus / promoted doesn't include things like that which are long-term threats to their core business even if it's selling a lot of ads short-term.
The only advantage a startup might have is that they could do completely new concepts, such as specifying what area you search in, allow you to modify their classification of your query and/or moderating sites you include - which is probably necessary anyway, since you'll hardly have the budget to fully index the web. I'm not saying it's impossible, but it's not going to be easy at all.
And after all of that, you still need a way to make some money.
The reality is more that some Google engineer will come up with an algorithm change that makes the result 40% better, but it will come at the expense of making that search 3ms slower so the change won't get merged. Or it will make the results worse for some niche set of queries that the business team really cares about, so again it won't get merged.
There are lots of consumers who would gladly pay $1 a month or whatever in order to use a couple extra milliseconds of compute power per per search in exchange for drastically better results, so there is lots of room for a startup to compete.
Google has a paid-for Search API, so they could do that if they chose to pursue it. And then they could let Google One users opt-in to the same thing via ordinary Search. I'm not sure whether Bing has anything equivalent.
This is hyperbolic, right? But they can solve spam in a split second, if they just admit they're watching you all the time.
[edit] /s thx for reading to the end, folks.
To put it simply, a new generation of the people who used to make the reliable niche websites that not just answered your questions but also helped you learn a particular topic have moved to youtube instead.
Google search is hollowing out as a result with the meat going and the SEO'd fluff that kinda answers the question but ONLY the direct question being asked with none of the wider expertise that more educated people in what they were searching for.
Of course google owns youtube as well.. so perhaps they just see it as an inevitable transition.
It takes time and effort to build up a spam site's ranking but it is trivial to blacklist those who get to the top.
The complaints I've read are from exactly the kind of generated content farms people are complaining about in this thread.
Really?
I can point you to Hard Problems that have been solved better at little startups than at Google - or, indeed, at any other bigco. That's why acquisitions happen.
Why does Google having 1000 engineers working on a problem automatically mean they are the smartest?
It's that Google has destroyed their own search results in order to continue to expand their revenue opportunities.
If Google:
- Enabled downvoting on results, like YT videos. (Has its own spam problems, just like YT)
- Allowed you to block certain domains from your search results, like YT videos. (If they added some kind of "coordinated network detection" and down-ranked domains coordinating with ones you've blocked, that'd be pretty cool).
- Allowed you to create your own custom search engines, like "Programmable Search Engine".
That would be incredibly valuable. They already have most of the tech. They could even create a subscription service around custom search engines if they really wanted. Plenty of people would find something like that incredibly valuable.
Anyhow, buried in there is your startup idea. Remember: your startup doesn't have to generate the same revenue or profit as the incumbent on day one to be successful.
Are there any search engines that do this? It's a great, simple idea.
Couple of positions up or down in google results for somewhat popular and valuable keywords can mean the difference in thousands of dollars per day of ad or affiliate revenue. I suspect it would get pretty wild if google launched something like this. There already are black-hat seo methods and services, but something so simple and direct would turn it up to 11.
They already have the tech to fight this on YT. They, in theory, are supposed to be doing the same thing to detect inauthentic behavior on ad placement and click abuse.
Why would they do this? Google's customers are the advertisers, not the end-users. And no one is going to pay for a search engine, it's been tried and has failed.
A company with an objectively superior search engine could make even more money with ads so now you’re back to the beginning
An “objectively superior” search engine, from an end user’s perspective, might have to make engineering choices that come at the expense of ad revenue.
But we’re all just talking hypothesis, it’d be cool to see someone launch a startup to get some answers.
It's not open source as far I know and there's only the free trial way to try it.
Always curious about things like this. I certainly would pay for this; it sounds like many other people here as well would. I'm curious if the constraint is that there aren't enough people to actually pay for the investment required for the service, or if there aren't enough people willing to pay to meet the standard VC notions of success. We seem to have a problem with building and supplying services for niche (read: "not expressable as an integer percent of the world's population") customer bases, and I'm never sure if that's a business problem or a cultural problem.
Nobody is going to pay for /just/ a search engine. But they might pay for, say, a /better/ search engine, plus additional features around gmail/gcal/gdrive.
Think of it more as subscribing "to google" and less as subscribing to "google search".
Regardless, the point isn't to "fix" google. It's to highlight a possible path for a new market entrant.
... If an existing player wanted to make a move here, I would say that both Mozilla and Apple are well positioned to add "personalized search" to a subscription service. Same with Microsoft. DDG could also make moves here if they expanded beyond search.
Google could easily use its current fingerprinting to constrain (to an extent) multiple votes. Even knowing only a portion of the population will participate in the voting they can use a Wilson confidence interval[0] or similar to properly weight votes.
Random sampling works here since you're not guaranteed one vote per user per page and the outcome in binomial, seen and downvoted or seen and not downvoted.
[0] https://www.mikulskibartosz.name/wilson-score-in-python-exam...
Internal Search owners can push for better algos, but what if the algo causes revenue to fall? Are there strong forces strong enough within the organisation to ensure that search quality prevails?
If this is the case, the problem is existential. It can only be arrested at the very top
This gets close to the real root of the issue -- attention is monetizable independently of the quality of content. There would be much less incentive to create SEO spam if search engines negatively weighted pages with ads and affiliate links, and if manufacturers were barred (e.g. by the FTC) from owning or imitating reviewers.
This is also something that Google can control if competitors come along.
i.e. If a reasonable competitor comes along that is willing to sacrifice ad revenue for better search result quality than Google, google can just adjust their search quality upwards to knock them out (and then adjust it back once the competitive threat is gone).
Perverse incentives from Google are all over the place - Searching for the delivery business "Just Eat" in the UK for instance returns an ad for their competitor Deliveroo above the legitimate organic search result for me - and I can also see that JustEat are trying to pay for their own brand name just to compete - and IMO this sort of behaviour is anti-competitive, borderline extortion considering Google is the de-facto way of searching for a business, and shocking from a search-quality perspective (where the wrong result is intentionally shown at the top because they paid more money).
Things have changed in the past few years, now that Google has developed advanced transformer models, but for a long time Google's question answering facility has been: "let spammers make 10^8 pages where the title is the question and the answer is in the page".
The trouble is that there's a fine line between "answer is in the page" and "word salad!"
They're going the OPPOSITE DIRECTION from this!! They recently removed all downvote visibility of YouTube videos from the user, so now downvotes only feed into their algorithm. So in the last line of defense of me ending up watching a shitty video, one of the most valuable tools has been removed by my betters. It's preposterous that people think that Google is doing a good job. They're actively getting worse, and ignoring everyone saying so.
All the big channels I follow are mad at this change, and there's a coordinated effort to bring it back.
Google's search algorithm is tuned up for searching the whole web. It turns out the heuristics you need are very different depending on the size of the collection.
When Gerard Salton was doing IR experiments with punched cards he was working with collections of as little as 70 documents and in that case you are going to be very concerned about recall and not precision. Maybe there is 1 relevant document and if you miss it you failed.
If you had 70 billion documents you might have 10,000 relevant documents and if you lost 60% of them you still have 4,000 documents. The end user gets more results than they can sift through.
Thus I always groan when I see a site is using "Google Site Search" because the relevance is usually worse than you'd get with the alternatives.
Connected with that is the tuning work: Google has sufficient data to tune up a big model for everybody but true personalized search eludes them because they don't have enough data from you to tune up a model for you.
you mean the dislike counter they just disabled to force people to sit through more low quality content and pre-roll ads to claim increase in platform engagement and viewership?
The only thing matters is revenue and Google had increases in acquisition costs in prior revenue reports. Expect to see the data points for the latter metrics to be highlighted on the earnings announcement, and a record quarter for YT coming out of the change.
Not convinced this would help. The spammers would just hire people to dislike competitors
> - Allowed you to block certain domains from your search results
This I would use. Never show me results form collider, watchmojo, ranker,
> - Allowed you to create your own custom search engines, like "Programmable Search Engine".
I think this would lead to people writing highly polarized engines. The Red Pill engine for example and we'd have a new problem, the proliferation of popular highly biased results. Of course that's not to say Google's results aren't already biased but they certainly are trying to cover everyone.
I would love for Google to build this in. Until they do, there is a WebExtension that does this: https://addons.mozilla.org/en-US/firefox/addon/hohser/ ("Block or Highlight Search Engine Results"). I use it to block stuff like W3Schools so when I search for something, MDN is always #1. Saves me a lot of time having to add "MDN" to the end of every query.
Blocking pintrest would be a dream come true.
This is def a user-engagement strategy -- but it has cons as well.
Part of the complaints in the thread were spam related, other were something deeper
These sort of things are easy to spot, but only if you actually have a basic amount of familiarly with the topic. It's hard to spot with "AI" or super-cheap labor.
(Disclosure: I work at Google, but not on search)
I still think the model could work if the algorithm is sufficiently different than Google's. Ideally, people would go "I did not find anything I cared about on Google, I know, I'll use Bing!" - but nobody does this, because the results are consistently worse.
Don't get me wrong, I like G as a company, I think they do worthwhile things! But they have left things slip and need competition into this field, I mean real competition, then maybe they would actually address issues.
Maybe the issue is also on the incentive level as well. I mean more searches means more eyeballs and more money for Google. If someone searches one thing and they are done that is less interaction! I hope they don't work like this, but it's possible.
And another possible problem is the opposite. Maybe Google is optimizing search for what it thinks people want, but it uses the wrong metric. Or it gives people what they want but not what they need.
I think a really good search engine would still algorithmically search it's index, but the content library should be human-curated with a goal of ingesting content via author, not via platform. Once a given author was human-approved as a quality source of information, content they produce could be automatically ingested going forwards, and conditionally re-reviewed by a human if there were reports the quality had decreased.
I think the death of the directory search dramatically dropped the number of self-curated, informative sites from a domain expert that were common in the early internet. Now instead of making a website, many people are on content silos like Reddit/FB
Assuming a method also existed for an author to authenticate themselves with the search engine, one could also enable an author to help identify their content across multiple platforms, as well as suggest other quality authors to consider.
That may be true, but I think one of the good points made on the OP is that it might actually be cultural constraints that keep them from solving the problem:
https://twitter.com/paulg/status/1477761335412809729:
> You might need to do a lot of manual spam fighting initially. That could be both the thing-that-doesn't-scale, and the thing that differentiates you by being alien to Google's DNA. (They must hate manual interventions; so inelegant).
Google has some very smart and knowledgeable people, but the things they do have to fit into certain boxes, which means there are some problems they just can't fix, e.g.
* Everything has to be automated at scale, which leads to consistent poor user experience (unappealable account closures initiated by inscrutable algorithms, SEO spam).
* You get promoted by building new products, not maintaining existing ones, which leads to self-defeating churn outside of core areas (e.g. abandoning Google Talk and squandering their position in the messenger market).
* etc.
Of course you can always hardcode the site you want in the Google search results but this is hacky and not very expressive.
The fact that this feature does not exist shows that there is something deep within Google's core that is preventing them from addressing SEO spam, just like there is something deep within Airbnb that makes it difficult to filter out Airbnbs with problem reviews.
Google has been coasting for a good long time and now major players are realizing they are wide open for disruption.
Google is smart so I assume they crunched the numbers and figured out they make more money from people filtering through crappy results that include viewing and clicking ads than by surfacing good content.
I think Google is optimizing for ad revenue, not for good search.
Why not try writing a search engine specifically
for some category dominated by SEO spam?
I like to compare search engine results and wrote this tool to make it easy:There in fact are many vertical search engines. You can click on "more engines" to see the whole list.
I have an exterminator who comes to my house every couple months, and sets up traps here, poison there. I don't have any rats in my house. I do see rats running across the yard sometimes. The exterminator explains it like this: Rat pressure. The rats overpopulate and there's "pressure" (like, uh, "memory pressure", which is also a fluid concept) so they try to get into your house more, through smaller holes, as a function of how much outside drama is going on, how many they are and how overpopulated, how scarce their food supply is, how cold it is outside, and whatever else drives rats into your house. (I love the dude who's my exterminator).
Anyway, this is the same problem every search engine faces. The more surface area they expose, the more pressure they have building, the more ways people have to fake out their systems.
We have to go back to the 1990s Yahoo! model. Curated content. A list of websites that are reputable. 1990s Yahoo is the future.
Carte blanche opening your system up to anyone to inject data seems like the wrong foot to start off on, whereas my curating a moderator, someone I personally know and feel good about, trust to whatever level, and hiring them - ideally making sure they're someone you respect and you're someone they respect, pay them well, and at scale will be able to pay for itself; this did just bring to mind however big pharma and pharmaceutical trials structure and how that system can be/is/has been captured - and so perhaps the pressures when dealing with multi-billion dollar market categories will always lead to shenanigans if ever trying to centralize too much, not allowing for de-risking and broader resource distribution via sales/profits to more parties than the "5-star" rated products.
In a thread on HN, I think it was yesterday, a few people posted about review sites where some product reviews are free - but others you had to pay for. A system to facilitate such organizations could allow a highly competitive environment, where organizations develop/build a brand - build trust for their brand as being competent and thorough - and so then over the lifetime of a reader/customer, perhaps they'll spend $1,000 buying reviews (say 333 big purchases over 40 years that you're willing to pay $3 a hit for) to make sure they're ; mind you there will be organized that could be captured to say promote one conglomerate of products over another, perhaps even regionally, but I'm beginning to think it's a necessary layer to combat the shit show that is Amazon (et al) reviews. Ideally these systems and how the reviews present the information, and how thorough - the technical depth and breadth and testing done - will help educate those who dive into using this system, which will sharpen themselves while keeping reviewers on their toes and arguably strengthening their organizations and competency as well.
Then Google came along and was a better search engine, for a time, that was a traffic leak for Yahoo! - and then Google has now devolved; I also thought Google had a good shot at competing with Facebook, but whomever's pulling the strings there, the launch of various platforms, they don't seem to understand it can take 5-10 years after the MVP of a product is launched for it to mature - but for whatever reason their executives or managers haven't been comfortable pulling the trigger, arguably because anyone with that entrepreneurial spirit just takes their idea and gets funding and owns a large portion of whatever they've done; but then you can never develop a full breadth, holistic ecosystem, that can grow into every crevice, nor as broadly, or nuanced as possible - so they're stuck being Search, Gmail, Calendar, etc.
I'm quite certain I've figured out the foundational MVP facilitating an "infinitely" growing system and that would allow 3rd parties to integrate, however I have severe chronic pain that messes up my executive function, so it's difficult for me to actually self-direct and execute - I'm stuck mostly in a low activity, stream of consciousness and go-with-the-flow life of routine - otherwise I would try to launch my plans, which I've done plenty of UX/UI for, as that is simple enough that somehow bypasses higher executive function (moving a pixel and then responding via visual feeling of it isn't complex) - but organizing to turn that into adequate specs to get solid estimates or fixed price quotes for work is extremely difficult to me.
On January 11th I do have a surgery that may or may not reduce my pain by 50%+, may or may not improve my executive function, ... I've even attempted to write draft "Show HN:" posts to explain what I am doing, the starting feature sets, the reasoning behind the design decisions I've made - but it just gets too complicated too quickly for me mentally then to be able to organize further or polish it. I think I have the perfect domain name for it too: ENGN (engine), what makes me smile every time I notice it in my layout/mockup of it is in the search input box it says "Search ENGN". My username on HN actually is an older incarnation of a plan I had, loceng being a short form of "local engine" - and ENGN being from engine, a name I brainstormed after Tumblr sold for $1.1 billion - and I realized that eventually I'd want to try to do my "local engine" idea but that that was too long for a brand name. Fortunately engn.com was for sale at the time, I can't remember if it was $2,000 USD or $4,000 USD - either way, not a bad price for a 4-letter .com that's pronounceable to something with meaning.
I've wanted to write a book too - on health, health systems, and on these systems we're talking about here. I'm 38 now and I taught myself to program when I was 11, learned SEO at 15, evolved to design as I'm more creative and programming became mindnumbing to me, and eventually thought I'd need (or want) VC money - so I started engaging on Fred Wilson's of USV.com's blog - AVC.com - so I have plenty of self-taught experience. The problem is even going back to my shorter or longer writings, or comments of mine on HN or other, it's nearly impossible for me to try to do the organization of it all - to compile parts, etc.
Maybe this surgery goes well and I can begin to do more, or maybe it doesn't; I've tried to hire people or get help over the years but 1) no one has been willing to engage enough as I'd need due to my executive dysfunction, and maybe that's a moot point as 2) it's extremely difficult for me to even manage someone or an ongoing project - whereas I could explain things and direct if people are initiating, if others are directing the conversation, then I could respond - but otherwise I can't do normal oversight and management any longer.
The most accurate odds I can give that this surgery will help (piriformis syndrome, my sciatic nerve goes through the piriformis muscles, rather than around it - so there's constant compression + that's worsened with use/engagement of the muscle) is 50/50. There's a high probability that this surgery t's not related to the primary source of my pain, which is from LASIK eye surgery I did 7 years ago - I got arguably the worst of the worst symptoms: central sensitization and hyperalgesia - a hypersensitivity to pain, where all sensations, pain especially, is amplified to as what seems as strongly as possible; and why I must highly limit my activity level as any little stresses on body, likely even normal natural muscle use which causes micro-tears, then compounds the problem and will take many days of very low activity to return to a still difficult-dysfunctional baseline. But perhaps there's a high probability that the sciatic nerve having been compressed for most of my life, my mind, nervous system, could handle that level of pain/sensitization - but then the damage to the cornea that happens in 100% of LASIK surgeries was what finally broke the camels back.
If this surgery doesn't go well, doesn't help - which it took me 1.5 years to even find a surgeon who does this type of specialized surgery - then I'm afraid I may end my life because this pain, the lack of productivity, of being stuck, of quite little social interaction overall - HN is likely the most stimulated my mind gets, only possible to write this fluidly when it's been at least 3 days of eating a very low inflammatory diet and very low activity level - and only if I've been mostly inactive, primarily sitting, since waking and getting out of bed - to not trigger any pain in my body. Anyway, it gets boring, repetitive.
I've thought of trying to find an Elixir/Phoenix/React/etc developer or agency on Upwork before surgery and try to struggle to get them on at least developing the initial foundation of ENGN, but aside from the struggles I listed above that I'll encounter, it also will cost additional money - and I've not worked in 5 years, I've spent $250,000+ on stem cell treatments to heal old high school football injuries, that I didn't even know I mostly had and only weren't tolerable after LASIK made my nervous system super hypersensitive - and to pay for this surgery my mother is taking $27,000+ USD out of her retirement; I'm in Ontario, Canada, but the healthcare system has been practically useless to me. Even if the foundation for ENGN would cost just $5,000 to $10,000 to get the ball rolling in terms of starting to get users to signup and bringing in revenues - it's more money, but even thinking about that additional stress would put on my mother, then adds to my already overwhelmed nervous system - so there's plenty of resistance there to overcome on its own. There's also always the potential I'd somehow hire a bad contractor or agency, bad in one or many ways, and then the MVP wouldn't get finished - primarily because of my own incompetency-dysfunction, and then I will just be reminded, again, of how stuck my life is and how it barely moves forward - personally or professionally.
I'm living a version of the Groundhog Day movie that keeps repeating itself, except where I'm in pain, and so far where I can tell people my story and ask as many people as possible to help and nothing happens. It's why last thing I try is this surgery, though I am supposed to do another PICL stem cell treatment - where they treat tissues inside of my neck - the first one did reduce my neck pain and migraine some - for whiplash related issues, in part from football - because they only treat one side of the tissues, not all of the tissues, the first treatment - and so the second treatment they target the remaining high yield tissues. But I'm certain I'll know after the surgery if there's any improvement or hope that my life can start to become different, and even though there is a stromal stem cell treatment that was developed at University of Pittsburgh - that had very successful human clinical trials, that were fast tracked under compassionate grounds in India, to heal/regenerate deeper corneal tissue for severe scarring and chemical burns - the ETA for it being FDA approved was 5 years to be clinically available in the US, perhaps less time before available in India, but I'll have nothing else significant treatment wise for further pain reduction to look forward to in the near term after getting this surgery - and so why I'm afraid I won't be around much longer if it doesn't help much.
Listen, first of all, do not consider ending your life. Seriously, you're way too smart for that. I'm sure you've got it worse, but I've had enormous sciatic problems in my life, I've had 3 herniated discs; they're behaving for the moment after massive doses of cortisone and without any painkillers, but I know what it feels like to cough them all out of my back at the same time. Not to be able to put a foot in front of another or turn your neck for weeks. (I'm a huuuuge fan of intramuscular cortisone injections, though. Like 5 or 6 large cortisone over a week, with some B-12. Every couple years. Not in the spine... fuck that. Alternating butt cheeks. You won't feel any benefit until the third day at least. If you can convince a doctor to give you that for a week, you will be fucking superman. They won't do this in America unless you know a doctor personally, but they'll do it in Mexico or Spain. I had it the last time my discs went out and it's been 6 years and the inflammation has not come back. They thought it would).
Anyway, before you off yourself, do try a fuckton of intramuscular steroids. The fifth day I levitated off a bed in the hospital; I hadn't walked in a week; I felt so good I went to a club; I got drunk and spent the night on a beach drinking and making out with an 18 year old model from Denmark. Seriously. There were wild cats walking around; it was winter on the Spanish coast. If you do one thing before you die, go get five cortisone shots in your ass, in a week.
I also got the hiccups for 24 hours and couldn't sleep, but that's neither here nor there. And I got temporary blindness in my left eye from fluid behind the retina, caused probably by too much testosterone. But. Goddamn it, I'm ok. You can be okay.
Enough about that.
About Yahoo and Google. That entrepreneurial spirit is, in my experience, way too often just about getting the funding and fucking off. We all know why these companies go downhill, but somehow it's always such a shock when they actually deteriorate in front of our eyes, huh? Google's search results, for instance. I would have expected their cofre business to stay more or less fine, not collapse a couple years after all the competition was eliminated.
It would be fine if they didn't grow into every crevice. Get search right, that's all we ask. I don't want Google to be my chat room or my shopping site. Why do they need to? Search is huge. They own 90% of the market.
>> but organizing to turn that into adequate specs to get solid estimates or fixed price quotes for work is extremely difficult to me.
That's always the worst. The business side. I've always just built things and hoped for the best. It sounds like you've got something interesting going there, although I have no damn clue what you're building, that's an exciting feeling. ENGN is killer. If you own ENGN.com, hell, money well spent.
I don't understand what you mean about "executive function", since you obviously have the capacity to write well-crafted email and think pretty clearly; perhaps I lack the executive function to discern your lack of executive function (I'm a brutally self-punishing alcoholic, but otherwise a damn good programmer)
Anyway I don't know if you're trying to ask for pointers to workers for this concept, I'm probably not it; I'm $200/hr and I'm already covered for the next year. This, however, should be your symphony. And I think you know how to do it.
and then the original vice documentary: https://www.youtube.com/watch?v=VaMjhwFE1Zw
and then tried the breathing: https://www.youtube.com/watch?v=tybOi4hjZFQ and cold showers and they've been helping a lot with inflammation!
There are tons of benefits of cold showers, so if you can start doing that, that might already help a lot! Here's a bunch of results you could check out: https://www.youtube.com/results?search_query=cold+showers+be...
Central sensitization and hyperalgesia is really a different kind of beast when it comes to pain, people including most doctors really seem to have a hard time understanding it. That added stress to my nervous system from opening my right eye and triggering/ramping up the sensitization/pain from just low activity and careful movement was enough to lock up my thinking, and arguably my emotions, but more specifically the locking/the eye pain and increased pain from movement makes fluid thinking have much more friction added to it; one good or as relatable as possible thought experiment for a "normal"/unaffected person may be to imagine the feeling when you get something in your eye and you desperately react to get it out because it's so painful: now imagine that foreign object feeling is immense because it's permanent and broad, it's "a lot of objects" in your eye(s) because your cornea and nerves were sliced across 90%+ of the cornea (and so abnormally/constantly signaling as such), and imagine how your brain/nervous system may try to cope from such an overwhelming/overriding system constantly firing to draw your full attention to what your eye is telling your body is an active/present moment object in your eye - potentially as if you just walked into a sharp object and your eye is signalling for you to literally freeze still because it thinks you're in the moment just done something to critically wound your eye and moving another millimetre or less would threaten your survival (as the evolutionary strength of the reaction has come to dictate).
Well, the answer is some people after LASIK get this severe reaction, their nervous system gets overwhelmed - arguably the more naturally sensitivity, creative, healthy, and grounded people will have stronger/more detrimental symptoms/reaction to the eye damage; central sensitization and hyperalgesia that LASIK for years completely swept under the rug as being possible and only recently admitting to being a potential "side" effect. In fact they purposefully mislabeled what should be called corneal neuralgia syndrome as "dry eye syndrome" to mislead people away from learning that the "dry eye" part is actually on a spectrum of symptoms caused by the damaged cornea/nerves that happens in 100% of their surgeries; non-LASIK done research done has shown up to 40% of people have permanent problems after LASIK, and in 2011 one of the expert FDA advisors, who votes yes to approving LASIK, published a letter to the FDA asking them to immediately recall LASIK because there was data they were ignoring that they should have never been ignoring, and that it should have never been approved in the first place. Part of the medical industrial complex where arguably regulatory and institutional capture has occurred in the name of profits.
So again, I'm writing this out first thing in the morning with my right eye shut still, my eyes having had a reduced level of pain while I slept allowing my nervous system to calm down some, along with recovering from the movement yesterday that compounds with the eye pain and worsens the mental/executive symptoms/dysfunction. There are a few tangents I know I lost track of writing above, I'm not going to be able to go back to insert them, so if what I wrote above doesn't flow well or seems to be missing something then that's why.
To answer you specifically - I've done breath work and cold showers in the past, daily to multiple times daily for weeks, and there was no net benefit and arguably it was added stress to my nervous system that's already overwhelmed from active overriding [eye] pain that can't be reduced. That same eye pain and the friction it added once I'd opened my right eye, and in part I had used a lot of mental energy already writing longer comments on HN - that mental energy otherwise goes to trying to maintain focus/distraction on anything to try to keep me from getting completely sucked in/lost to the pain, the friction that will block mental (including emotional abilities/processing), lead to that small inconsequential reminder or level of frustration triggered from a well-intentioned person offering advice - where I was too blocked, too much friction/resistance from the pain at that point that it took me at least 20 minutes of going back and forth to reading parts of what you wrote then having to leave because of frustration/irritability being triggered because of the very slight pressure/stress that was added to my already overwhelmed nervous system; I'm not sure I'm describing this well here but it's best I can do at the moment.
That's why logically, emotionally I'm already certain it's the compassionate thing to do and I'll forgive myself for not being able to handle the burden nor for being able to handle knowing the burden/consequences I'd be leaving behind, logically I'm only doing the piriformis syndrome surgery as the last thing I try - so if it doesn't have a dramatic reduction in pain and cause a noticeable reduction in my executive dysfunction (my argument or hypothesis is the eye pain, a/the major source of pain could be compounding with another potentially major source of pain [piriformis syndrome, similar to sciatica], so eye pain could be say 33%, the PS could be say 33%, and the compounding could potentially cause runaway/cascading/feedback loop pain of 33%+ - for potentially a 66%+ reduction in pain/sensitization) then I'm gone. Of course I have arguments and proof points from my prior experiences with healing injuries with stem cell treatments as to why the surgery may or may not help, so the outcome is completely uncertain if it will help enough, if my executive function will improve at all, if my quality of life can improve at all, and so the most accurate odds I can give it is a 50/50 shot.
I've also done plenty of water fasting, I did carnivore/high fat red meat only diet for 8 months - now I try to stick to just organic red meat, kale, and raspberries. I've also done many Ayahuasca ceremonies, MDMA therapeutic sessions, massage, acupuncture, etc. There are pros and cons to all of it, some I've been able to maintain because it's at least neutral - other things like acupuncture are intolerable, for example, as after sessions to clear my Wood meridian energy line [3 end points at right eye] then my body is completely calm but then the pain is completely localized at my right eye and then for the following 6-8 hours I feel an intense burning sensation at my right eye - meanwhile feeling the pain and strwss referring from my eye and building up the stress into my body until there's an equilibrium of sorts where the pain level/hypersensitivity/sensitivity level in my body matches the level of pain in my eyes.
My nervous system is very healthy, it's the pain that's overwhelming, short circulating or disrupting different processes. I can't do anything more to reduce the eye pain because the stromal stem cell treatment is up to 5 years away from being clinically available, everything being delayed due to pandemic as well, so my last hope is the piriformis syndrome surgery - for which I have no real or solid reference for how much it might be sensitizing my nervous system - all I know is "sciatica" symptoms bothered me enough 15 years ago for me to first try to investigate it, but that my nervous system could handle that pain along with vast majority of high school football related injuries that I wasn't even aware I had, and it was only after LASIK that my nervous system and therefore mind was disrupter.)
Not necessarily: https://yacy.net
Now, one of the deficiencies here has been examples. Try this: "best miter saw". you will not find any websites that actually discuss the answer to this question, despite it being a product category with a lot of price variability and performance tradeoffs (weight, capacity, power, cord vs cordless, accuracy).
Nearly any product reviews for large purchases follow the same pattern unless consumer reports has decided to dig deep (e.g. washing machines).
How about guitar strings? Sandpaper? Printers? google's algorithm has allowed profit motivated websites to displace the commons to too great an extent.
Marketers/SEO people are starting to infiltrate this as well, but since they can't control and SEO the content on Reddit nearly as much, this still works pretty well for now.
Fun story: I knew a person who worked for a home building website part time. She got paid to write stories on home renovations. She had never done _anything_ she wrote about. Mostly, she gathered up other blogspam and recycled and rewrote it without citation. Sometimes she went to forums, sometimes reddit, sometimes youtube. But, the universal part of it was that she had to produce 2x pieces of content per week endlessly. Just for a local LA builder. Most of the content wasn't "wrong" but, it also wasn't exactly incisive and didn't include any details that would have been useful. Instead it was just filler. The worst part is that it consistently improved that company's search ranking.
Content farming needs to die.
There’s definitely an element of “we got our profit so fuck it” with respect to the search engine advertising business and Google’s incentives to make search quality better, but that doesn’t change the difficulty of the underlying problem. If I wanted to pay Google $5 per search for super high quality results, even $50, they can’t just make this product better to get my money, they are fighting an ongoing war against adversarial SEO which prevents this from ever being better than stalemate at best or more likely due to economics, the slow slide into declining quality we see due to the SEO side having more money with which to pay for engineering brainpower.
Use tracking especially and JS generally as a weight in ranking so sites that contains much of any of these needs to be exceptionally high quality to float to the top.
This means sites with limited ads and tracking, typically enthusiast driven pages float to the top.
Now always when someone discusses a novel way to combat webspam someone will immediately counter: if this becomes popular SEO hackers will immediately start doing this.
Well - if reducing page size and removing tracking becomes a leading SEO trick I can deal with a bit of SEO hacking :-)
Yet for some reason I feel Google won't start using this very simple metric :-)
/!\ The term "," contains characters that are not currently supported
/!\ Try rephrasing the query, changing the word order or using synonyms to get different results. Tips.
sample: https://search.marginalia.nu/search?query=site%3Astackoverfl...
Please make your engine ignore the comma, it shouldn't affect the search.
Either ignore the site:... expression or filter sites accordingly.
Thanks a lot for creating Marginalia!
site:-queries are supported, but only at the first domain level. (e.g. site:marginalia.nu; not site:search.marginalia.nu). I might tune it so that it strips subdomains automatically, that is pretty trivial.
e.g. this works fine: https://search.marginalia.nu/search?query=site%3Amemex.margi...
If you just do a site:-query without any search terms, you'll get a dump outlining the crawling status of the site. Like this:
https://search.marginalia.nu/search?query=site%3Astackoverfl...
You can see it's crawled 80 pages but not gone much farther since SO ranked so low.
https://searx.space to learn / get started. Find one and visit it, click Preferences upper right then follow your schnoz.
Any new niche search engine will go through a small window of time where they have the luxury that none of the sites they are indexing are spending all their effort trying to reverse engineer your signals, and optimize against them. I'm incredibly skeptical that they can remain useful once people all the SEO efforts of various marketers start to be turned against them.
Maybe the main observed motivation. But I'd argue that a lot of that is just a fraudulently-profitable front to much more devious problems.
And pintrest would die.
Block results from specific domains on Google or DDG:
google.*##.g:has(a[href*="thetopsites.com"])
duckduckgo.*##.results > div:has(a[href*="thetopsites.com"])
And it's even possible to target element content with regex with the `:has-text(/regex/)` selector. google.*##.g:has(*:has-text(/bye topic of noninterest/i))
duckduckgo.*##.results > div:has(*:has-text(/bye topic of noninterest/i))
Bonus content: Ever tried getting rid of Medium's obnoxious cookie notification? Just nuke it from orbit: *##body>div:has(div:has-text(/To make Medium work.*Privacy Policy.*Cookie Policy/i))It would be nice if Google would ask you the simple question: "Did you find what you're looking for?" Instead they rely on the assumption that users only stop looking when they've found what they're looking for.
These days, there's a reasonably high chance that I quit looking because I gave up in futility--not because I found what I was looking for.
It's also the case that there's no way to train Google not to omit search terms or generalize them to the point of uselessness.
I really wish abusive SEO were the only problem but it's far from the case. Search results being crappy is a cumulative effect. You could solve SEO spam and I'll still not be able to find a USB SuperSpeed cable because it gets generalized to "usb cable" and there are a gazillion more charging cables than there are SuperSpeed cables.
Used to be that you could quote things to indicate that you really meant it. That's fuzzy now too. Every time we figure out how to circumvent the bad results, features are removed.
They're friends with the king, so don't hold you breath.
All the various suggestions in this thread plus far more complex and insightful solutions are known to Google. Most of it boils down to using automated user feedback to improve or measure search result relevancy.
Google doesn't need to solicit user upvotes / downvotes to improve rankings. They can monitor user clicks on results in addition to analytics on the sites the users visit to determine which sites are relevant to which searches.
Google doesn't optimize for search relevancy.
Travel - Expedia, Hotels.com, Kayak, etc.
Consumer Goods - Amazon, WalMart, EBay, Etsy, etc.
Automobile Purchase - Cars.com, Autotrader, etc.
Career/Job - Indeed, LinkedIn, etc.
As Google continues to lose search volume on these big revenue categories it is going to make spam much more difficult as they are working to sort out long tail spam. Way harder.
The same with Amazon. You'd think that with all the armchair search quality experts cropping up lately there might be more vocal complaints about the fact that Amazon's own search can't find basic consumer products sold by Amazon itself. If I want to find stuff on Amazon, I search Google for it.
Mayo clinic, Harvard health, and pubmeb do a great job with health info. IMDb for movies, Goodreads for books, *gearlab.com for reviews, booking.com for accomodations.
I think the biggest threat to Google isn't a better general search engine, it's user behavior switching to more domain-specific websites as the top of the funnel. E.g. people going directly to Amazon to search for products instead of first searching Google.
To some extent, Google has figured this out, which is why they now have a dedicated flight search, hotel search, product search (Google shopping still exists and it's pretty good!), etc.
In any case, I'm not sure that competing in search is a very attractive notion. AdWords is the only meaningfully profitable search and business. Even if you steal 10% of Google's market, that absolutely doesn't translate into 10% of the revenue.
That said, recipes. Someone make a search engine where the top results don't start with 500 words on the history & etymology of butter, because that's what Google want.
Recipes have become a bellwether Internet problem. In the past, your great-grandmother had a card file with a bunch of 3x5 index cards with the ingredients and instructions on how to make everything, and they pretty much all fit on one side. There was a great deal of domain knowledge required (e.g. "whip to stiff peaks"), but these things reveled in their terseness.
Internet recipes all begin with 9 paragraphs of the author's first time encoutering the dish in a Moroccan bazaar in 1997, and the life story of the chef. There are two embedded 10-minute videos of the lifecycle of the vanilla bean. And then you get to the ingredient list. Then two more 10-minute videos, then instructions.
The drive to make recipes full-contact Internet content has changed what it means to be a recipe. This is similar to how cooking shows evolved from Julia Childs working on a sound stage to a carnival barker presentation with vivid personalities dominating the scene.
I'm not sure there is any technological solution to a problem that has fundamentally changed what it means to be a recipe, short of establishing a new informational silo in the form of a new Web site devoted to recipes only. You could encourage an RSS-like format for recipes, but that requires buy-in from places that profit from the new evolution. This new status quo may be good or bad--you can make the argument either way--but it is what it is. A cultural change is required more than tweaking algorithms.
(Unless tweaking algorithms can be foundational to cultural change, in which case we really, really, really need to take a hard look at the corporate behemoths and their algorithms, and sooner better than later.)
We need a semantic dupe filter: If it doesn’t add new facts or new ideas, treat it as an identical copy.
The technological solution would be to stop rewarding them for these monstrosities. One of the main motivator for turning a short recipes into a 19 page essai about the chef's life is that more words = better ranking.
Which is the same as there being no technological solutions.
I’m curious what would happen if those products were split up into 2 separate companies.
Over COVID, I did the whole fitness thing from a few different angles (overhauled diet, trained for a marathon, now lifting weights a lot). I found I could only find good info by going directly to a trusted source - literally, typing http://www. like I’m in the 90s or something. This is the exact issue a search engine should solve, but Google doesn’t.
If not, at least one of us is going to hate this new search engine.
What sorts of unique things did you do that Google failed at? Maybe you read through discussion sites or got tips from books or something like that
Importantly, this was bottom up: it was largely recommendation engines suggesting people I then filtered through for what I wanted (running, bodybuilding) vs didn’t (traditional weight loss). I couldn’t specify what I wanted, or it would be garbage SEO spam.
That helps a lot! Thank you.
The simplest way to compete with Google is to create a DIY Webrings site that disallows harvesting of data by Google. Charge curators to create a webring, and let curators select three hashtags and a description that represent their list of fifty or fewer sites. Use the revenue to pay a human to curate the list of hashtags, and let users tip a webring curator in gratitude with an Apple Pay button.
This is how to make a million dollars, Pinboard-style, out of the ashes of the original curated Yahoo idea and the information structures of hashtagging. It doesn’t work if you allow free-for-all infinite-sized lists, it doesn’t work if you allow free-for-all hashtags, but with clear limits and moderation of tags (instead of webrings), it would thrive. By moderating tags, users can keep the webring they paid for, and SEO rings will be stick out for having no shared network with any other rings, which allows for easier detection and culling of malicious non-participatory actors. Plus, with the curation networks in place, it becomes possible to bubble up rings that have unusual content for positive human moderation activity.
I tried to find some good podcast lists yesterday and each site I visited had a really interesting cross-section, but there were so many duplicates. I wish the ring site existed, so that it could remember what it had shown me already, and I could say “show me rings that intersect with this podcast and have something new I haven’t seen before”.
That’s where the theory of pagerank and the practice of curation and the capabilities of search align, and given that moderation of hashtags scales very cheaply, is a billion dollar opportunity that Google and Amazon cannot compete with if handled properly. It’s not about trying to get a cut of every visit’s revenue potential. It’s about giving human beings a directory that respects their time and remembers what they’ve seen.
This isn't going to happen again. Google isn't going to sit around twiddling its thumbs while a competitor develops a better algorithm.
You have to attack the problem from a different angle entirely (make something that looks nothing like a search engine), I don't think a niche market is going to be enough.
Perhaps you just want to make something that scares Google into acquiring you, rather than actually bettering the situation. If that's the case, I implore you to think of doing better ways to spend your life.
To displace a giant gobbling 180 billion dollars a year? Yeah, no kidding. But nobody is asking for that. They are just asking for decent search results.
This is the exactly the kind of thing that Google cannot fathom manually doing. As if entering facts into a computer were morally wrong somehow. They'd much rather launch the equivalent of a shell script that harnesses face-melting amounts of computational power, processing literally trillions of webpages in bulk, signal and noise together, junk, spam, and misdirection alike, to learn bad associations and then serve them up with no human review and then put the full force of their reputation behind a results page that apparently people never check and certainly can't correct because of the inscrutability of a machine-learned model that has few to no levers to adjust.
I had a chance to debrief people who had left their relevance team and they told me things that were outright contradictory to what rank-and-file Google employees have told me. (What they told me did make sense in terms of my experience as an IR system developer, SEO publisher, etc.)
Microsoft bought a company called PowerSet that had extracted a large database of entities and relationships from Wikipedia and used the technology to make the "Bing" search engine.
Earlier Microsoft engines were a joke, but Bing was so good that Google saw it as a threat so they bought Freebase to get a similar kind of database, then they killed it to incorporate it into the "Google Knowledge Graph".
For all of their hating on semantics note that they hired R. V. Guha as their chief scientist, who worked with Doug Lenat on the notorious
No, DDG bang operators don't let you do this. I want an SERP, not a shortcut to a single site's on-site search.
> I’m pretty sure the engineers responsible for Google Search aren’t happy about the quality of results either. I’m wondering if this isn’t really a tech problem but the influence of some suit responsible for quarterly ad revenue increases.
Please no more of this. Two men, Page and Brin, together have basically unfettered control over Google.* If Google does something bad then, unless it's genuinely something small enough that those two could not be expected to hear about it, it's happening with—at the very least—their acquiescence. And low overall search quality is not something that some "suit" is successfully hiding from Good Czar Larry. They could fire the "suit", or command him or her to make other decisions. This is—again, at the very least—something that they have chosen not to do. The responsiblity lies with them.
* There is the risk of lawsuits from the minority shareholders, I assume. But IIUC this is not realistically that big a restraint on what shareholders with a majority of votes can do. However IANAL.
And to be clear I want to be able to control these myself, not algorithm trying to guess my preferences. No guessing, just do what I tell you to.
Multiple search profiles with different priorities would be nice too.
I would like the search algorithm to be transparent, I should be able to tell why I got a certain result and how I can avoid such results in the future.
In an ideal world, you build a thing, and it's done. It runs automatically and prints out money.
In reality, you still need human labors to do manual tasks, even in tech industry.
[1] https://www.searchenginejournal.com/google-eat/quality-rater...
Crowd-sourced humans are making Google appear more intelligent than they actually are. I always envisioned that spam efforts would just immediately set off an alarm that would be handled by a bot to blacklist you without a human even knowing your site existed, but there still seem to be at least a few ways to game Google's search ratings.
Showing results for searchterm
No results found for searchterm
Followed by an unending list of random celebrities I don't know nor care about, businesses I've never been that sell items I have absolutely no use of, and random foreign news articles.Failure to recognise the typo is unexpected but forgivable. But then, rather than helping me with my search, they attempt to distract and lead me away from it - using triggers that you'd think they should have known wouldn't work.
I really don't understand how this is even possible, and it's not a rare occurrence.
Also, why hasn't Apple built a search engine yet? It baffles me that they chose to go head-to-head with Google on Maps, yet outsourced their search engine. I would've liked it the other way around: Google Maps and Apple Search.
I agree they should make their own search engine. But currently they’re being paid a ton of money not to.
But this only works for certain industries. It’s much less common to see this kind of tactic if your searching for say a coffee shop because they sheer number of local results let’s Google be “hyper local” with these kinds of results.
For licensed trades the solution is to go to your official trade licensing body (for the UK it's the NICEIC for electricians and the Gas Safe Register for gas/HVAC technicians), for unlicensed ones it's more difficult. There are "review" sites that claim to provide good results but their business incentives & vulnerability to spam/fake reviews are unknown.
(Disclosure: I work on ads at Google, speaking only for myself)
Honestly what I think is every search engine sucks these days, and Google manages to suck a little bit less.
The reason is because how easy it has been to publish low quality content. It's rare to find high quality contents. The issue with search engine is that they don't show these rare contents. These aren't recommended by default. These are hidden.
What happened is the recommendation system is broken! If there weren't any neural networks making decision, it would have been different issue. But with modern search engines deploying recommendation system, I think it is all about rich gets richer scheme. You can't recommended new or fresh but quality content because it was never visited! So, when the entire backend is relying on user data, the system is being fed crappy data because users don't care and those who do are few in number.
As long as system is making revenue, it will be this way. Most people never care at all and would never be bothered because they only care for simple queries. If anyone deviates from the norm, Google search results are pretty bad like every other search engines.
Money's no object for him, so he wanted to outsource the filtering, ranking, and interpreting of results. Would be even more useful today (albeit a tiny TAM.)
you mean Larry Page?
If such a demand exists, should we not be seeing an active “secondary marketplace” for people offering to do 15-minute human meta search & research tasks?
A human can always do better: take Google's results, then remove SEO spam/duplicates, extract more relevant snippets, combine results from multiple nearby queries, etc.
Demand exists, but someone has to build it. And it's unclear how big the market is.
One of the solutions is to hire people remotely that will build a list from searching and verifying that the information is correct. There are a ton of scaling issues with a person building the lists, and you will still run into errors, but it has been the best way to verify and have the most up-to-date data.
A problem I noticed is people are so use to getting tons of results when they use Google or another database site, so they expect the same with human-powered results. SO I think a lot of expectation setting and also just working with the TAM that understands the impact a curated list will provide over a outdated database. Both have their costs but for a higher impact the human curation will always win over automation. (I have an example of 3 different lead gen sites having a different city for 1 company, the AI on the websites mixed up the location, a Human would be able to catch that)
This is why I am very hyped on Brave search's goggles feature , it will let you share exclude / include site lists to use w/ the search engine. Hopefully it will empower these niche communities to curate a list of non-spam sites ( like the ad-blockers do with ads today )
On the other hand, Google needs to maintain the ballistic trajectory of its revenue growth. So how can they fix search quality when they've minmax'd themselves into this situation in the first place? If they were to make the ads background yellow again, that would have negative short term effects that I doubt any career exec can stomach.
The other moats are lock-in to advertising networks, website metrics, and effective control over Web standards through the Chrome browser.
An alternative search platform might provide better search. It would be fighting Google on at least three other fronts. It might have some success, but it would be challenging. (As history largely demonstrates.)
Even a rival tech monopolist, Microsoft, barely holds even with its own search offering (I use that indirectly via DDG), and scrapped its own web-browser development
If you have a good search engine, people will flock to it, and search ads will be valuable. That's it. That's how Google became Google.
The fact that Microsoft couldn't do it honestly doesn't mean much. Microsoft also couldn't do a phone OS, a portable music player, and many other things. They have a complex web of conflicting interests that $SEARCH_ENGINE_STARTUP does not.
It's just that "search" is really a web of interrelated services, capabilities, and revenue streams, and they tend to reinforce each other strongly. I'd like to see the monopoly disrupted.[1] But I don't think it's just a matter of "build a better mousetrap^Wsearch engine." Attack one corner, and Google will snipe at you from the others.
And with the AdWords cash cow, they've got an immense revenue stream.
________________________________
Notes:
1. Well, mostly. Google's acquired so goddamned much personal data that the premise is frankly kind of terrifying as well --- a weakened Google with neither the revenues nor talent to defend that pile.... And I'd really like to see the toppling occur without simply raising a new monopoly in its place.
Unfortunately, defaults matter and Google is spending billions of dollars yearly to make sure they are the default search engine wherever they can. Most people don’t switch from default.
my guess is the rate of spam content production far outpaces the rate of original content creation. so the power law concentrates even further in the tiny percentage of OC and a moat forms around them (highest ad $, highest authority/authenticity).
where do we end up 5 years from now? further consolidation and the continued return to aol style portals (telco/media giants and fast-lane to own content?) pay-to-access silos dominating the internet?
[1] oversimplifying a bit of course, there was a novel ranking method that was more than accurate enough, and it scaled, which allowed for the search ad business to go gangbusters.
Sure a search engine that specialises on narrow area of knowledge without much money in it, can be very relevant and bullshit-free.
But there's no way to make it work for the general web search. People hack things. If they didn't we'd have Communism built by now (yes the "good" - classless, stateless one).
So, a curated strategy where the users can UP/DOWN sources they trust for particular topics. Relevancy and user ranking of answers determines score. Of course, the user can search any sources they like, but a voting system would control the defaults.
I think this avoids slanted results as well, because topics are objective. The subjectivity will be in the comments, where they belong. That's in contrast to today where SEO scams can determine how high up results are.
So, to game my proposed search engine, you need to infiltrate the users. I believe this is harder to do across numerous sources compared to the current system, which is game the algorithm.
I jumped onto youtube, I put in the model number, same thing pretty much. I get unrelated videos mixed in with the model number I put in. I'm pretty sure some videos even though the model number is the same is not being shown. Ironically there's a video explaining a certain solution and warning people not to fall victim of another video scamming people, which the dislike has been removed, comments obviously deleted so some people may be calling and getting scammed.
I will give you my two cents. I have used duckduckgo, bing, searx etc. for extended periods of time and hated every one of those things. The problem is that what you search seem to be essentially the gateway to wild west of internet. I understand the proposition of spam control in search engines, but atleast to me I think the early days of google without DMCA and copyright bans made google the best.
I fear SEO spam control will only bring the worst of the moderated internet. It will not be the first time big tech tried to douse a gasoline fire with more gasoline because they taught the more fire meant the previous fire will get suffocated by the lack of oxygen(?). Rather than using "AI" as a crutch to solve SEO as a problem, I want to see an option that is true to 2005 era google.
Reputation, non-fakeness etc. can be derived from it for anybody - you just list identities you trust/follow (with weights?) and anything you look at can be scored.
Virtual identities can also be created, ie. identity listing all links mentioned on HN (with positive sentiment only?), links from wikipedia etc. so people can follow those to create their reality graphs.
The interesting part is that it doesn't claim universal truthness - depending on who you follow your results will be skewed towards their opinion of the world. Ie. if you follow MIT, Wikipedia and E. Musk you'll see different view of truthness than somebody following FOX News and Flat Earth Society for example.
It could be interesting to focus on "dislike" marking (only?) as it may be much more lightweight to approach it from blacklisting side.
It was big during the dot com days, but withered after Google.
Interestingly, I do think that that model may need to be revisited.
Edit: I feel that Reddit is filling some of this need, at least for things like Vaccuum and Espresso machines with dedicated spaces.
It's easily one of the largest SEO players out there, and they've been on a buying spree recently with their purchase of Meredith. The quality of the content has gotten better, but it's still a monster of an SEO optimized content machine.
I just remember the concept from way back.
I'll start: "how to rent a car" [0]
[0] worth noting that I personally get somewhat reasonable results for this, with a 3rd result from nerdwallet.com and a 4th from wikihow.com, both of which seem to answer the question in an unbiased way
They don't exist without a search engine.
A major complaint is that there used to be good free reviews of commercial products that could be easily found.
That is not "all the information". Information about the current round of commerically advertised products is something like 5-10% of all commerce (or less).
And we are entering a world where "all the information" is what we do all day, what we say, how we react to different stimuli.
That is the real review sites - why do people take this train and not that, why is that park safe and this one full of muggings.
We need to solve the Google problem not because we want blogging like it's 2009 but because epidemiology is about to open humanity's eyes. And it's going to hurt if we don't make it free and open.
The problems being discussed today (and yesterday in the similar thread) come from the fact that for Google user != customer.
When you have incentives that are misaligned like this, you can only go so far! We seem to have reached that point with Google, where there is not much more that can be done on the search experience front without jeopardizing customer experience (ad revenue).
Disclosure: I’m working on a paid search engine to solve this problem on a fundamental level, by aligning the incentives and making user also the customer so we can best serve them and their needs. It is called Kagi and is currently in closed beta accepting beta-testers.
Same thing is happening today, there are just more of these actors doing it. They just game the algorithm for terms related to products. Notice you still get decent serps in Google for terms that don’t relate to something that can be sold using an affiliate link.
Fairly easy problem to fix, but Google would have to hire a black hat to help solve it. But the good ones ain’t gonna work there.
It is not to be "un-Google" but because I get better results.
For example searching in a good subreddit can be more fruitful, giving answers from genuine people in moderated parts of the internet. If you get crap then try another subreddit - some mods are better than others.
Is this a business opportunity - I think so, although I have no idea how you would go about it. Maybe a decent search engine for programmers would be a good start! E.g. "Exception Message XYZ" + site with decent answers.
c0nGrats-You_HaVe_Won_ThE_Pr1ze!
..Or some silly variation of this that takes literally 0.1 ms for a human to discern that it's spam. Yet something happened to Gmail's spam algorithm in the last couple years that has been consistently letting these through. To be fair, it does catch most spam but it's only batting something like 75% and the spam it does catch is often times much less obvious to human eyes than the stuff it lets through.
I wrote a bit more about search engines competition and problems/opportunities here - How Alternative Search Engines Can Win Users https://konaraddi.com/writing/2021/2021-08-05-on-search-comp...
Come up with criteria to determine which websites are "better quality". Measure them, rank them, put the ones that fit the criteria best at the top.
On the other side, there's the people promoting their websites. Do what you can to get as close to Google's ideal as possible through whatever means. Profit.
At this point the criteria becomes useless for any real quality analysis.
I'll be switching back to Google
I wonder how difficult it is to compare the main body of text in search results, then say if it is over a 95% match with another site (I.e. it has been copy-pasted), demote it in the search results. If a site generates too many of these demotions then it gets blacklisted from the index.
My LSH implementation is here: https://github.com/loda-lang/loda-rust/blob/develop/script/t...
Example of the 100 most similar documents: https://github.com/neoneye/loda-identify-similar-programs/bl...
There can be false positives, so after LSH then do a more in-depth comparison.
Wonder if someone can throw light on to why this isn't effective.
Doesn't Google already consider that if a user returns to the results page (or clicks a second link) then the first link visited was not satisfactory. Seems like a pretty elegant solution.
And the upvote/downvote would be very tricky to implement in a way that the SEO crowd couldn't just game it horribly.
Immediately returning to the results page is essentially a downvote.
You can't really crowdsource this stuff, because the problem of brigading and other forms of abuse is way too high. Just imagine what the crowdsourced results for "trump won" or "trump lost" would look like, or hydroxychloroquine, or ivermectin, or to go with some older cults of personality, Hitler or Ataturk.
When I run web sites I frequently look at the log and find a large fraction of the traffic is from search engines. This is a problem because it costs me money to serve that traffic. It might not be initially obvious but it costs more than serving real users because the search engines will scan everything and break the cache.
Google sends a significant amount of traffic. Bing sends a detectable amount of traffic. Baidu's crawler might be more active than the two of those together but I never get hits from Baidu. Other crawlers deliver me trouble instead of value: even if I'm not interested in hosting pirate or plagiarized content, a crawler that is looking for trouble is only going to bring me trouble.
I hate doing it but I turn off crawlers other than Google and Bing both at the robots.txt and web server level because I just can't afford to serve Baidu queries.
I'd like to sign an exclusivity contract with a search engine such that they get exclusive access to crawl it and in turn I get a privileged position in search results. This would give the search engine and myself an incentive to deliver end-to-end quality results.
I mean no mincing about- recipe sites that are ads are blocked. Results with pixel tracker etc are blocked. Hell, results that are paywalled are blocked because they're useless.
And the web now looks like a 1500-2000 word listicle with 3 images becasue that is what thr ranking algorithm favours.
If you find the info you need and leave quickly that actually down ranks the page. That is is idiotic. Pages that give you what you want quickly are punished!
Moreover, they are more related to a specific type of search query, that likely a result of broad based ranking algorithms that loosely are the most efficient ranking system.
Most of the results are poorly formatted content "gathered" from stackoverflow, github, quora, etc.
And from a "person who wants to see an image" perspective, Google is purely a gateway to Pinterest or Gettyimages.
Just block any domain containing the word pinterest
That would be valuable to me.
I can show you a demo. Just to show I am not screwing around: if you don't like the demo I will pay you $500.
This is a good perspective. Where can Google not go? Places that don't lead to profit. They will try (cough Wave cough) but will give up.
E.g., Linus wasn't looking for profit, and Linux ate the world.
Kinda like they tried with YouTube Heroes?
But then, who’s to say you won’t get the same kind of backlash?
Someone (who probably doesnt have a website) said that comment moderation on your own website is to much work. Perhaps the whole internet is to much work?
But i like the spam search engine by and for spammers as a way of finding the latest and greatest affiliate marketing and blockchain swindle.
But with 3x memory needed for the indexes, the server costs probably aren't going to be bootstrap'able.
Especially for a "small" crawl of a billion web pages, event at just 10k per page.
What sort of memory and space do you have on the single server?
What's the average document size that you index?
Genuinely curious on how doable a modern search engine is on modern hardware.
The server has 128 Gb RAM and the index currently fits on a single 1 Tb SSD + an Optane drive of 480 Gb.
I find the average document to clock in at 7 Kb, in terms of raw HTML. In the index that's, dunno, probably less than 1 KB/doc.
How would this be limited to "initially"? Wouldn't it be a lot, initially, and then only get worse?
This is very true. How many times have I clicked on a site met with ads so bad that the browser slows down, and after 10 seconds the page gets covered up by more and more crap and then a paywall shows up sometimes too. Now here's the thing - a competitor to Google might detect you clicking back and then pop-up a special set of controls near the search result that lets you say: "too many ads" or "paywall".
However, if such an engine were to start beating Google, I'm sure Google would implement it in their own way: automatically detect why you clicked back in such a short timespan.
Perhaps ML will help in detecting such campaigns.
Do you seriously believe that Google doesn't use that as a datapoint already?
God please no. YouTube premium shows what Google would do, i.e., they would further ruin the free experience by ramping up the amount of ads you see to "incentivize" the premium search.
Is there any particular reason why internet search has to have a distorting gatekeeper to the global commons (that pretends playing Maxwell's demon). For chrissake, the stuff being indexed is public.
People care about UX not about technology remember that unless people are willing to sacrifice good UX in order to have greater security and privacy. These things are tricky and there is no right formula.
For a couple of years, Blekko ran a "3 card monte" game where we white listed the results from Google, Bing, and our own index. For every "contested" query, Blekko consistently beat the others by a significant margin. If the query wasn't contested, Bing and Google did about the same, and if the query was obscure, typically Google did better than Bing or Blekko.
What is a "contested" query? That is one where there is a lot of money on the line. My favorite one was "best credit card" (which is search engine shorthand for "What is the best credit card?" because the stop words "What", "is", and "the" are removed).
Why is it contested? Because if you put an advertisement into the results of that query, and the person making it clicked on that link and signed up for a credit card, you could be paid $50 or more. For a single click. Other queries that advertisers would pay well for getting the traffic of the user were, car dealerships, hotel chains, jewelry retailers, and university "referral" services (like the one that was busted for getting people into Ivy League schools by faking academic records).
Extremely few people click on an ad put onto a page of search results for the query "what is shoe rubber made of?"[1]. However it is required to serve queries like that so that people will come back when they are looking to spend money on something.
So using the same exact idea that Paul proposed Blekko built an English language index which allowed you to curate the crap out of your search results and return much better data. The "value" of that was not considered to be high enough to insist on people logging in to use the engine. Knowing an id for the person making the query allowed for user specific blacklists of spammers (so if for example you never wanted to see a Pinterest link in your results you could make that happen).
Without sufficient traffic, using the feedback loop "of these documents, which one was clicked as the 'best' answer?" type algorithms for ranking fail to converge rapidly enough for decent ranking.
Without a credible threat that if your site is not included in the index, your traffic will be greatly reduced, it is difficult to negotiate with web sites to permit crawling, rather than deny your crawls with the robots.txt file.
Blekko's best customers and most ardent fans? Reference Librarians. Yup, people who needed web search to do their jobs, not to find the movie times for the latest feature. Blekko never did try to create a subscription service, but I think such a service that is somewhere between free and the $$$ of LexisNexis has a shot, at least as a lifestyle business. You still need to get rights to the data and that gets harder and harder.
[1] Okay, bots do, but humans don't
IF anyone wants to see a demo please email me.
Is it viable competition?
PageRank seemed to borrow from the concept of citation count. The idea that "importance" could be measured by the number of times a webpage, like a paper published in a peer-reviewed academic journal, was referenced by other webpages, like other papers published in peer-reviewed academic journals. The initial name for the project before "Google" was "Backrub", referring to the reliance on "backlinks" to quantify importance.
An index of a commercially-oriented www full of sites supported by online advertising is nothing like Web of Science or some other database collection that allows ranking by citation count. The www has no peer-review and no limits on commercial activity.
Google succeeded in creating something highly profitable and sometimes useful, but the founders never delivered on their original promise. That was a search engine in the academic realm, where the technical details were public, and one that would be free from the influence of advertising.^1 Instead the project was turned into an online advertising business. A 180-degree pivot.
The moral/ethical debate went from the question of being advertising-supported to the question of invading the personal privacy of users, for the benefit of advertising. Whatever ideals the founders held in 1998 were overtaken by the lure of pure financial success. Once oppposed to idea of using cookies for advertising purposes, the founders were persuaded to purchase DoubleClick, ground zero for the explosion of online ads, for $3.1 bilion. Not sure what if any moral/ethical debate remains today. While the company is being sued simultaneously by hundreds of plaintiffs, including the US government, one of the founders is "hiding out" on a small island in the South Pacific. Whatever motivations he had to make an open, academic search engine free from the influence of advertising, they seem to be gone.
In sum, the world still needs a decent web search engine free from the influence of online advertising.
1. https://infolab.stanford.edu/~backrub/google.html
Excerpts:
"Up until now most search engine development has gone on at companies with little publication of technical details. This causes search engine technology to remain largely a black art and to be advertising oriented (see Appendix A). With Google, we have a strong goal to push more development and understanding into the academic realm.
Appendix A: Advertising and Mixed Motives
Currently, the predominant business model for commercial search engines is advertising. The goals of the advertising business model do not always correspond to providing quality search to users.
For this type of reason and historical experience with other media [Bagdikian 83], we expect that advertising funded search engines will be inherently biased towards the advertisers and away from the needs of the consumers.
Furthermore, advertising income often provides an incentive to provide poor quality search results.
[T]here will always be money from advertisers who want a customer to switch products, or have something that is genuinely new. But we believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search engine that is transparent and in the academic realm."
PageRank seemed to borrow from the concept of citation count. The idea that "importance" could be measured by the number of times a webpage, like a paper published in a peer-reviewed academic journal, was referenced by other webpages, like other papers published in peer-reviewed academic journals. The initial name for the project before "Google" was "Backrub", referring to the reliance on "backlinks" to quantify importance.
An index of a commercially-oriented www full of sites supported by online advertising is nothing like Web of Science or some other database collection that allows ranking by citation count. The www has no peer-review and no limits on commercial activity.
Google succeeded in creating something highly profitable and sometimes useful, but the founders never delivered on their original promise. That was a search engine in the academic realm, where the technical details were public, and one that would be free from the influence of advertising.^1 Instead the project was turned into an online advertising business. A 180-degree pivot.
The moral/ethical debate went from the question of being advertising-supported to the question of invading the personal privacy of users, for the benefit of advertising. Whatever ideas the founders held in 1998 regarding the influence of advertising on web search were overtaken by the lure of pure financial success. Once oppposed to idea of using cookies for advertising purposes, the founders were persuaded to purchase DoubleClick, a company with a terrible privacy record that uses cookies and purchasing data to profile users as ad targets,^2 for almost double what they paid for YouTube. Not sure what if any moral/ethical debate remains today. While the company is being sued simultaneously by hundreds of plaintiffs, including the US government, one of the founders is "hiding out" on a small island in the South Pacific. Whatever motivations he had to make an open, academic search engine free from the influence of advertising, they seem to be gone.
In sum, the world still needs a decent web search engine free from the influence of online advertising.
1. https://infolab.stanford.edu/~backrub/google.html
Excerpts:
"Up until now most search engine development has gone on at companies with little publication of technical details. This causes search engine technology to remain largely a black art and to be advertising oriented (see Appendix A). With Google, we have a strong goal to push more development and understanding into the academic realm.
Appendix A: Advertising and Mixed Motives
Currently, the predominant business model for commercial search engines is advertising. The goals of the advertising business model do not always correspond to providing quality search to users.
For this type of reason and historical experience with other media [Bagdikian 83], we expect that advertising funded search engines will be inherently biased towards the advertisers and away from the needs of the consumers.
Furthermore, advertising income often provides an incentive to provide poor quality search results.
[T]here will always be money from advertisers who want a customer to switch products, or have something that is genuinely new. But we believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search engine that is transparent and in the academic realm."
2. https://www.nytimes.com/2000/02/17/technology/us-investigati...
https://slate.com/technology/2005/11/why-web-surfers-love-to...
I never considered the possibility that an incubator would support a specific product, then later on call for alternatives that would essentially freeze out the original product that they supported. I'm sure this very rarely happens, but it's interesting to see a real-world example in action.
I view it as something like the rich folks who call for additional taxation of the rich. They're not going to just pay extra money that they don't have to under the current tax rules, both because it's not particularly fair and because one person paying extra taxes, even if they're very wealthy, isn't going to make a big impact. That doesn't mean they can't lobby to change the rules and be totally fine with it if everyone is paying additional taxes.