The world needs a non-profit search engine
daoudclarke.net
daoudclarke.net
You conflate two meanings of value: monetary value, and intrinsic value. Search engines are intrinsically but not monetarily valuable to users. Search engines are monetarily, but not intrinsically, valuable to advertisers. You can get into trouble when you conflate these two meaning of "value".
In fact, right here is the pivot on which the internet goes from an idealistic shang-ri-la for geeks, to a commercial hellscape for the unwashed masses. It is surprisingly easy to create intrinsic value with computers! You see it all day every day on HN: some geek had a thought, spends a weekend making it, and then deploys a solution.
It is surprisingly hard to extract monetary value from an intrinsically valuable solution. In fact, I believe that creating artificial scarcity is the hardest part of building an internet business, requiring invention on par with the intrinsically valuable part - and yet its the very thing that idealists rail against.
(And making something artificially scarce does seem morally repugnant. And yet I don't see any other way to pay developers. Full stop. Open source software + consulting fees is a good way to go, but that can't apply to hosted search for the public. Well I guess it could, you could teach businesses how to game your own engine!)
This is like saying that our entire economic system is backward ...
Even your "sell water on a hot day" idea probably won't sell a lot if you set up shop right next to an enormous promo stand from a global bottled water company that gives away bottled water for free. (And due to the magic of the internet, every spot you pick is right next to a global competitor with deep pockets)
A better way to deal with this is to discover what the intrinsic value for the product is, allow the creator to give it away, and then subsidize the creator the value they generated.
The nice thing with this system is that we can transition to it really easily. There are already many Open Source projects that exist, all we need to do is ear-mark a certain sized pot, figure out the weights of existing products, and hand it out. And as the economy gets stronger, allow the pot to grow, and eventually there might not be a thing as closed software, as open software will always be able to generate more value than closed.
Obviously the problem is that those in charge of giving away money are very passive, and in fact put up large barriers (the onerous application process) for anyone wanting the money. This, to me, is absurd. The admins of such a fund should themselves be active in one or more areas of interest, such that they themselves should be approaching people like Justine and telling her, "Here, please take my money."
I don't know much about grants, but like most things that are broken I suspect there are perverse incentives here. For example, someone is probably investing that grant money, and if they give it away the money leaves the market, and that will make the admins sad. (I don't know if that's the case, but that's the kind of thing I would expect to see.)
Oh, another surprisingly effective thing people are doing are subscriptions to people you like. You see this particularly on twitch, and sometimes on youtube. People can and will give money to the creators they like! It's kind of amazing to me. I suppose the analog for people like Justine is Patreon. It's fascinating though that people will happily subscribe to an entertaining personality, but will happily ignore someone who toils at making incredible things behind closed doors (but still use their tools). The obvious solution is that all open source projects should become a source of entertainment.
Instead, you should seek to discover the intrinsic value as it is occurring, and therefore no risk to the granters. I believe it is possible to do so through a modified Vickrey auction. A Vickrey auction is one in which the winners pays the price the second place bidder placed. In this modified system, the top 90%(or some other optimizing number we can work out) of bidders win, and they all pay the price of the highest non-winner. A certain subset users of the Open Source software are asked to give their true value for some sub-set of open source products as they use it and using statistics we could extrapolate that to the rest of the population.
We would then use these weights to give away all grant money. This might mean that Firefox can be funded exclusively through this pot, and they would no longer be beholden to Google for default Search. It would mean the grant writing process would go away. It would mean, as ads make products less valuable, that ads go away.
So, fund people based on github stars. I could think of worse solutions, honestly!
Oh, something else I remembered: Princeton's Institute for Advanced Study! This is an interesting idea. Basically they tell a scientist "You've changed human history, the least we can do is give you an income for the rest of your life that has no strings attached. You can teach, or not, research, or not." It would be cool if each of FAANG had one, or they pooled together to make one, or had a virtual distributed version. "Dear Fabrice Bellard, we will be depositing 10k euros a month into your account until you die. Yours, Google."
Google provides unlimited searches to users without charge. On the other side, google provides human attention to advertisers, and charges for it due to the scarce nature of attention. Sort of like a marketplace business.
I”d like to offer a different view, one that thousands of subscribers at Kagi search can hopefully stand behind too. (Kagi founder here)
Searching has a monetary value and a cost, the only question is whether the user is paying for it or a third party is paying for the user (a cringy thought when you think about it).
This question is answered by the very business model of the search engine, which determines who its customer is. It can either be its users (like in the case of Kagi and mwmbl) or advertisers (like most other search engine).
Although it is really hard to break through the habit of getting search for "free", we are at least happy to be able to offer this choice to the consumer today. I am expecting to see many more paid search engines in the future.
Based on what I’m seeing in google search quality… their days of unquestioned dominance died a few years ago. The king is dead, the prince (bing) is on vacation somewhere, and the crown is up for grabs.
The people who build stuff with open source are very smart, so they usually don't need foss devs as consultants to explain to them how to do it. I make some of the most popular projects on this website. Been doing it for years. I get plenty of donations because my work makes people so happy. But no one has ever offered to pay me for something in return, since I don't have anything with any economic value, other than me myself. So I get plenty of job offers from people who would love to be able to say I'm their employee. Since controlling people is about as valuable as controlling the services people need.
It'd be nice if our cultural mythology about earning your keep doing an honest day's work through fair trade was how the system worked, rather than me needing to depend on the gift economy that's funded by control. But I think there was just so much abuse of the classic models of human cooperation that they just couldn't transition to the digital era, and as such we need to find a way to adapt.
Programming reminds me of other professions with high inputs but low per-unit-cost outputs: teaching, music, movies, art, journalism, etc. -- basically anything that is create-once, share-often. All those (for the most part) are things that America's shortsightedly capitalist economy fail to adequately incentivize/reward unless you happen to become a celebrity.
There is already, and will continue to be, a class of human labor output that is intrinsically valuable but which our economy is unable to adequately price. I'd argue that's more an issue with the speculation-driven economy that we have than with the labor in question. It'll only get more drastic we we automate more and more and further amplify human creativity.
It'd be cool to see a nonprofit search engine/email service/office suite/whatever funded in a similar style as NPR, but of course that'd run into political issues at every level.
In an utopian cyberpunk future, what if there were multiple voluntary nonprofit "shadow governments" that you can choose to tithe every month, almost like churches? You can choose between e-governments red, blue, green, yellow, purple, gray, whatever... give them 1% of your income a year in exchange for a suite of services run and staffed by professionals who are salaried but own no equity; they work as a form of civil service, not a wealth-building scheme. Syndicates for the public good, I suppose. Lol, in all the video games, these usually turn into private military companies with killer androids, but what if they just, uh, provided really good email (and automatic online driver's license renewals) instead?
Such systems would probably never be able to attract the best talent (unless they turn into something like Mozilla, which is a big enterprise masquerading as a nonprofit), but you often don't NEED the best. Wikipedia, the sum total of human knowledge, is also the sum total of human mediocrity (with plenty of redundant efforts, infighting, unoptimized problems, etc.). But having stable collaborative communities is something that is critical for producing works of intrinsic value, and developing slow-trickle funding streams for such communities -- the kind that can sustain without turning them into potential get-rich-quick schemes, or subject them to violent boom-bust cycles -- is what allows them to both keep working effectively, AND keeps away the exploitative get-rich-quick types looking to subvert public labor for personal gain.
We need a funding model that provides enough income to attract people who want to do it for the public good, but not so much that it also attracts people who want to turn it into personal wealth. In high-input, low-per-unit-cost services, optimizing for maximum (as opposed to sufficient) monetization will too often mean that the users themselves become the product, as we see time and again with Google, Facebook, Twitter, and basically the entire modern online economy. It's the difference between Amazon and your local library, Oracle and Postgres, EvilOS and Linux, etc.
Things don't have to be artificially scarce if we can learn to ask, "How can I make this as widely accessible as possible while ensuring I have my own basics needs covered?" instead of "How can I make this as monetarily valuable as possible?"
There are always going to be people who want to just do things to make the world a better place, just as surely as there are always going to be people who want to optimize for personal profit. We have every system to fund the latter, but not so much the former right now. Is the for-profit business model the only, or best, way?
Right now it is extremely difficult to build your own web crawler that would compete with Google. And that is not because of the technology, but because multiple sites will prevent your bot from accessing them if you're not Google or Bing - either through robots.txt, or through directly banning your IP if it's trying to crawl and it's not a confirmed google-bot.
Having a non-profit, open source, crawler that keeps an up to date index (or web cache) of the web would help competition spring up.
That's a excellent idea! In the spirit of open-data, and people can do with it what they want.
i am probably missing something but can you give an example where this happens?
(I don't know if this happens for this specific example, but Google does this for some searches)
Part of why sites participate in the infobox program is that in practice you do get quite a lot of hits from it: many people click through to see the answer in context.
A question to ask could be: how often do users care about information from a few minutes ago, compared to information that has been available for a longer duration of time?
- a few thousand news-sites (like nyt.com, bbc.co.uk),
- a few thousand very popular blogs (based on what influencers people search for),
- a handful of social media sites (e.g. Twitter),
- a few hundred databases in areas like weather, airlines, sports (like ATP for people who look for Wimbledon results today)?
Nah, not a small task but you can break it down into well understood problems that have known solution.
The hard part is ranking everything.
If from experience - how did you get around multiple sites disallowing crawlers other than google or bing?
I don't recursively call links found in the pages. I expect the user to give me the URLs to crawl and save.
In order to "find" new content, I let the user specify where they want to search for things the engine hasn't "crawled" yet. So, a search for scooters to buy might end up searching Amazon directly, then lets the user "save" the site by passing the Amazon URL for the scooter they like to the system for imaging.
I use GPT-3 or other ML models to do some of the heavy lifting for adding labels to the pages or documents the user uploads.
This ends up being a "curated" list of documents important to the individual user, not an exhaustive crawl of all things which are important to all users.
Wouldn't Google just use this too? Which would give Google in greater dominance over alternatives...
Regardless of search engine design, there's HUGE money in SEO. Any successful search engine will be gamed. Do you have the developer power to go red-queen against all the large companies in the world?
A search engine like Google isn't just a search engine as the author describes it. It is a very integral part of the economy of the internet and just labelling a simplistic interpretation of the present state as "Evil" with an academically poor write up of what a viable alternative is does little good.
As google is now a turd, not just no longer capable of delivering this service but actively destroying the good part of the web by refusing to index it.
It is my attention, it doesn't belong to anyone else.
My access to information and educated opinion is a far more integral part of the greater economy.
Google is like a screaming man at a town meeting making sure no one else can get a word in. The meeting is now pointless.
Your point about wanting to read articles written for you by others is certainly possible. The very fact that such a desirable outcome drives you to Google and nowhere else should suggest the complexity of the problem they’re solving and how there isn’t really anything else out there doing so well.
>Ideally as few as possible should profit from this process.
Why?
When you accepted a job offer in the software industry, did you stipulate that your mission is to write code for your employer and you will be charging as little as possible for that privilege? Minimum wage should get you by just fine, right?
I hate fully grown adults behaving as though anyone except them making a profit is somehow evil.
Yes and I'm not impressed.
> > Ideally as few as possible should profit from this process.
> Why?
> When you accepted a job offer in the software industry, did you stipulate that your mission is to write code for your employer and you will be charging as little as possible for that privilege? Minimum wage should get you by just fine, right?
I'm not a good example as I indeed live wonderfully on minimum wage and write software for free.
> I hate fully grown adults behaving as though anyone except them making a profit is somehow evil.
Don't worry, my philosophy is not that superficial. We have people who make things, people who organize the making of things and people who organize the things made.
It can be true that the meta data is more valuable than the data it self and organizing an effort can be much more intense than any of the tasks involved. But lets not pretend that is always the case.
Before money and before the written word we had the exchange of thoughts, observations and ideas. I believe this to be somewhat like the foundation on which everything else we did is build. I want to see this process benefit from technology.
You wrote your comment perhaps a bit limited by the ropes of the platform but sincerely, free from any agenda, you wrote pretty much what you think.
Now if we [beyond HN] add additional layers of agendas between our exchange, each interested in maximizing their profit from it perhaps not you but many others will resort to self-moderation.
You wont be able to state it simply like: "I hate fully grown adults behaving as though anyone except them making a profit is somehow evil."
It could become something like "I don't understand why some people don't like others making money" stripped from how strong you feel about the subject. You could also chose not to say anything.
At that point we are messing with the very fabric of our collective reality.
If I had to chose between freely communicating and the economy it wouldn't be a hard choice.
That doesn't necessarily make it right. Online scamming also supports a lot of people's livelihoods.
I’d expect you to know the difference.
Do you think a Wikipedia style team of human reviewers with upvote/downvote capabilities can help reduce the SEO spam?
Obviously it’s hard to review the reviewers as well but Wikipedia seems to have done that.
You know how it works here. We would strip you of your internet points if you start being nonsensical.
The hn machine is not all that smart and can easily be gamed, and would, if the stakes were to become high enough. Farming hn karma by making pointed statements on crowd favorites (parent named two, oss licensing or privacy also spring to mind) gets you your 10k, no originality or honesty required, in no time.
What's protecting hn is a lot of moderation + relative irrelevance. If those 10k were to systematically bring you enough eyes (by driving search results), you are in effect printing money. There is no reason to assume the number of people doing it would not scale with the return attached to doing it.
Though I do not know how the initial allotment of karma/points could be distributed for the pioneers and for the new growing community, maybe allot 'n' points for new user after a year...
Playing defense is exhausting when playing offense is extremely cheap.
HN is quite a large crowd, but not extremely big. An attacker must penetrate a lot of small crowds to be successful.
I bet the (niche) product as such wouldn't be as hard to build as it would be to scale. Imagine every user constantly tweaking (directly or indirectly) their search result settings, and having that impact millions (or more) indexed items, for every user.
Interestingly summarized on a Lifehacker some time ago. https://lifehacker.com/the-psychology-of-a-fanboy-why-you-ke...
I like my friends but I don't want to get my news or politics or shopping advice from them.
https://searx.be/search?q=what%20camera%20would%20HN%20recom...
Results 1-3: SEO "best cameras" from DDG, Qwant
Result 4: Ask HN from DDG, Qwant
Result 5: Ask HN from Google
Results 6-8: SEO "best cameras" from DDG, Qwant
Result 9: Ask HN from Google
Result 10: Camera forum from DDG, Qwant
Edit: if you're not familiar with SearX, everyone who visits will get a slightly different result based on dynamic results. Even if the same person refreshes a few times the exact results and ordering will vary; I've learned to try the same search a few times to get better results, it's just a quirk of how it works and how each remote engine reacts at that given instant.
See: http://getoutfoxed.com/node/46
Some approach like this could still work, but it’s incredibly hard to maintain/define the right “network”, and across different domains. (E.g. HN probably not so good a community for latest fashions or sports trivia)
Have 4 different types of list.
Whitelist - highest scoring
Not listed - these websites are not ranked.
Yellowlist - show but keep in a separate column
Blacklist - don't show
https://search.brave.com/help/goggles is interesting.
Still waiting for Mastercard to change their incredibly ‘racist’ company name. /s
It could be useful for power users, though.
99%+ of people in the tech industry currently do not care to do the extra steps required for data neutrality, and privacy.
99%+ of people are lazy to the point of harming themselves and others.
1%- of people examine how the 99%+ do things and pioneer harm reduction tactics in spite of everyone constantly reminding them that no one wants their help.
this isn't true. 170 years ago, maybe. handwashing became a thing in the late 1800s after Semmelweiss and Pasteur
SEO is basically another form of advertising, so evidence suggests to me that people would use something like this.
[1] https://en.wikipedia.org/wiki/Ad_blocking#:~:text=users
[2] https://earthweb.com/how-many-people-use-ad-blockers/
[3] https://www.insiderintelligence.com/insights/ad-blocking
I believe the addon page is average daily users over a week or 6 days?[2]
Firefox monthly active users (unique over 28 days) is around 200 to 210 million worldwide (26 to 22 million in the USA)[3].
It's probably wrong to compare those numbers but a very naive 200 / 28 = 7.41 million average unique users per day. 5 / 7.41 = 67% (probably wrong)
[1] https://addons.mozilla.org/blog/firefoxs-most-popular-innova...
Make a customizable and privacy-respecting search for us, power users.
Also, I have noticed that Firefox internal search works good for sites you had visited. So when I want to visit a page I have seen earlier, I can go straight to it skipping Google.
Also, you can click on any search box and add it with a prefix, so that you can search MDN or Wikipedia directly, again, without informing Google.
[1] https://source.chromium.org/chromium/chromium/src/+/main:com...
It’s worth the experiment.
Wikipedia shows it can work at scale.
[1] https://static.googleusercontent.com/media/guidelines.raterh...
Simply put, search engines have been at war with SEO for over 30-years, which has significantly raised the bar not only being a search engine, but producing content; not to mention knowing how to search for information. With the introduction of machine generated content, information wars between countries, global dependence of online commerce & information, etc — the speed of change shows no signs of letting up.
In my opinion, for the average person, knowing how to search for information is the real issue, not that the quality of information available has become worse or that Google has become a worse search engine. If anything, Google has reduced its advantage search capabilities not for financial gain, but because average user is just too lazy to learn how to search and keep up with changes required to continue to be an advanced searcher.
But also yes the users and the UI both fail. When I used to search for something I would type in something like “gutter clog clean” but slowly started noticing that Google likes longer sentences like “how do I clean a clog in my gutters?”. In pursuit of making Knowsmore (from Ralph Breaks the Internet), Google lost the power user features. Search would be infinitely better if they actually fucking respected literal mode and stopped trying to treat me like an idiot with no attention span. Having search results that contain one out of like 8 words in my query and asking me if I want to include others and then when I say I do still showing me results without them is broken UI and not a user problem.
I agree, which leads me to the conclusion that subscription is the best way to avoid this conflict of interest. Unfortunately, most of the world won't subscribe to a search engine, and doesn't seem to mind ads - to a degree. With Google looking more and more like AltaVista before its demise (to Google), my conclusion is that Google will strangle itself out of existence and give way for the next "new, streamlined, not-full-of-ads" competitor.
There have been lots of no ads (for now) attempts. DDG had like one small ad at one point. But people didn’t leave in droves. It’s almost like people are ok with ads.
In the 20-30 searches that I do in a day, I still have to google about half of them. Either because it's stuff Google does well (currency conversion, for example), or Kagi just doesn't get what I'm trying to search.
I remember starting out with the Internet searching on Altavista and Yahoo and Lycos. The information that was present was nowhere near as now, and it was more "exploratory". Nowadays people just kind of know what they want and just wants to quickly get there.
Currency conversion is not technically a search. It is question answering and Kagi capabilities are still being built. Google only has a 20 year headstart. Can you report all such cases to kagifeedback.org so they are on our radar?
The other examples are a bit harder to describe and I can't quite describe how Google gets it right. I think I might need more time to describe it out, as it involves search in another language.
Use Verbatim?
Beyond that, complaining Google does not do XYZ misses the point. Google is a search engine designed for the average user and the average user does not want verbatim search. They also do not want: advanced search operators, true Boolean search, regular expressions, API access to search, open source code, real-time streams of pages Google’s crawling, etc.
What they do want and always have is natural language based searches in there language of preference with clarifying responses from the search engine in natural language; that is, they want to treat a search engine like a person and be treated like a person; which was odd that they referenced Knowsmore, since Knowsmore [1] used keyword based searches, not plain language searches.
Google is not the primary problem, the average user is the issue. Unless people realize that — they’re fighting in a war they do not even understand.
To make it even more clear, Google is easily able to detect and block users blocking ADs, but they do not. More than 60% of users still don’t block ADs; not because they love ADs, but because effort to figure it out simply is not worth it to them, they like ADs, etc.
I agree with you but Google is not yet at that point where it can act and serve people like an Answer Machine that knows everything; both the people's preferences and the perfect answers.
>Google is not the primary problem, the average user is the issue. Unless people realize that — they’re fighting in a war they do not even understand.
Again I agree that casual users are the problem but how we can help them? This is the The Innovator's Dilemma[0] where if we ask casual users what new stuff they want from Google Search, they will answer "nothing". Because even they themselves don't know how their UX can be or should be improved and on top of that they are satisfied with Google's mediocrity. They would just respond "Google is Google".
>Beyond that, complaining Google does not do XYZ misses the point. Google is a search engine designed for the average user and the average user does not want verbatim search. They also do not want: advanced search operators, true Boolean search, regular expressions, API access to search, open source code, real-time streams of pages Google’s crawling, etc.
Complexity of constructing "complex" search queries needs to be simplified so casual users can use such features and queries.
As for the parsing of queries, that's probably based on how most users use search. Not everyone is familiar with keyword -based search. I expect they've done tons of A/B tests to determine what kind of query interpretation makes most users get better results. We're just not "most users".
This not only means you exert massive selection pressure on the shape of websites. The SEO spammers don't need to be good or know what they are doing, they just need to be lucky once. If they get it right, they float to the top, and can iterate on that design. This effectively is saying that no matter how secret or smart your algorithms are, it doesn't matter if you're in Google's position. The numbers are stacked against you.
To make matters worse, any company with that sort of a market share has serious handcuffs in how heavy handed and "unfair" you can be without risking litigation for anti-competitive practices.
I think the best thing that could happen for Google is ironically serious competition in the search market. This would help both problems at once.
Only because it's profitable for them to allow it to be gamed, like all the spam sites now when you search for SO, Google allows them to be ranked because they're filled with Google ads. But it'd be trivial to just delist them all, that'd be beneficial for the user but not the search engine.
It's not a matter of 'developer power' just flip a boolean somewhere and delist the site.
https://www.ted.com/talks/eli_pariser_beware_online_filter_b...
The last link someone clicks on, is usually the right link to associate with the search term for future users. Its not hard to make a search engine!
I'm not a Robot!
Let's try? It's also an interesting research topic in itself and might be a topic for academic research. At the moment Google is a black box and their incentives are not really aligned to stop SEO. It's good for them to show more ads, it's good for them to show you copycat pages of github/stackoverflow with ads. Not saying that Google is doing this on purpose - I doubt it - but we don't know. It's surely possible to create an index and ranking that prefers different things than Google.
Let's try something at least. It will be gamed, will be worse probably but it's open and can be a playground for academic research.
Best that can happen is that there are ways to for a better ranking and Google was dishonest to maximize profit. If it's still gamed at least the mechanics can be studied and analysed and maybe someone can figure out a switch like 'be unfair in ways Google can't' - would love a 'no ads on the page' switch that would probably solve quite a few problems.
Trust our black-box and you won't have enough devs for this is just a bad answer for such an important problem. The amount of stupid simple redirection spam in my results in the last few years also looks like that Google just doesn't care alot about this anymore.
While they're not wrong about how the way Google determines ranking has its issues, this way has its own set of problems. If you explicitly use user ratings as part of your rankings in some way, people can punish sites they don't like, ala review bombing on Yelp, Steam, etc.
Not saying it's necessarily a bad idea because of that, but I hope they don't fall victim to the mentality of, "let's just trust the users" as an ironclad rule, because that doesn't always work out well.
This particular example might be about the Republican party-wide support of echo chambers and offline mob behavior that led to an invasion of the building containing politicians certifying the vote of a newly elected leader and fueling disinformation about elections to weaken the trust and integrity of the system?
The moment you have community rankings on search, and your search gets popular, you land in a war zone with bots trying to mangle those. Reddit is kind of good dealing with that, but it is very resource intensive.
However it could just be reputation based like on Wikipedia
I'm sure you could still manage to make "fake" accounts but it would be much more difficult, and linking them together would be much easier.
Of course starting a site like this would be very difficult. But maybe you could start without it then add it in once you get to a decent popularity such that many people can find a referral if they need to.
In China, all social accounts must be associated with a phone number, and phone numbers are tied to government identities. It doesn't stop any manipulation of scores and rankings.
> then keep track of their credibility?
It is very likely China will do that too soon. I think you can already imagine the ramifications.
source: worked in ads and search for decades, incl google.
i have a guy in digital marketing tell me that his friend does SEO and he does wonders with obscure keywords and shit. that friend is a freelancer and earns a good payday.
When you want to insert your brand in every fucking imaginative keyword as opposed to people "searching for something",
why does internet advertising revolve around everyone assuming every person googling something "WANTS TO BUY SOMETHING"?
The truely garbage content is produced as cheaply as possible (scraped, generated from a data source or generated via “ai”) to capture advertising revenue, often via sub prime advertising networks (or a number of middleman networks).
But to your point, not everyone wants to buy something, and not everyone needs to.
Much of the content out there is simply trying to capture your attention and make you available to some of the worst advertising and ad networks (read scams, lead gen, fake buttons, affiliate crap).
Because people googling SOMETHING are more likely to buy SOMETHING than people googling SOMETHING_ELSE.
Rather I think it’s because everyone who buys ads has something to sell.
Search ads are “direct action”. You click a link to do something. Ads on eg. TV are more about “brand memory” - reminding you they exist. When you watch tv you’re passively taking in information, but when you’re searching you’re actively trying to click something already. It’s a better fit behaviorally.
Instead of writing facebook.com in the URL-bar, they search for facebook and click the first link...
https://www.abc.net.au/news/2022-06-21/scammers-using-text-m...
https://www.channelnewsasia.com/singapore/scam-bank-google-a...
I think the only solution is education.
Perhaps it has improved recently, but it used to be a plague in crypto - people getting ads for phishing sites instead of legitimate ones, losing money, and Google being unresponsive to reports.
The reasons this is actually potentially bad are pretty deep in the internet wonk weeds, where you get into questions of gatekeepers and provinance of information and it shouldn't be surprising most people don't care about those things: those of us who do have failed to provide them with better tools.
On some level it's a little like saying "my dad sent me an email and he didn't use pgp! Can you believe it!??"
The results are fine for me in day to day usage and I find that Google will not provide me with better results if I cant find it on DDG.
Would it be better if they had their own indexer? Maybe...
You use a proxied Bing. If you want a proxied Google (what I prefer), you can use https://startpage.com
Difference to Google is how they position themselves in regards to privacy, and that Google actually built a search engine. Both make their money by providing ad space.
I did no such thing but looked around and found the company had been sold to an advertiser/tracker. It used to work much, much better than ddg:(
And this is educational content, text only, no ads or popups, no SEO hacking. Bing's analysis tool told me only that I was missing the "lang" attribute from my HTML tag. So I added it, but of course that wasn't the issue.
I reached out to them, and they replied saying that the page didn't meet the requirements for listing, but didn't elaborate.
It certainly makes me wonder what content their broken algorithm is missing.
And it sucks because it means DDG is missing that content, too.
This is especially dangerous because it propagates an illusion that there's dozens of engines to choose from. The reality is these three companies control more and more of humanity's ingress to information, censoring what they see fit for political/financial gain.
My own personal website for example is not even listed in the search engine I use. Microsoft support see the issue but have no explanation and escalate and quote “quality requirements” for weeks now. A human review yielded positive results regarding quality. Meanwhile, I see lots of SEO spam on DDG when searching for generic technical terms.
Not so in Bing, it's simply not there. The page exists since March 2021. Bing Webmaster Tools reads "Discovered but not crawled. URL cannot appear on Bing", giving no further reason. Also: "Last crawl attempted 01 Feb 2022 at 19:35", which means that Bing did not bother to retry for months, despite me submitting it manually on a regular basis. Clicking the "Live URL" tab results in entirely green checkmarks along with "URL can be indexed by Bing".
Another example would be my personal site: https://herrbischoff.com. Same issue. That one is listed on Google for more than 10 years.
My working theory is that Bing’s selection algorithm is biased towards large and already popular sites. In the server logs, I don’t see Bing even attempting to crawl the sites I mentioned, except requesting robots.txt and the root page. Bing appears to be excruciatingly slow to update anything but high traffic sites.
Again, Microsoft Support was unable to explain this behavior even after manual, human review found everything to be in order.
I tried deleting robots.txt entirely and got only Chinese crawlers and SEO bots, but still no Bing crawl. All organic traffic comes from blogs linking directly and Google.
—————
Thank you for your patience!
After further review, it appears that your site < http://herrbischoff.com https://ipbl.herrbischoff.com/> did not meet the standards set by Bing the last time it was crawled.
Bing constantly prioritizes the content to be indexed that will drive highest users satisfaction. Please follow Bing Webmaster Guidelines to better understand criteria for most valuable content.
They did that for a while without considering the obvious downside - the incentive to mass blacklist your competitors.
It's really useful!
If I were to make a search engine, I'd definitely give users more control over their results. Block crap sites, vote up your favourite sites, vote down questionable sites, maybe different context profiles, because if you're searching for Java in the context of vacation or news events you want different results than if you're searching for it in a programming context.
There's so much that search could do better than what Google is doing, but I'm not doing it because it's way too much work, and it requires serious resources to index everything.
Seriously, people, if you're writing about anything at all, making assumptions is always a bad idea. If you're writing about a product, make it more than easy to get to it. Provide plentiful CTA's (that's Calls To Action, defined so as not to make the same mistake of assumption) - links, bittons, a big banner at the top: ("I'm building a non profit search engine called Mwmbl! Find out more").
K, thanks, </ moan >
The tricky part, if you want people to use your search engine for more than the novelty factor, and what most Google competitors struggle with is drawing the rest of the damn owl. For example, commercial searches, local businesses, that sort of thing. As much as Google flounders with some queries, the overall package is still really good.
How are the batches of URLs to be crawled generated/discovered and posted at your API?
How do you deal with duplicate crawls?
And since you control what URLs need to be crawled, you protect yourself against rogue clients sending arbitrary URLs.
There certainly are a lot of elegant ways to reduce spam for this particular problem imo.
I'm not worried about the URLs, but the content of the URLs sent back.
Say the server tells a client to crawl a CNN article. The "hacked" client sends a fake CNN article back.
Some might, because of A/B testing or news updating, but even updating news will get a positive similar page and those that don't should probably fall into an exceptions category until it can be determined what can be done about it. Maybe a flag in the URL to give you a static page or just accept that it changes often enough that even faked pages won't last long?
There is a challenge for sites that serve different content based on GeoIP, A/B testing, dynamic content, etc. So some human review of the diff may help check for malice. If there's literally spam, human review would clearly detect this and that bot is distrusted.
Plus I can now cause you to have to run your own crawler anyway and either slow progress or cost you a lot of money.
You can buy a trustworthy residential IP for low cost, you can buy them in bulk in the thousands. All of them are real residential IPs from any ISP of your choosing in any country. You can rent Chrome browsers running over those IPs, directed via remote desktop and accessibility protocols (good luck banning that without running awful of anti-discrimination laws). You can do all that for under 1k$ a month for like 1 million clients.
My workplace has been at the other end of DDoS attacks directed by such services, best you can do is ban specific Chrome versions they use but that lasts until they update.
It's an uphill battle that you will loose in the long term if you rely on client trust.
I think in this specific case, the spammer is on poor footing. The spammer wants to inject specific content, ideally many times. With double processing of URLs and the spammer controls 50% of the clients then there's a 50% chance that a simple diff would show the injected spam. The problem is that the spammer needs to do this many times, so their injection becomes statistically apparent. If the spammer can only inject a small number of messages before they are detected, then the cost per injected spam will be quite high. Long running spam campaigns could eventually be detected by content analysis, so the spammer also needs to rotate content.
Obviously you can play with the numbers, the attacker could try to control >>50% of the clients. The project could process URLs >2x. The project could re-process N% of URLs on trusted hardware, etc. It's not easy by any means, but you can tune the knobs to increase the cost for spammers.
You make it sound easy. ;)
It changed a bit in the implementation.
I have a long standing bet that, at some point, some company will be "globalized" (operated under some common funding by many different countries, like many research projects or defense organization or aid funds, etc...), and the "search engine" part of google is the prime candidate.
That being said, I'm from Europe, so "sharing the cost of something useful" is not culturally untolerable.
Far fetched and controversial opinion, I know. We'll see.
But I do wonder if the national archives or the library of congress could be a good host for this sort of project. Not sure I agree it should be run by a government… most don’t have great histories when put in charge of gatekeeping access to information.
This is what really nails it for me.
There's far too much black box in pretty much every major search engine out there. Maybe it's by design "so that people can't game it". Even so, it's not working very well.
I'm excited for the next 10 years to see what we (humans) come up with to solve the state of the internet, because something's gonna give at some point.
The overwhelming vast majority of people just don’t care about the internet outside those major few silos, so as far as “humanity” is concerned the internet is working as intended.
It pisses me off what’s become of the internet, but I don’t personally see it changing.
If so, I don't agree. Your small minority's job is to deliver those alternatives, and to feed the flames while the rest of the world makes the transition. Which they will do because the next thing is clearly so much better.
Have more faith in most of humanity and doubt yourself and similar others for having failed so far to disrupt this industry with better technology.
We (humans)? Are the fish up to something? Never did trust fish, especially regarding SEO stuff.
Case in point: www.forkandspoonkitchen.org
The first search engine that provides community curation and manages to get most tech-savvy people on board, classifying the content for free, is going to reign in the upcoming decade as Google loses its grip.
I see the largest internet companies, including Facebook, Twitter, and also Google, fight spam and other harmful content only to the degree absolutely necessary to stay somewhat usable. Which makes sense because it's costly and does not generate profit.
I would expect a non-profit, however, to focus much more on fighting harmful content because it centers around the user experience, hence quality of the content.
I don't see a guarantee this works in practice, but the respective incentives seem clear.
When a for profit entity is more successful at fighting abuse, their users are happier and they sell more ads and so can devote more resources to fighting spam. When a nonprofit successfully fights spam, they don't get more resources, and the spammers upgrade their toolboxes, because they do have a profit incentive.
If you create a search engine where users can report spam and get some form of karma for valid reports that is shown in their social network, then it's quite likely that the users have enough momentum to get ahead of the spam.
There are very few incentives better than profit..This is going to be a US v/s USSR fight during the cold war..
https://en.m.wikipedia.org/wiki/Knowledge_Engine_(Wikimedia_...
Also worth noting that Google is a significant donor (and now enterprise customer) of Wikipedia, but unclear if this had any impact of Wikipedia’s choice not to continue the project.
There frequently non-profits that use excess funds to unnecessarily expand beyond the original mission, for example Wikipedia — or that pay staff, especially executives way beyond what most donors realize.
To me, being a non-profit is what it is, I don’t read too much into organization being a non-profit.
I searched for "Elephant," and I got a Wikipedia page about a specific Elephant statue at Coney Island, a UK elephant charity, and a blog post about Haskell ("the elephant in the room").
It's unfair to poke fun at a very small project that admits that it is far from done yet, but it's gotta figure out a way to crack the "which pages are most likely to be relevant" problem or else it's not going to be useful.
Amen. How many times I searched for an astronaut's name - or any other person with significance - and instead I got results for some pro sports guy.
I guess 'significance' is subjective and I should search more specific.
SEO and all it's results seem to be immanent to the system.
What you want to do is put the SEO monster in front of your cart and make it do useful work. You basically got an army of hard working people with money to burn who will do anything. What is there to complaint about?
(previous suggestion here: https://news.ycombinator.com/item?id=31585340)
If one users reputation is higher, there's redistribution, but algorithms would need to carefully weight reputation linearly in all decisions, making redistribution not advantageous, solving any issues with sybil attacks[1].
One might be that the typical user can sensibly elect a few individuals to trust -- it could be developers (which are a natural choice for trust), to activist and publicly visible individuals (even close friends). Then presumably you could adopt his trust model (such individuals could be roots in independent conservative trust webs/graphs). I think a very large number of such webs might be computationally expensive, but hopefully you'd be able to find someone you trust or start your own independent graph (if you trust no one, you'd effectively lose all anti-SEO measures I guess). This very naturally leads to a decentralized reputation system!
I believe you mean 'inherent'.
immanent:
adjective
existing or operating within; inherent.
"the protection of liberties is immanent in constitutional arrangements"
synonyms: inherent, intrinsic, innate, built-in, latent, essential, fundamental, basic, ingrained, natural
Great, detect the affiliate links and downrank them. Ban it completely if the problem is severe or the products are fake/etc.
> earn money with ads
Similar, detect them and downrank. If a particular ad network is knowingly malicious, ban it entirely.
Google’s name was so ubiquitous it became a verb, Duck Duck Go is a smart memorable name.
Mwmbl is a challenging product name, even if the .org domain name was available.
I was also confused by the name. That's not necessarily a good first impression.
- Access your data for all websites - Monitor extension usage and manage themes
Spam site operators have a huge incentive to get users to click on their links & provide them with ad-revenue
For Google, they could make things so much better by down-ranking sites that show ads
I believe one of the biggest impacts toward breaking up Google's monopoly on search is making them open up access to their index, even requiring Google to provide direct API search access for others to build alternative search products. They have a search API today, but it is prohibitively expensive to build on ($5/1000 calls).
I built a fairly popular search engine a couple years back, but the cost of Google's search API and increasing number of bot attacks make it difficult to reason keeping it online.
I kept forgetting about Kagi. I have a login for that.
Yep.com has a different model, I haven't read into it far enough to decide if they actually do as they say they'll do.
If you want to search through an incredibly limited % of the web then yeah it can be a solution, but even the lamest search engine company out there would outperform a GPT-2 like model running from your laptop.
Feels like that would be good info to share, once it's depersonalised.
It's usually easy enough to add the site, by simply writing its name, but it could be easier.
I'm thinking right now this:
- Compile a browser without the cross origin limitations
- Make a site that uses iframes with all the answer-providing websites in it
- Simply focus the text input in the site/iframe you want and search away
- Have a way to open the results in your main browser or just use that patched browser
Like one of those internet explorer toolbars, except they cover the whole area.
You could also build your own, most search engines have a specific pattern to how they encode the search term in the URL. Although, I suppose that doesn't support auto-complete
I still remember helping a friend finding informations on the accounting balance of Rome's, Italy, public transport, and finding the most relevant link buried deep at around page 20. The first 15 pages were almost completely news websites with completely irrelevant news to the search query but they would consistently rank much higher.
Google will have to reinvent itself or it will eventually destroy itself with negligence of its core business. There isn't yet critical mass of casual users who think Google sucks, all they think is that the Google is internet. That's their intellectual level.
Hopefully some day soon the internet will be searchable again.
Thanks to everyone involved in attempting to make this happen (preferably in a non-profit-maximized way).
(Said before at https://news.ycombinator.com/item?id=32034390)
Youtube frankly offers better monetisation and most importantly easier user retention via subscriptions.
I tend to think the quality of the content tends to be better when people don't think about stuff like user retention or subscriptions, but rather how it will actually reach people that care. Good search/curation is a key component for that.
Of course such a world free of implicit monetization will require it to be explicit (Patreon-style), but that should massively realign incentives.
I bet Google spends at least as much to combat that, and it's extremely hard to deal with while being open-source. It's useless to call for a non-profit search engine without tackling this very core issue.
Google's search results are so bad I can't really alledge incompetence here but have to wonder whether there's some different motivation. Maybe it's that low quality search results tend to be plastered with ads, which they get a cut on.
All the mainstream search engines' priority is to maximize ad clicks/impressions (or collect data to target future impressions), either directly on their own property, or indirectly when linking to websites that embed their ads.
There's no reason why they can't detect ads or analytics and use that as a negative ranking factor (so that all other factors being equal, a non-ad-infested result would rank higher than the ad-infested one), but this would go contrary to their business model.
> The paid subscription model > Donation funded, non-profit model
No! There is a 3rd! You could do a search app eco system where you leave the unlimited overly complicated puzzles a search engine could address as an exercise for the user.
I always have a bazillion ideas but couldn't think of a single good phone app before mobile phones. I mean, should I want my phone to be a gaming console? It seems ridiculous. Writing is writing books, all other kinds are watered down. Do I want to write books with an onscreen keyboard? It all sounded idiotic, nothing worth using.
But the idea you mention, typing an overly popular domain name without extension should take you to the website directly... What you are trying to say IMHO is CLI! Search is just the failback if the provided query/instruction doesn't make sense to any of the apps.
I cant think of many but there are no doubt thousands of activities that could benefit from an at least somewhat themed search engine. An app could be a biochemistry web directory that ranks results from a chosen sub folder above the normal results.
Any FOSS or other company could create a web dir tree with the few or many pages about it self. A check box lets you pick the ones you want to query. Normal results go under those results. The biochem wont bother you when searching for pokemon.
People love my stores. What they really want is to see illustrated results from my inventory above all other results. Uncheck the box if you are not in the mood. (edit: I'm joking of course but I do have a good fews shopping apps that I actually use)
The ad supported free internet is one of the most important business models the world has arguably ever seen. Very few can argue with the fact that poor kids in developing countries over the past two decades and longer have had their lives changed beyond anyones wildest dreams thanks to the free resources at the tip of their fingertips.
On the same note, much of the wealth accumulation in the developer community has been on the backs of this very business model. The immense demand for dev talent and the astronomical salaries paid out is a consequence of the difficult financial choices made by so many before us.
When i read absolutely low-effort activism such as the text in the link about how(paraphrasing) 'sEaRcH eNgEnEs mAkE mOnEyY" and thus they are bad. I'm astounded at how intelligent people who can write code can simultaneously be so fucking moronic in their grasp of economics.
The web is an ecosystem. There are always going to be incentives that don't fit your moral compass that are getting optimized for and against. The answer isn't to burn it all down and shit all over a business model because it apparently doesn't fit your childish understanding of the ideal. By all means, compete, but atleast try to understand the various actors and participants in this complex web of entities and what role they're playing in the flow of investment, content, data and economic activity that is far more nuanced than "wEb rEsUlTs wIll B beTtTeR iF nOT oPtImiZeD fUr $$$ "
Face fucking palm
Absolutely nothing surprising about that. Intelligence without knowledge is not very helpful. If all you know is how to write code, you will suck with other things, even if you're intelligent.
The bigger problem is that people tend to downplay the knowledge that's required to do something, simply because they do not know how much they don't know. It gets worse the more intelligent you are because you're more confident in yourself then.
Case in point: name of this project. I've read it like 10 times on this page already yet I still can't spell it from memory. I could paraphrase:
> I'm astounded at how intelligent people who can write code can simultaneously be so fucking moronic in their grasp of marketing.
(but won't, since, as I said, it's not astounding at all; just for illustration purposes)
Great idea, and "instant search the web" would probably a better pitch then "non-profit search engine". Interesting argument that google doesn't do this because it isn't compatible with their ad model, but that doesn't mean a new ad-funded search engine can't do this. For google it might be billions of dollars in lost revenue while they adjust their ad model, a new ad-funded search engine wouldn't have this problem.
> Frictionless ... For example if you are typing “facebook” or “hmrc login” you could go straight there from the address bar.
No thanks. I sometimes do search for "company name" looking for the wikipedia article for the company, or news about the company, or information about the company in general. If you used facebook before, then it's going to autocomplete as soon as you type "face" in your addressbar, and you won't need the search engine. So if someone searches for facebook, they're either using the browser for the first time, or they're looking for information about facebook. Latter seems more likely.
- on the one hand I really want free, open and non profit services to succeed
- at the same time I greatly value the user experience
Don't get me wrong: These two things can go hand in hand. There are tons of good examples out there.
But, the closer you get to classic user-centric applications and leave the software developer bubble, the greater the discrepancy becomes in my experience. Brave, DuckDuckGo, Firefox and so on are desirable. But I always feel like I am missing out on the UX.
Google still yields better search results FOR ME(even with all those ads and clickbait).
Firefox still feels a bit dated and slow compared to Chrome.
I value the positive effects of free software so much that I am willing to accept limitations in usability in the hope that it will improve over time. But I feel like it should not be this way.
I can't support every project financially or contribute to its success as a contributor. My time and financial resources are limited.
I haven't really found a solution for this problem. My best guess is that the government should intervene in the free market and install market barriers to tame giants like Google. But this is repugnant to the liberal in me.
There is no way that design choices (especially the ordering of results) can be made in a way that pleases everyone. So either you dumb it down to the point of meaninglessness OR you enforce a mainstream-only ruleset.
The route forward, and what should be advocated for, is a distributed network of search engines, each for a specific vertical. If it operated as a cooperative they could share expertise and technology, they could then build a “meta” search engine for the co-op that combined all the results from the specialist niches. Each member basically “owning” the “franchise” for a specific type of search or category.
So, I don’t believe a single non-profit is the answer. More a co-op type arrangement where the co-op organisation (which may be a non-profit) has a mission to advance internet search through it’s network and strategic investment.
I think the pertinent question though, is what's the best way to demonopolize search. Maybe the answer to that is non profit, maybe something else.
Google has a most search users. They have an even higher (much higher) portion of search revenue and essentially all of the sector's profits. One advantage a non profit might have is going after the low profit parts of search. Use cases where Google is likely to be under-serving users.
Also, search isn't just websearch anymore. It's a way of calling a calculator, translating, etc. It's a text box that does stuff. The newest gen of language models may be the technical catalyst for some rapid evolution in the "clever text box" space. Google is obviously super active in this space, but shifts are a good time to get in.
Where would you skate, if you were skating towards where the search puck is going?
I wonder if something like that could work today, only with the index being shared across the user base.
The benefit would be that it’s a decentralized system. No giant infrastructure required which needs to be paid for by a big corporation. Basically, the infrastructure needs would be outsourced to millions of devices. And for websites, users and crawlers would be the same thing. Which is to say, you cannot block one without also blocking the other.
It could also add feedback mechanisms. Active ones, such as commenting on pages and discussing them, as we do on HN. But also passive ones such as tracking how long the user interacted with the page, to score the value of pages/domains and improve the ranking algorithm.
Edit - ah, he means the search engine should be a non-profit. Not what I thought he meant.
Not knowing what to do with $4m means the failure of the education systems.
It's weird that more people don't know about it.
But also there should be more than one decentralized search engine.
Ouch. I wish you the best but that statement makes me lose hope. Employees are expensive. Servers aren't exactly cheap either. And unexpected mistakes along theyl way cost a lot.
Google has too many of them and it's probably why Google Search is not improving.
1. https://addons.mozilla.org/en-US/firefox/extensions/category...
Edit: You can make a bookmark and add a keyword, but that doesn't help me use it as a search engine. It used to be you could just create one from the contextual menu, what is this convoluted process.
1. Go to kagi.com
2. Right-click on the search field and choose "Add a keyword for this search..."
3. Fill in the keyword (e.g. kagi, as I did)
4. Right click in the top url+search field, and choose "Add Kagi search"
Now you can search via keyword, you can change to kagi while typing, and you can set the default search engine to kagi in the preferences.
How would Google know how long you spend on the site? It only sees what links you clicked and doesn't know what happens next. (Unless the website uses Analytics, but Analytics doesn't affect search ranking.)
https://www.ghacks.net/2021/03/16/wonder-about-the-data-goog...
This doesn't mean that this data can be used to inform Google Search ranking. That would be very shady and potentially illegal. I work for Google, and even though I do not work specifically on Search ranking, this doesn't sound like something that could be happening.
This relates to the "how can they know how long I look at a web site for" asked above - if not specifically they do at least know the answer stochastically.
EDIT: Ok, I see that this is about a search engine structured as a not-for-profit, not as a search engine for nonprofits.
Is there evidence that they do this?
How many advertisers outside of that country are going to ask YouTube to play their ads in that country in a language different from the dominant language? Probably not many.
Star trek imdb
First result is startrek.com, second result is Star Trek into Darkness IMDb but 3rd is xkcd.
It then goes off into Q and William Shatner Wikipedia links and Muppet Movie IMDB in Russian.
I tried putting a plus in front of IMDB and quoting Star Trek. It doesn't seem to be able to find Star Trek on IMDB. I admire the concept, and it is extremely fast.
I will help!
Are you looking to build a trusted network who can verify and validate other users responses on an undefined period?
[1] https://github.com/mwmbl/mwmbl#how-do-you-pronounce-mwmbl
His maths is correct, just an easy-to-misread phrase
obviously good search trump's the dark mode tho.
>I found it very slow the last time I tried it
Maybe you used a node that was running on a rpi ;)
That is, even if a site wanted to, there's no way for it to declare "I have content related to X". Even better would be if these indices could then be distributed in a cache-and-forward model similar to how DNS (another distributed discovery index) works. There was some exceedingly rudimentary attempt at this through elements such as keyword meta tags, but even at best these referenced a vanishingly small fraction of the actual content of a site or article. Sitemaps also address a component of the problem, but again, only in part.
Some might see a few immediate issues. One is that not all site are sufficiently dynamic to know what content they actually contain. To an extent this might be addressable through extension to the webserver protocol such that a server would be aware, or become aware, of what content it contained.
Another is that a site might in some instances be inclined to misrepresent what it contained. This may be hard for some to believe, but I'm given to understand it occasionally does occur. To help guard against this, there might be vetted indices, in which one or more third parties vouch for the validity of an index. These reputation-sources could of course themselves be assessed for accuracy.
But if sites were responsible for reporting on what content they actually contained, and could be constrained to doing so accurately, a huge part of the overhead in creating independent search engine, and breaking the seach-engine monopoly, would be eliminated.
One might imagine why certain existing gatekeepers over Web standards might oppose such an initiative.
There would still remain other problems to solve within search space. It's possible to divide General Web Search into a set of specific problems:
- Site crawling: this includes determining search targets, any exclusions from such lists, and performing the actual crawling. Self-indexing addresses part of this problem.
- Indexing: Mapping of actual contents to keyword and query terms which might address that content.
- Ranking: Assigning a preference / deprecation to specific sites. This is essentially a trust / reputation assessment, with a canonicity / authenticity assessment (e.g., where did a specific item or document first appear).
- SEO: This is the Red Queen's Race issue in addressing insincere / malicous actors. Strong and durable penalties for abuse, and long-term reputational accrual, should be useful here.
- Query interpretation: There's a considerable art to figuring out what a question actually means. In some cases queries should be taken strictly verbatim. Quite often, however, interpretation is necessary. How those alternatives are posed might vary, with an option not often employed presently being to suggest a range of potential interpretations or related queries which might produce better results for specific query scenarios.
- Presentation: This is generation of the serch engine result page itself, incorporating several of the other considerations listed, but also addressing usability, accessibility, clarity, and other concerns.
- Revalidation: As the editors of the Hitchiker's Guide observed, the Universe is not static, and circumstances change. Revalidating, revisiting, and revising results and reputational assessments is necessary.
- Monetisation/Funding: I'm partial to a public goods model, or perhaps a farebox role via ISPs, pro-rated to general income/wealth within a region. Advertising, as a famous Stanford research paper prophetically observed, forces disallignment with searchers' interests and objectives.
Have you noticed how newspapers systematically do not supply a clear source for their articles? It's especially prevalent on political cases where there are easy-to-link paper trails. This makes it a lot harder to find the source for their article, so you end up just taking their word for their angle on the story.
A great recent example is Biden's Executive Order on the protection of women. When the newspapers writes about his EO, they're never doing it form a neutral standpoint. In this case they're either pro or anti abortion. But if you want to know the contents of Biden's EO for yourself, then you're forced to search for it. And depending on the search engine, that might also be hard because also search engines are politically biased.
Just so we're clear, this post isn't pro or anti abortion. Instead it's an example on how newspapers systematically force you to take their word for their angle on any given news story. So if you want to know source material, then you're forced to search for it. And when you do search for it, you're then at the mercy of the political bias of the search engine.
For that reason I'm not so sure a non-profit search engine will make political biases go away, especially when you consider what happened to Wikipedia. While not a search engine, it is a non-profit and communal project that set out with the ideal of being truly neutral, but in the end it failed at that, and some would say spectacularly. And the main reason is exactly bullshit, or rather the BS that comes with political bias.
Don't get me wrong, it's still a great source for information, but when you search for any topic that is in any shape or form politically sensitive, then you have to know about Wikipedia's clear political bias beforehand, or else you might take their angle as gospel.
This is especially insidious when it comes to search engines and also social networks, because most people assume that what is shown to them there is neutral, or at least coming from a friendly party. But then it turns out, that's not always the case.
When you systematically get biased information, then it's a democratic problem, because it prevents people from making up their own mind about political topics. Thus when people finally vote, the risk is that we get a society that does not reflect peoples actual opinions.
I think most people in here has been on the receiving end of that, no matter which side of the aisle you're on. And the result is always resentment and bitterness which in turn does not make for a healthy democratic environment.
Instead the political bias should be more clearly visible and out in the open on both newspapers, encyclopaedias and search engines alike. And while a non-profit search engine would certainly save you from corporate interests, it still won't save you from political ones, though it might be a good trade-off to save privacy.
Linking an archived copy of the content has merits, though increasing The Usual Suspect (the Internet Archive's Wayback Machine) itself has difficulty in preserving complex content. (The rather less transparent Archive.Today is often superior in quality if not necessarily in trust or reputation, and I say that relying on it heavily myself.)
I know that I've noted news organisations which do seem to link sources relatively freely, and thought I'd commented on it previously. I'm not finding any previous mentions by me here at HN (my usual personal quips trove, amongst other benefits). Though if memory serves, the New York Times tends not to provide links where it really ought to. I believe the Los Angeles Times might have a greater tendency to. Perhaps also NPR and/or The Guardian which I tend to rely on, though I'm not certain that's been my earlier observation.
In the 19th century and through a good part of the early 20th, newspapers would not only reference specific documents or speeches but publish them in whole. In large part, this is because that was the only way to distribute and reference the material. That practice has waned tremendously, and we often see material only through commentary and reference, rather than in original form. I've come to view this quite dimly.
I'm also fairly certain that copyright is a major consideration and factor here, and another way in which it's proving a disservice to the public good.
For numerous reasons, it seems that there's an "is-ought" disconnect here, as is often the case.
Business tends to be exceedingly spooked by risk, especially long-tail unconstrained risk. And copyright litigation presents an excellent example of same.
There are also other concerns. In an era of physical print and shrinking "news holes", the actual textual content of newspapers tended to shrink, perhaps establishing a tradition of no longer printing speeches verbatim. With the attention economy of the Web, the risk of sending readers off-site is a concern I've heard voiced many times both in print and public discussions and privately amongst people I know in the press. It's quite unfortunate, but real.
The case of US Government documents, in which there is no copyright concern is especially inexcusable. I'd agree with you strongly there.
Increasingly when I find such an opinion piece, regardless of the publication, I look for the source document and try to read it first. (I don't always follow through, but it is if nothing else an aspirational goal.)
If not, I think the way some WEB3 projects are funded maybe an interesting inspiration (not talking about Ponzi scheme here). Many projects are "non profit" and sale tokens before the service is 100% ready. It funds the amelioration and scaling of the project. And the possibility to resale the tokens at a higher price in the future sometime attract token holders and often increase the "motivation" of the token holders / supporter of the project... fuel the community (money is only part of the motivation). Here token could be associated to symbolic "privileges" (badge, access to early releases), or governance (taking part of some votes).
This system have clearly some drawbacks, but allows sometimes to increase the number of early users and supporters, and get more funding while staying a non-profit.