Digital subscription rates have been rising, but not fast enough to subsist most publications without additional ad revenue.
My concern is that such a search engine would allow private interests to create "news/media/info" sites directly that qualify while long-respected publications are ostracized.
Great idea if the primary intent is to index other types of content outside of that sphere, but based on the history of the modern web I think we could expect that to be gamed pretty quickly in a detrimental way.
Even just writing this comment I keep coming into thoughts and it's making me realize what an interesting subject of discussion it could be. There's a lot of questions in there!
The best balance of privacy and user experience seems to be display ads related to search term context that don't do any tracking (like billboards in the real world). But even that isn't an easy proposition and will need a custom network built to generate those type of ads.
A lot of people say they will pay for a search engine, but it is such a small niche that I believe the cost would be prohibitively expensive, and the search engine would probably still be subpar to Google in many respect. Would you pay $10/month, what about $29? Those are most likely where the monthly fees would have to be for this type of product.
Anyhow, I'd love to see something like this. I've switched to DDG but often have to reach for Google if I can't find relevant results.
Excerpt from an old comment of mine [1]:
Mozilla Research put the entire web's advertising revenue at $12.70/month per user. In other words, if they are right we are living with the consequences of advertising for a mere $13/month, $13 dollar they still get from us anyway because it's baked into the prices of the advertised products.
---
2. the purple for the links makes me think i've already clicked and seen everything
Two links came back:
"Facts on Farts" "NiceJewishMom.com"
Not a good first impression.
The simplest is "search for doc that contaons" . But we're used to "search for document about concept" by now. ISTRC Bing called themselves a "decision engine".
I don't know what Google is/trying to be now, but it often thinks it knows my business better than myself, excluding critical words.
I don't think there can be a universal answer. The corpus gathering is a huge barrier to entry, but having a common corpus would still allow room for competition on diversity of querying methods.
Atleast the title in firefox says:
> Wiby - Search Engine for the Classic Web
That's also probably why you didn't find any information on the svelte framework.
Huxley "The Doors of Perception" (1954):
>each person is at each moment capable of remembering all that has ever happened to him and perceiving everything that is happening everwhere in the universe. The function of the brain and nervous system is to protect us from being overwhelmed and confused by this mass of largely useless and irrelevant knowledge, by shutting out most of what we should otherwise perceive or remember at any moment, and leaving only that very small and special selection which is likely to be practically useful.
We've all heard the comparisons between the brain and the internet. And we've all been overloaded at one point by the vast amounts of useless crap on the internet. So what about a site that allows users to customize their filters of reduction? You could have popular profiles premade and ready for tweaking, or you could go raw and witness an endless stream of information raked from all over.
What happens if you're a top result on the add free segment and you start incurring some level of costs? Do you drop off if you add an ad?
Granted I really like the idea, I'd love to see it tried, but I wonder what all the unintended consequences would be / skeptical of the value of segregating "ad-free" vs. "has an ad of any sort" vs. the sites that really are a mess.
I believe fastmail is like that.
Donations/patreon for one, then subscriptions. It should be possible to offer a contentless index for download (I estimate it could be around 50-200GB) and that could be a paid option.
I think that such sites would be in ballpark of a few ‰. That would enable me to offer the contentless index for download. With delta updates and torrent for distribution it could be not that expensive, but that's a thing that I could charge for.
My intention is to use AdBlock rules like easylist to check whether or not indeed the page.
My initial code is fine in Go, but I lost enthusiasm for Go lately and careerwise it's not a good fit for me (I don't have much time to learn something not as useful for me professionally). So I started to rewrite it in Rust, while learning it, you can laugh now (Rust Evangelism Strike Force el oh el). It has an advantage with ready to use rules parser from Brave [2] and presumably high quality tokenizer from html5ever [3].
I want to use a tokenizer instead of a full parser to be able to do stream processing bringing costs down.
Common Crawl data lays on S3 so the processing must be done initially on EC2 to keep it low cost.
[0] Current Go code: https://github.com/hadrianw/abracabra
[2] https://github.com/brave/adblock-rust
[3] https://docs.rs/html5ever/0.25.1/html5ever/tokenizer/index.h...
EDIT:
Also for the search part I want to use something more stand alone than Elasticsearch to offer desktop search with downloaded index. When I started with Go I wanted to use Bleve [4], now I'm not sure, but I think that Bleve is getting mature enough. I will worry when I will have some data to search through.
One of the challenges with this whole enterprise is a small need of JavaScript parsing. There is a common pattern, that for example Google Analytics uses, that uses a snippet of JavaScript to insert a proper script tag. But those snippets are very short so I think they may not need a full JS VM, maybe even a tokenizer would be good enough. Browser AdBlockers base on the site executing JavaScript already.