I'd rather that web robots use this information to build useful indexes than to have to worry about generating yet another feed in the hopes that it helps people find my content in a search engine.
Besides, a web robot can determine how much other sites link to my content and help determine its overall ranking in results. Adding another type of index file to my site will do nothing to determine how it relates to other sites.
I don't see this as a barrier unique to startups.
We have had embedded metadata in websites for decades. In the beginning, Search Engines did even use them. Until someone started stuffing unrelated keywords in it to rank higher.
How is that any different from requiring a crawler to index XML sitemaps?
> At a minimum, adding some metadata content to XML sitemaps
The purpose of a sitemap is to tell a web robot what resources there are, with some minimal metadata about page titles and last modified date.
Google has some extensions for identifying images and videos.
But that adds more work for site maintainers, who have to duplicate work.
1. Implementation (sites do not need to have a sitemap; or those that have it, may not have an accurate one)
2. Discoverability (finding sites in the first place, you'll need a centralised directory of all sites; or resort back to crawling in which case sitemaps are not needed)
3. Ranking (biggest problem in creating a search engine)
1. This would be up to sites, to your point, major question would be best way to create incentives.
2. This is solvable via a number of approaches, but the search engines themselves would be mostly responsible for finding the right approach for their business. I know how I would do it.
3. Indeed, which would be the main point of this decentralization, to let search engines focus on their hardest problem.
Edit: would Kagi not benefit from having to worry about crawling / indexing sites?
It would, but sitemaps do not provide that function as we discussed above. However if EU Open Web Search succeeded, that is something we could probably use to some extent.
Not only are some sites malicious -- mostly unimportant ones -- but many good sites are simply incompetent.
So you mean that a search engine is supposed to ignore what you are asking for and instead give you what it thinks you really meant?
Actually, they aren't really natural language queries. They are just ordered lists of search-terms. Goo provides no mechanism for saying "This is an english-language question". And even if Goo could parse my natural language, and rephrase it as something like "Are you looking for a list of books published by Douglas Hofstadter?", when you turn that into a query on the index, it stops having anything to do with natural language.
All search engines that attempt to be useful will have to filter out the junk. You just have to trust that the search engine you are using isn't withholding results from you that it considers "bad" (eg: "misinformation" (i.e. stuff somebody disagrees with)).
And to me, that is the crux of the debate really. Nobody wants spam for search results--everybody agrees with that and there is no real debate about filtering that crap out. The argument really is should a very large company that has a huge market share get to decide what constitutes "fact" and what is "misinformation". Based on 2.5 years of experience so far, what was once deemed "misinformation" has a sneaky way of becoming "factual information". Labeling and hiding "misinformation" because it goes against some narrative pushed by incredibly powerful entities is very scary and there was a hell of a lot of exactly that going on during this covid crap.
I used to fall on the side of "private companies can do whatever they want" but now I'm not so sure. Companies like FB, Twitter or Google play a huge role in shaping politics and society. I'm no longer convinced it is okay to let them play the role of "fact checker" or anything like that. Filtering spam is one thing, but hiding "misinformation" is entirely different.
I think we live in a world now where we are so used to a few tech giants mediating everything for us that we can't even imagine other solutions to this problem, but it's also how we got to this point in the first place.
Why is it not enough to punish sites that abuse the keywords?
You need a trustworthy core by which you can judge the vote of new users. You can incorporate them until somebody complains about a result that is out of place.
This doesn't have to fully scale. There are many pages without monetary value that won't be manipulated. The tags are an additional signal that can be used where they work. If they don't work, they can be ignored.
But it will scale because there are far more consumers than producers.