EU Open Web Search project kicked off
openwebsearch.eu
openwebsearch.eu
Most of the negativity seems to come from the following points:
- EU-funded project cannot succeed in tech because previous EU-funded projects have failed in tech in the past (and generally government-funded project in tech are suspicious)
- Search is hard, and therefore it will fail
- The project is underfunded
Even if the points above can all be valid (though the obvious US-funded startup launched by a bunch of uni students seemed to have fared pretty well) it seems we are missing the point of this project.
The proposal is to contribute to the creation of open building blocks necessary to enable others (including private US companies) to make better search products.
Better search product are needed.
Shall we remember HN of some of the intense conversations that happened here about Google failing us:
- Google Search is Dying [1]
- Every Google result now looks like an ad [2]
- Google no longer producing high quality search results in significant categories [3]
So while, yes search is a hard topic, we should welcome initiatives aiming at improving the ground infrastructure needed to lower the barrier to entry on this subject and hope it will allow many companies to build better search products (or inspire other initiatives to contribute in similar and even more successful way)
1: https://news.ycombinator.com/item?id=30347719
My 3 going on 4 kid was watching a cartoon on Youtube: Curious George.
An ad popped up promoting some show, featuring foul language and sexual intercourse.
Yagoddabekidding.
Something like, perhaps, the same standards and sense as in traditional broadcasting.
> Everything else is just a funnel
A funnel promoting pick up trucks and financial services to toddlers? Do they count that as an "impression" in the statistics that they feed the client? It seems like borderline fraud.
Youtube Premium has no ads.
"My /own/ children don't see inappropriate ads; therefore there isn't any problem."
> Before complaining about broadcasting standards, maybe first up your parenting standards.
"Before whining for ice cream, maybe first eat your dinner!"
How can otherwise smart people accept being treated like brainless idiots by similar services is beyond me. And how can they teach their kids that's fine too is something I'll never understand... but hey its your kids, everybody sees quality of their parenting first hand (and then complain how kids are unruly and eat junk these days... gee I wonder where they took inspiration and who built their character).
I'm using it in Japanese. Using the on-screen keyboard in Japanese mode, you can enter hiragana only. It doesn't do any word recognition to turn into kanji or katakana, which results in the search results being poor.
Nobody uses onscreen keyboards on Android TV for anything beyond entering a Wi-Fi password; it's a nonstarter.
YouTube Kids for Android TV ... ironically, just a toy for now ... with a well-deserved, accurate 1.4 star rating in the Play Store.
An interesting thing I noticed is than initially when I used the app)(with the same YouTube profile) I would get worse video suggestions. These days they are the same in YouTube app as well as smartyoutubetv.
Another interesting observation is that if the app breaks and I start watching YouTube in their app the frequency of their ads is quite reasonable in the beginning. Few times I even thought. Hey, if they show me an ad every couple of videos that's not that bad right, I might switch back to YouTube. Then as the amount of content watched increases the amount of ads increases to the point of making it unwatchable (4 ad breaks in a 20min video, sometimes even double 40s ads one can't skip). This makes me go back.
A) Youtube's terms of service clearly state "not for kids under 13"[1]
B) Youtube has produced a product for the younger age market [2]
C) folks in this thread are reasonably complaining that full youtube isn't appropriate for children.
Am I missing something here? (Like maybe your child was using youtubekids and still got the unacceptable content?)
[1] https://blog.youtube/news-and-events/children-youtube/#:~:te.... "While we permit users between the ages of 13 and 17 to register for an account with parental permission, we do not allow children under the age of 13 to create an account"
Not sure about software, but all (well, Apple is getting there) phones now use the same connector to charge because of the EU.
The correct strategy isn’t to hand this over to a giant bureaucracy, but to create an atmosphere where we can have a dozen alternatives to Google.
The only search product that comes close to this is Amazon search or booking.com (or maybe YouTube).
All multi billion dollar businesses. And when smaller ones emerge like the flight scanner one, google buys them.
I don’t see google going anywhere soon. They have a lot of faults but they’re still too good.
What is so great about those search examples?
However, there's no way Geizhals will ever expand into general search. The reason why they are so much better than the competition is that they are focussing on a small, profitable niche and presumably use manual data entry to ensure they have the best data.
[1]: there's a UK version too "skinflint.co.uk"
I don't think I've seen a website that does that before and I think it's worth praising.
But at the end of the day the resulting dataset would be sold to third parties so they can rank the results appropriately. Which to me seems to be the only sane way forward. Only a government could run something the scale of the Google crawler and succeed at doing so. And then everyone can build search engines on top of that.
That bit troubles me. If the index is maintained by a government agency, and every search engine is using the same index, then that's a massive censorship avenue. I wonder how "open" Open Web Index is going to be.
The vast majority of the net is, after all, now indexed, so you can run your own indexer to cover whatever they didn't cover.
> improving search relevance by x% is tremendously socially valuable, but probably makes Google's bottom line go up by a thousandth of x, and they have a very direct understanding of the connection.
I'm not sure this is correct. In a vacuum of real competition, the most profitable ought to be how it is right now, when search results are kind of ambiguously bad, so you need to click and skim through a few results to maybe find what you want, multiplying the number of ad impressions.
We can talk about the Importance Of This until we die of old age. Google failed us, Google is Borg, Yandex is FSB, Bing is ... Bing, etc. However, the fact that there is a problem to be solved and It Is Important doesn't mean that the EU will solve it. If anything, it's just the set-up for another political play that will have damaging consequences to the internet as a whole, just like GDPR, EU's poster-child "internet project".
They made GDPR just strong enough legislatively to be annoying, but not strong enough to actually change anything. Companies can still store EU citizens' data anywhere they want and do whatever they want. There's no insight into this. It is an unenforceable law and the only artifact of it existing is that I have to have "I don't care about cookies" installed, so that Avast Antivirus can eventually decide to silently collect and sell on my data.
For all intents and purposes OpenWebSearch is most likely not meant to succeed at anything either, and is just going to be a political stepping stone towards legislature that will be awkward and make the internet worse for all users. EU has a long history of creating or hanging onto laws that betray a misunderstanding of how the digital age works, how the internet works, how data can be copied or moved around for free.
Here's an example. Every country in Europe has some sort of legal construct in place that will prevent you from secretly recording a conversation you're having and uploading it to the internet. So for example, take Austria. They accomplish it by preventing you from publishing it on the internet. However, they don't prevent you from recording it. There are laws against secretly recording a conversation you are not a part of, but there's no such law for the situation when you are part of the conversation. So you can record and upload, just not publish. However, you can get on a train and go to a different country which has laws that prevent you from recording, but has no laws against publishing. Then you're in the clear. Or you can just use a VPN so that it looks like you uploaded and published it from the other country. Or you can just upload to YouTube, which will not show where the thing was published from - and claim that you did so during your tourist visit to Vietnam or whatever. And if someone brings a civil lawsuit? Good luck trawling the Vietnamese legal system for clues about that specific issue. Hope you know this rare language with virtually no legal experts who at the same time speak Vietnamese and your particular local European language. The costs would be on the order of tens of thousands at least, which is out of reach for anyone but the wealthiest EU citizens.
Want more? Egregious copyright related laws are known by everyone.
So are dns blocks of shunned sites. Or the recent Austrian project to block cloudflare IPs which literally broke the internet.
The fact such laws are still in place from before the internet - or are even still being put out - and are effectively unenforceable while making the internet worse for everyone - makes it clear that the governments in place simply have a misunderstanding of how the internet works.
None of that will stop politicians from coming up with BS excuses breaking the world with fever-dream laws in order to push their latest political agenda. The real question is: if such clearly unfit legislature is being put in place for political folly in industries we understand - how much of that is happening in industries we don't understand? Health Care, Food, Agriculture, Civil Engineering, Patent Law, and so on. My guess? You can tell what my guess is.
As for Open Web Search, what everyone should really be asking themselves is: ok, so what's the scam that's going to be pulled here?
As I'm one of the people who is working within the Open Web Search project please allow me to feel strongly about your statement. I've been involved in campaigning for this project since around 2014. This project did not originate as a 4d-chess move of some political game. It exists because of the hard work of a group of people, some of which are researchers, some of which are working at smaller search engines and some of which are involved in civil rights organisations. Currently still sitting in the kick-off meeting I can tell you that we are actively discussing how to get this project to produce useful results for a european open web index. Getting the EU to draft problematic legislation is neither in our power nor in our interest.
Google's priority with ads is to make sure ad buyers are happy with the results they are getting. This means making sure the people who click the ads go on to buy the product being advertised. That's what ad buyers measure.
That's a major fallacy you're opening with. If something is important to us (as a society), we will find ways to make it serve us well even when it's under attack. The rest of your post even makes a similar point: the social value of well-working search is greater than the economic value for even the biggest search monopolist on the planet. So why not socialize it?
Being idealistic about it won't change the outcome.
As far as I know, this hasn't really been tried, definitely not with the kinds of resources this project will have available.
PageRank was initially much less prone to SEO shenanigans because it relied on signals from other sites (incoming links) to decide how important a result was. Of course, as Google became more popular, people started sharing links on other pages and so on to cheat the PageRank algorithm. And Google have been caught in a fight with SEO ever since.
AFAIK, nobody has ever tried that because of the dangers of vote buying and coercion. Essentially, gaming the system.
(I don't think this proves anything! Simply wanted to suggest that comparing search to democracy doesn't significantly change the analysis wrt opacity.)
It hasn't escaped the wider world that quality open-source search is desirable, and it's hard to think what this new EU project brings to the table that isn't already available if others want to contribute to existing efforts. I wish the EU project the best of luck of course!
This is another Gaia-x, remember, the EU big tech cloud killer?
Do we distinguish direct EU funding from funding by governments of EU countries? If not ASML would like a word, and I'm sure people from other member countries can come up with other examples.
Like the Internet, or the WWW?
I'm not suspicious of all government-funded projects (back in the day my own PhD was government-funded!) but I can't help be suspicious of claims such as:
"an open European infrastructure for internet search, based on European values and jurisdiction"
and
"The project will be contributing to Europe’s digital sovereignty"
Q1: Who defines "European values"? Is that done by Qualified Majority voting or would - for instance - Hungary have a veto on any proposed definition?
Q2: Which treaties regulate "digital sovereignty"? Recalling that the 27 member states each "remain sovereign and independent"[0], is that digital sovereignty being handled in BRU, in the 27 states, or a mixture of both?
jurisdiction = GDPR will not be cheated.
European = We mistrust America, because they do crappy stuff that meant we had to pass GDPR.
>Europe’s digital sovereignty
Means - Europe will not be ruled by interests outside Europe, the interests inside Europe can fight it out via EU procedures.
Are there (m)any global search engines with European users who are "cheating" GDPR? Which ones?
> Europe will not be ruled by interests outside Europe
That's a very bold claim, especially considering the geopolitical situation right now.
> the interests inside Europe can fight it out via EU procedures
Would that be the interests of the estimated 25k-30k lobbyists who work in Brussels on behalf of their corporate paymasters?
>Would that be the interests of the estimated 25k-30k lobbyists who work in Brussels on behalf of their corporate paymasters?
probably, as well as the various governments that exist in the EU.
But we still mistrust our citizens that's why we introduce Chat-control 2.0, so we have a "better then China" control over private communication.
European values ;)
That's why People trust evil-company's more then stupid[1] Governments.
[1] never attribute to malice that which is adequately explained by stupidity.
I think the fundamental problem here is that the people that are interested and in grants and are capable of writing grant proposals are different than the people that are interested in building things. There's very little overlap. So the money goes to the people capable of writing proposals, and the people doing the work do it for free in an obscure corner of the internet.
It's sad really, but I suspect it's a side effect of the huge bureaucratical machine that is the EU. One way to make this better would be to simplify the access to grants so that technical people can do it without needing a class in "EU funding speech".
One of these days I’ll actually implement the above assuming nobody else does. I figured if I can at least get the basics done and a reference implementation that’s easy to run it could prove the concept. If anyone is interested in this do email my in my bio.
What I worry about for this project is that it becomes another island which prohibits remixing of results like google and bing, and its own index and ranking algorithms become gamed.
I wish the creators best of luck though. I am also hoping for some more blogs and papers about the internals of he engine. So little information is published in the space that anything is welcome, especially if it’s deeply technical.
Good luck! I'll be watching your progress and cheering you all on!
I'm not trying to be dismissive, it's just my feeling from working on search.marginalia.nu is that nearly every aspect of search benefits from locality, not only is the full crawl-set instrumental in determining both domain rankings and relevance signals on a term-level such as anchor tag keywords; but the way an inverted index is typically set up is extremely disk cache friendly where the access pattern for checking the first document warms up the cache for the other queries, but that discount obviously only exists when it's the same cache.
Enough people add domain specific search endpoints, with perhaps a taxonomy to say "hey send those sort of queries over here" and you have a compelling engine that self heals should someone stop running things, or starts spamming.
You can also integrate search results for which you cannot have the index, like social media APIs, another reason.
You could also mix and match search results from various topic-oriented indices. That's a research question, whether that is really better than building one unified one. But we think it is the way to bring index fragments to the edge, with the obvious privacy advantages.
I imagine a binary, with a simple Admin UI allowing you to crawl some domains recursively would be enough to index your own website, and then have those results shared.
Where I could see this being really useful, is let someone who knows everything about pokemon provide the index for searching pokemon information. Then when they federate, provide a taxonomy saying "for queries that have these words, call me". Suddenly you have a very high value search source for pokemon.
Throw in some zero click info information boxes and you have added a lot of value.
They can use a solution which already integrates the search. Forums and CMSes are a good target for that. Then you can say "I'd like my search to look at widgetlovers.com too" - and you get their sitemap + featured external links, because they run FooPress that supports it.
Kind of the same as sitemap we already produce for Google.
That's basically Usenet killfiles and, yes, I think they're totally due for a comeback in one form or another. Usenet may have had its issues towards the end (although it still exists), but killfiles weren't one of its problems. The simplest one you could just discard sources you didn't want to read anymore but the more advanced you could assign weight/rankings based on various factors (keywords / usernames / if you did participate or not in a discussion / etc.).
There are also some web extensions available so that you can fill it with more data.
For me it made more of an enterprise-grade use case (e.g. for building a search for your own file servers or confluence) so I only tested it out a little. It's a huge java project, that's why I decided to go with searx back then...cause yacy was pretty hard to setup.
> a decentralized p2p websearch and collaborative tool.
> It relies on a distributed collaborative filter[6] to let users personalize and share their preferred results on a search.
Do you see any major Mastodon nodes interfacing with Truth Social or Gab? I certainly don't. If federation barely works for a social media app, I fail to see how it would even matter for a search engine.
Even so you could base it on activitypub I suspect. It would need to be extended for sure to implement the sorts of things believe would be required.
From the webpafe that half of the time shows "Resource Limit exceeded" to a technology stack diagram on the bottom of this page https://openwebsearch.eu/the-project/ being completely unreadable due to bad scaling.
It is very disappointing really. Another example from the top of my head. Here in Poland we have ID cards(as every other EU country) . Those ID cards have to be renewed every now and then (10~15years). In last years an online system for government services was implemented including for renewal of those cards. One could take a photo with a mobile phone, submit an application and pick up a card from a gov office in few weeks time. Unfortunately, EU made a law that ID cards applications have to be acommpanied by biometrics (fingerprints) so this system has been thrown away. One has to physically go to the gov office, scan their fingerprints, apply for a new id card and then go again to pick it up...
Ok, so what happens in 10 years time? They should have the fingerprints already, right? No. They take the fingerprints, they store them only until one picks up the id card and then they are deleted. There are no fingerprint database, they are not stored anywhere. The fingerprints are used only to ensure the same person that submitted the application picks up the document. It makes zero sense, other than to break the previous online system. Thanks EU.
The government already does... just saying.
Our (Polish - and other EU states as I understand it) government is not allowed to store everyone's fingerprints unless they are a criminal. I as well as other people here have pretty strong feelings against it.
They are. They are stored on the card itself, and they are NOT stored in government databases. This is a good thing...
...that we know of.
Yet.
Also, using fingerprints to routinely authenticate people presents a whole new lot of problems. No one voted for a party that proposed such idea.EU simply decided to mandate this and it has to be implemented no questions asked. What about people that have trouble using the fingerprint scanners govs use? I used to play base guitar when I was a teenager. I can't wait to find out how well/bad this tech will work with the thick skin on my fingers.
edit: I mean you can renew a US passport purely by mail without ever interacting with anyone and I haven't heard of massive issues caused by that.
Also one needs an ID card plus 2FA authentication(usually connected to a bank account, a physical smart card, or mobile phone) to login to the government services portal in the first place. This portal is seen as a huge accomplishment (in comparison to the inconvenience of having to do every little thing in person, queue for hours etc). It is not just taxes, it is health service, local councils, national(and health) insurance, building permits, basically almost everything one can do in person can be done via this portal - except renew an ID card since the stupid EU rule came in force...
As for being corrupt and anti-EU, I have to say in recent decade at least there has been no other group of politicians more incompetent, corrupt and anti-EU over the EU commision itself. From the botched/corrupt "green new deal" that resulted in complete dependency on Russian hydrocarbons and resulted in another war in Europe, through the "pay Turkey for the problem to go away" mediterran refugee "solution", to complete ineptitude at the first 6 months of the pandemic and basically leaving Italy on its own, culminating in illegal witholding of funds to member states that elected parties opposed to the current option in Brussels.
However, what truly destroys EU is not even the above, but the lack of respect for the rule of law amongst the top officials. They have their goals and no matter what, they will do anything to reach them. For example they want more integration and a federal state. They proposed it fairly some years ago as an EU-Constitution and it was demolished in referendums. Instead of giving up, hearing the democratic choice and going the direction the sovereign(the people) told them to they then proceeded to implement it another way over people's heads(It was supposed to be implemented in the treaty od Lisbon). However, those treaties have to be unanimous, and some countries didn't want to essentially be ruled by the biggest countries so it got watered down back then. Then they realised it is impossible to implement this goal in accordance with the rule of law, so what they are trying to do now is twofold. First, throw out unanimous voting in favor of majority vote so smaller country objections can be disregarded. Second, bully countries that disagree by illegal witholding of funds.
We're very near the end of the EU, and it does make me sad because I still believe in the ideas that led to it in the first place: free market for goods, travel and work, shared values and work towards common goals between member states - not bully eachother or sell other member's state's security for financial gain.
At least in my generation 10 years ago I would think 95% people would consider themselves very pro-EU, now, unless the current political class GTFO promptly I don't see EU being a thing in next 10 years.
The primary goal for academics is to publish new findings, while what you need to build a search engine is rock solid CS and information retrieval basics. Academically, it's not very exciting. Most of it was hashed out in the 1980s or earlier.
> 7 countries.
> 25+ people.
There are literally dozens of them!
The problem is that you need people who actually know how to architect complex software systems much more than you need revolutionary new algorithms. For that, professors are the wrong people. A professor on the team, sure, that might be helpful. Not half a Manhattan project's worth.
I disagree on the budget though. It is basically pocket change.
A shoestring budget keeps the costs down by design and by necessity. A large budget virtually ensures the search engine becomes so expensive to operate it will never break even.
And the EU just solved that problem.
Have no fear; all of the actual work will be done by PhD students straight out of undergrad, and most of the actual leadership will be done by a string of recent PhD grads who need results in 6 months because they'll be full time job marketing for the 6 months after that ;-)
Who is doing product management?
Who is doing product marketing?
etc
This is all applied engineering at this point, not R&D. How does it at all fit into academia's strong suit?
Also, tell me you wouldn’t love to work on a large project that wouldn’t be subject to the arbitrary whims and promises of the marketing department.
Which results in an interesting engine nobody uses. Products that start with the tech and then think of selling it fall on their faces for a reason.
Is Google proves anything, it's that greed is real.
Heh, so, funny story...
>A second grant—the DARPA-NSF grant most closely associated with Google’s origin—was part of a coordinated effort to build a massive digital library using the internet as its backbone. Both grants funded research by two graduate students who were making rapid advances in web-page ranking, as well as tracking (and making sense of) user queries: future Google cofounders Sergey Brin and Larry Page.
>The research by Brin and Page under these grants became the heart of Google: people using search functions to find precisely what they wanted inside a very large data set.
https://qz.com/1145669/googles-true-origin-partly-lies-in-ci...
But nobody will implement the 'boring' features needed to make the thing generally useful.
Curious what the salaries will be on this one.
I’m just skeptical if the EU bureaucrats will put the money in the right place, and if this is even the right approach.
Parent commenter own search project Marginalia Search [1] could even benefit from it, or even maybe collaborate with it.
It is not a winner-take-all situation, and we need various open initiative in this space to get out of the current conundrum we are in with Google stronghold on search.
Overall I think there are better ways to improve search from an EU perspective by doing what they are supposed to be doing:
- create a fair environment for companies to compete in, e.g., take a look how Google, Apple or Meta's assets are set up to make it harder for competitors, break that up
- improve standards in eduction – it doesn't really make sense for all member countries to think of and maintain a good CS curriculum and they all seem to be pretty bad at it
- make it easier to build something and get funded, and reward creating prosperity, don't tax it to death
Just tested marginalia's random mode btw. Pretty cool, reminds me of the internet when I was a kid
(edit: formatting)
Enjoy! It's a great story.
(Plus: for who might not know, DARPA is US defense research, and heavily influenced by the intelligence services needs. Which is not necessarily bad! Just good to understand where and how Google originated. And wrt DARPA, they funded the creation of the internet itself, for whatever matters.
In Europe, things often go slightly different. The Web is a result of CERN, who are also a project partner of OpenWebSearch.EU. Why? Well, better search can also be beneficial for better science, not just for end users wanting to find their way or buying something.)
I salute your efforts and endorse your search engine. I also recognize that you know what it takes to build a search engine.
I don't think you have deep familiarity with EU academia. - The primary goal for academia is to influence society. Publishing is a route to that. - Being head of the EU search engine project would give high academic status - There are hundreds of articles which you could publish on this project - Rock solid CS. Would someone like Knuth count as "rock-solid"? Who is better at CS, the person who can implement quicksort because they practiced leetcode, or the person who invented quicksort? - information retrieval basics. Again, these basics were probably developed in academia.
The skills you say are basic to this endevour are more prevalent in top-quality professors and post-docs than they are in industry.
This is requires far more CS than you'll find in your usual software development effort, for sure, and many CS professors absolutely fit that bill. However, to the same degree it also demands far more on the software engineering side. People out of academia in general, from every time I've seen them build software, have not been all too impressive on that side of things.
Web search has an incredible demand for being well rounded, beyond anything else I've encountered. CS isn't the hard part bottle-necking everything else, it's just one of the many hard parts.
Worked for Google.
So you mean that a search engine is supposed to ignore what you are asking for and instead give you what it thinks you really meant?
Actually, they aren't really natural language queries. They are just ordered lists of search-terms. Goo provides no mechanism for saying "This is an english-language question". And even if Goo could parse my natural language, and rephrase it as something like "Are you looking for a list of books published by Douglas Hofstadter?", when you turn that into a query on the index, it stops having anything to do with natural language.
Not only are some sites malicious -- mostly unimportant ones -- but many good sites are simply incompetent.
I'd rather that web robots use this information to build useful indexes than to have to worry about generating yet another feed in the hopes that it helps people find my content in a search engine.
Besides, a web robot can determine how much other sites link to my content and help determine its overall ranking in results. Adding another type of index file to my site will do nothing to determine how it relates to other sites.
I don't see this as a barrier unique to startups.
We have had embedded metadata in websites for decades. In the beginning, Search Engines did even use them. Until someone started stuffing unrelated keywords in it to rank higher.
How is that any different from requiring a crawler to index XML sitemaps?
> At a minimum, adding some metadata content to XML sitemaps
The purpose of a sitemap is to tell a web robot what resources there are, with some minimal metadata about page titles and last modified date.
Google has some extensions for identifying images and videos.
But that adds more work for site maintainers, who have to duplicate work.
1. Implementation (sites do not need to have a sitemap; or those that have it, may not have an accurate one)
2. Discoverability (finding sites in the first place, you'll need a centralised directory of all sites; or resort back to crawling in which case sitemaps are not needed)
3. Ranking (biggest problem in creating a search engine)
1. This would be up to sites, to your point, major question would be best way to create incentives.
2. This is solvable via a number of approaches, but the search engines themselves would be mostly responsible for finding the right approach for their business. I know how I would do it.
3. Indeed, which would be the main point of this decentralization, to let search engines focus on their hardest problem.
Edit: would Kagi not benefit from having to worry about crawling / indexing sites?
It would, but sitemaps do not provide that function as we discussed above. However if EU Open Web Search succeeded, that is something we could probably use to some extent.
All search engines that attempt to be useful will have to filter out the junk. You just have to trust that the search engine you are using isn't withholding results from you that it considers "bad" (eg: "misinformation" (i.e. stuff somebody disagrees with)).
And to me, that is the crux of the debate really. Nobody wants spam for search results--everybody agrees with that and there is no real debate about filtering that crap out. The argument really is should a very large company that has a huge market share get to decide what constitutes "fact" and what is "misinformation". Based on 2.5 years of experience so far, what was once deemed "misinformation" has a sneaky way of becoming "factual information". Labeling and hiding "misinformation" because it goes against some narrative pushed by incredibly powerful entities is very scary and there was a hell of a lot of exactly that going on during this covid crap.
I used to fall on the side of "private companies can do whatever they want" but now I'm not so sure. Companies like FB, Twitter or Google play a huge role in shaping politics and society. I'm no longer convinced it is okay to let them play the role of "fact checker" or anything like that. Filtering spam is one thing, but hiding "misinformation" is entirely different.
I think we live in a world now where we are so used to a few tech giants mediating everything for us that we can't even imagine other solutions to this problem, but it's also how we got to this point in the first place.
Why is it not enough to punish sites that abuse the keywords?
You need a trustworthy core by which you can judge the vote of new users. You can incorporate them until somebody complains about a result that is out of place.
This doesn't have to fully scale. There are many pages without monetary value that won't be manipulated. The tags are an additional signal that can be used where they work. If they don't work, they can be ignored.
But it will scale because there are far more consumers than producers.
[0] https://en.wikipedia.org/wiki/Quaero
[1] https://www.dw.com/en/germany-pulls-away-from-quaero-search-...
We are very aware of the Quaero/Theseus history :-)
These search engines will then have the freedom to define their own search product experience, business model, even ranking of results.
So.. Who's going to create the index? Indexing the web is expensive, and its offset by the ads the indexer runs on their search website, such as Google, bing, brave and others.
Bullseye.
Regarding privacy the bar is significantly higher than what Google has to deal with. This will come at some cost in quality and/or speed.
t. was involved in one of those plans as part of the research team in a public lab
Anyway, it's not a "European value", it's a piece of legislation, so it falls under the "European jurisdiction" category, not the "European values" category.
That complying with "European values" is part of the project goals is unfortunate. Who decides what are European values? There's no Declaration of European Values. We don't all have the same values. There are just European laws, and then a whole bunch of opinions that not all Europeans agree about.
Still, it's a good illustration that it's a somewhat unique European value. It's definitely not an American value.
do you also think you’re not allowed to remember people’s names without permission because of GDPR?
Let's switch that around: you're referring to their right not to be listed in my database any more.
I don't see what difference it makes whether I'm incorporated or not; I'm not OK with the idea that certain information cannot be shared. Like, if the information is false, that's one thing; but a legal requirement that some true information must be hidden, that seems like pure badness.
I think it's reasonable to label China and Russia as control freaks, but the US and Europe aren't much different at this point too given the creepy focus on trying to shut down whatever they want by calling it "misinformation" and pressuring social media to comply.
If you are arguing that being truly unbiased cannot ever be realized, I counter that neither truth nor justice can ever be truly realized, by this should not stop anyone from taking them as their values.
signed by: an expertI prefer reporting that wears its biases on its sleeve. Regrettably The Guardian squeezed out most of its most interesting writers, apparently by repeatedly spiking their stories.
Yeah, Truth and Justice are platonic abstracts; nobody thinks they exist in the real world. But people do believe that unbiased reporting is possible. It isn't.
> Resource Limit Is Reached
> The website is temporarily unable to service your request as it exceeded resource limit. Please try again later.
Original URL might be more resilient...
https://web.archive.org/web/20220920183027/https://openwebse...
I'm dreaming too much, I need coffee..
PageRank made user preferences as signalled by hyperlinks the key signal of quality. Who really hyperlinks any more? And how many, by percentage, are non-automated etc.?
The web this works for is one of forums & personal websites. It's hard to say that today there are any properties of websites that are a reliable signal of quality.
Hence the proliferation of voting sites (such as HN, etc.) which are little more than search engines augmented with reliable signals of user preferences.
the commenter is not referring to the search results, they're referring to the interface not having all the wank that google packs in to optimise for profit
>Who really hyperlinks any more?
who doesn't?
> Once the index has been created, the next step is to develop search applications.
> The team at TU Graz will be particularly active here in the CoDiS Lab and will work on the conception and user-centric aspects of the search applications. This includes, for example, research into new search paradigms that enable searchers to have a say in how the search takes place. The idea is that there are different search algorithms or that you can influence the behavior of the search algorithms. For example, you could search specifically for scientific documents or for documents with arguments, include search terms that have already been used, or include documents from the intranet in the search.
I didnt look for all but I did look who is partner from Slovenia.
The most privacy invading ISP/Mobile company in our country selling statistics about their users (although anonymized, but walking on a thin line).
I just hope it wont go down that drain.
Regarding all the negativity about the government project.
If we can run CERN we can surely do a web search project.
The remainder of the search problem seems to just be collecting relevant trafficked sites for listing in results. Today Google et al seem to be doing this BY HAND. And it's not even obfuscated.
Recently, for the first time in my life, the wizard behind the curtain seems to have been exposed. I feel strongly that one could probably start a small index that catered to a fairly large audience.
And honestly, for other queries, just tell the user to search that site directly. I think you could even market it to users as not a technical limitation, but behavior that should be considered fuddy-duddy.
Like, really, you're going to search me? You know they have their own search right?
Even Yellow Pages faded into obscurity eventually.
I have the exact opposite experience.
To wit: searching HN via the algolia link at the bottom is way worse than searching on Google with a site:ycombinator.com restrict.
Same thing for YouTube, where the search engine is tuned for maximizing watch time and strictly not to return what you're looking for.
No research institutes from {France, Italy, Spain, Greece, Portugal, etc ...} involved.
Who is defining those "European values"? Is it the European Commission?
As a liberal person who still dreams of mature citizens who form their own opinions well-informed from a rich debate, I now see the "European values" quite critically.
In Germany - if you criticize the Corona measures you are called a Nazi, if you criticize the Ukraine war you are a Russian troll. Are those then the "European values"?
The European Commission has already established some projects to weight the information according to its will (SocialTruth, PROVENANCE, EUNOMIA, etc.).
When a government agency talks about truth, then all the hairs on the back of my neck stand up.
When governments speak of disinformation, it is usually only in the sense that the information does not correspond to their interests. I had to learn that the so-called fact checkers don't check facts, they just sell a counter opinion as a "fact".
So for a European search engine, a filter is then placed upstream that filters out all disinformation?
In the past, that was called censorship. Today, it's more like citizen service.
These notions are propagated by the US to the rest of the world. While EU politicians seem to go along, europeans do not tolerate it much and that much is reflected in society.
If the EU or more gov sponsored search engines pop up, i have no doubt they will want to control them in their own way - i'm okay with that. Right now our only option is US controlled entities that answer to US govs and rules.
At least here we can point fingers and hold politicians accountable if they try to influence things the wrong way. We can't do that with US companies.
Germany is currently governed by a chancellor who has in fact been convicted of lying, has selective memory lapses as a co-responsible party in one of the biggest banking scandals, and was also instrumental in the political decisions to push ahead with dependence on Russian gas.
So sorry if I doubt that politicians can be made accountable for anything.
> the researchers will develop the core of a European Open Web Index (OWI) as a basis for a new Internet Search in Europe.
That sounds as if building an index is at least within scope. But isn't the design of the index one of the key differentiators between different search engines? Perhaps the OWI will just produce tools for building transparent, privacy-preserving search engines, rather than crawling the web itself. It's not clear (to me) from reading the site.
We further deliver components to make search engines on top of this index. The project vision is that there will be many different search engines, not just 4 worldwide. Hoping to lead the way!
Oh PageRank (R) can be done using real semantic links (actual human comments) rather than easily spammed hyperlinks. And you can group all similar pages and serve those as a single expandable/refinable entry.
For example, given 10 similar articles where one has a well-known domain but few links and the others are all cross-linked only among a known spam-net, then just show the well known page (stackoverflow) and not the useless copycats.
It seems doable for a computer to also navigate the page like a human and try to negotiate various cookiewalls and similar. If a page requires jumping through hoops to enter, rank it poorly. There should be a single decline/reject if there is a button at all, and the content ranked should be the content shown after rejecting. If it shows two buttons "Accept" and "Options" then just don't index that page.
Oh, how I wish that search engines would simply give up on encountering a paywall/cookiewall, and refuse to index the page.
How can something be unbiased yet somehow stop showing illegal search results?
I'd really like to see them match the 20+ years of search quality fine-tuning that Google built into their search engine.
Not that Google is as good as it used to, but still, catching up with them is way more complicated than just building a big crawl + index piece of infrastructure.
And all of that on a government-funded shoestring budget.
Mmmh.
Good luck to them, but I'm not holding my breath.
My understanding is google used to keep a lot of data around for "long tail" queries, and stopped doing that at some point. This seems to be the issue around search declining outside of "food near me" type queries (which they absolutely excel at).
Search needs help its true, but 8 million are not even gonna move the needle and that other engines have that much money invested, just to maintain their network switches.
One of the tech lead feed [0] looks quite biased towards Ukraine though. I hope it doesn't interfere with the search engine.
What made Google such a game changer was that they based their index not just on the contents, but on how pages linked to each other.
A tiny image with a few nodes and stakeholders, including third party indices and monetary services ("Third Party Enrichment Services"). In which decade are they living? https://openwebsearch.eu/the-project/
How much of what google spends on "search" is strictly for search, vs business goals related to search?
How much of that google spend is salary? What are those salaries? How do they compare with EU post-doc salaries?
--
I doubt this will take off. I mean they investend more in funding and marketing instead of starting to built something. they should've started with code (agpl3 of course) and invited more and more people. at the moment this is more buzzword bingo bullshit than anything else. it's basically always the same problem, instead of focusing on the product, they fous more on the message.
Where is the EU own and operated 2nm foundry?
Well, I would rather like to have visibility into the ranking algorithm, and to be able to control the resultset ordering.
When I said I wanted control over the ranking, I meant as a user. Tell me what parameters I can rank on; give me a UI that makes it reasonably easy to express a ranking. That's it.
I hope UI design and stuff is orthogonal to the construction of an open index. I think this could be very interesting.
The actual website of the project (with some concrete info) can be found here: https://openwebsearch.eu/
Besides it's a 8.5 million EUR project, it's literally nothing, it's payroll for a few people. The money is being invested into people who then spend most of it, so it's a triple investment.
Just because it crawled it doesn't mean it stored it.