Brave Search Goggles: Alter search rankings with rules and filters
github.com
github.com
Here's the source code for a sample "Hacker News" Goggle. Essentially, it will prioritize domains popular on Hacker News. I could even see a browser plugin that lets you add or remove domains as you visit them.
https://github.com/brave/goggles-quickstart/blob/main/goggle...
And here's the language syntax.
https://github.com/brave/goggles-quickstart/blob/main/goggle...
wonder how they came up with that list? That list it self is a goldmine on it's own.
Would be great to complement with social media as we’ll. For example, make the domains in posts within my networks, more relevant.
That way, through controlling/curating my networks, I get to decide what’s more relevant to me in search results.
Would be awesome to add stackoverflow in there too :)
The cynic in me can’t help but wonder if this is just dialing up an echo chamber if we include social media proximity as a filter.
I later learned that it's actually against Microsoft's Bing Search API use and display requirements (https://docs.microsoft.com/en-us/bing/search-apis/bing-web-s...), which states that you're not allowed to:
> Modify the results content (other than to reformat them in a way that does not violate any other requirement), unless required by law or agreed to by Microsoft.
> Reorder, including by omission, the results displayed in an answer when an order or ranking is provided, unless required by law or agreed to by Microsoft.
> Display content that was not included within any part of a response in a way that would lead a user to believe that content is part of the response.
And there are many other restrictions imposed upon you. I'd argue the reason you see almost zero innovation in this space (the only thing I can think of is !bangs) is because non-Google/Bing(/Brave) Search Engines have very little control over the received data, and with the little control we do have then there are many restrictions that prevents us from making these 'innovative' features.
Based on my pessimistic reading then we're not allowed under any circumstances to merge API results from multiple search engines (e.g. Bing + Yandex + Google), and for every site we want to exclude/boost/derank/rewrite sites/urls then we need explicit permission from a Microsoft representative (which would prevent us from implementing this kind of feature). We're also not allowed to display third-party ads on any pages that displays part of the API results (which is funny because I feel it's near impossible to join Microsoft's Advertising network even though my site in my opinion have a good purpose and provide a high quality UI/UX experience - https://ask.moe).. I've later switched to Google's Programmable Search Engine and even though my site has been approved by Google as a non-profit, as well as approved for AdSense, then they will only display ads in the search results if I turn off revenue-share.
This industry desperately needs legislation that makes it possible for everybody to obtain access to Google/Microsoft's crawled data, and this access should not be limited to expensive APIs that make it impossible for anybody besides Google/Microsoft to do any sort of innovation.
I wonder if the LinkedIn scraping case would apply to someone crawling Google search results.
Google has never had the appetite for letting users explicitly improve its results. Their reasoning was that pure algorithm is the only possible solution that scales. Even the minor features, such as "I never want to see this site", were removed.
I am glad Brave is trying out something different. Crossing my fingers that it works, the current state of search is so meh..
Hmmm, can't find that discussions tab.
EDIT: I take this back. If I put the term “discussion” in the search along with “toyota sienna”, I see a “discussions” section in the search results if I am on the “goggles” tab. I don’t see the discussions section if I omit the “discussion” term though.
I have no idea why they aren't able to flag and block spam domains anymore.
If it is so easy, why don't you build a dominant search engine that manages to complete avoid the "totally easy" spam you talk about?
If perhaps, you have extensive experience in the anti-spam section of a dominant search engine (or some similar position), I'd give your comment more weight. Do you have such experience?
Incidentally, if anyone has insight into how to improve that aspect of my job, that would be much appreciated!
Unless you provide me with your qualifications and/or previous experience that would lend credence to your claim, I will view your claim as one from an armchair critic/engineer.
And yes, I'm being "an armchair critic/engineer", like everyone else criticizing Google or any other big company that seems to be getting worse at what they excelled.
Since you haven't provided any qualifications/past experience, I am not convinced at all that it's as easy as you say it is.
Ranking is already treated as predicting which of the search matches the user is most likely to want to see. Clickbait gets highly ranked because users do in fact click on it. Goggles (which sound like a great idea) apparently let you somewhat control the prior probability distribution as an input to the ranking algorithm.
I wonder what features it lets you select on. E.g. we had a thread here recently about finding useful product reviews. It would be great to have a goggle that downranks any site containing amazon affiliate links, since those are almost always basically shill sites despite being "reviews".
[0] https://linustechtips.com/topic/1270087-linus-media-group-ma...
Additionally we hide information behind proprietary platforms like Discord that cannot be crawled.
Ironically Googles research of text AI undermines their search algorithm because now spammers use said AIs to get good page ratings. A lot of sites have ever the same content with a few different words here and there. Especially new media content get immediately swamped with such sites since there are no established sites under certain key words. A domain filter sadly won't work as well since there a so many scamming sites.
They used to ask some users for feedback back in the good days: https://searchengineland.com/google-feedback-experiment-whic...
Brave is doing what Mozilla wont't do.
Also with Goggles you can finally block all the stupid Automated comparisons sites, Brave already had a forums section, and this makes it even better.
This is what I see: https://imgur.com/onlrA5I
Goggles allows to specify thousands of rules, not only on domain but also URL patterns and in the future matching on elements of the page as well (e.g. titles, etc.)
It is also not clear that Lenses predates the Goggles whitepaper. We’ve been playing with the idea of Goggles for a long time.
Disclaimer: I work at Brave.
Goggles white-paper was released more than a year ago, long before Kagi was even announced to the public.
Additionally, before Brave acquired Tailcat (Jan 2021) I had the pleasure to share the draft of the paper with Kagi's founder.
So no, there is no prior art.
Let me add that I do not claim that Goggles is prior art of Lenses either.
One of the key features of Goggles design is that the instructions, rules and filters are open and URL accessible.
A Goggle is not so much a personal preference configuration, but a way to collaborative come up with shareable and expandable search re-rankers.
Very different goals if you ask me. Of course, Goggles can be used for personal preferences exclusively, but that's not the use case we had in mind.
For a bit of historic accuracy if it ever matters for future readers:
Kagi was founded in 2019 and we have operated for years in private beta with thousands of users before public beta release this June.
Goggles were not inspired by Kagi”s lenses and I can confirm seeing the whitepaper before we got the lens feature out last year.
Kagi”s Lenses were inspired by Blekko’s “slashtags” which is probably the original “prior art” for this kind of feature.
Looks like we arrived to similar idea, but different execution. Kagi”s Lens feature is osimple to create filter for the web, that anyone can make with a few clicks plus a bunch of powerful built-in lenses like “noncommercial” or “discussions” search.
It seems to be better in every other way too, but that is actually the single reason why I pay for it.
Our own innovation in this domain go back to 2005 and was called Personal Search Engine. https://web.archive.org/web/20060220233451/http://www.mojeek... (first capture Feb 2006).
We are currently bringing this back (in Beta). RollYo innovated too (private Beta August 2005). Google Custom Search launched in October 2006. So there were at least 3 services that predate Blekko (2010).
For the record, when I said "long before Kagi was even announced to the public." I meant exactly what I wrote, not that Kagi did not exist, it did.
Kagi was announced to the public in 2019, long before public beta release this June. I understand it is kind of hard to track small, bootstrapped startups with no mainstream exposure, but as I said this is for historic accuracy.
My extension has existed since 2012: https://chrome.google.com/webstore/detail/search-filter/eidh...
In June 2020, I added support for external (collaborative) filter lists which are incredibly similar to Goggles.
Not saying anything was copied, just that it might be nice to cite/reference prior work.
This sounds like an opportunity to create white and blacklists in cooperation among various projects.
# Make these domains stand out in results +en.wikipedia.org +stackoverflow.com +github.com +api.rubyonrails.org # SPAM - never show these results experts-exchange.com # Pull filters from external source @https://clobapi.herokuapp.com/default-filters.txt
This default list is the only one I distribute but users have come up with own lists.
It would be nice to have a Github repo with such lists (or meta lists: the @ syntax works recursively, allowing lists to import other lists).
Your suggestion of having a standard for the list syntax is interesting.
I hadn't heard of Kagi before until recently, and only just seen this Brave Goggle feature, but I had a very similar idea with my own search https://namusearch.com/ - Search using user defined url lists which can be publicly shared
I think I initially had the idea by thinking about uBlock but for Search.
This paper proposes an open and collaborative system by whicha community, or a single user, can create sets of rules and filters, called Goggles, to define the space which a search engine can pull results from. Instead of a single ranking algorithm, we could have as many as needed, overcoming the biases that a single actor (the search engine) embeds into the results ... Such system would be made possible by the availability of a host search engine, providing the index and infrastructure, which are unlikely to be replicated without major development and infrastructure costs.
Unironically, a multitude of biases in search engine results is just what we need. And the difficulty in building your own index is just why we don't have it. I would love to have my own version of google-without-the-stupid-stuff according to people I specifically identify and respect.I don't think we yet appreciate the importance of good bubble hygeine on mental health, and this is a bespoke bubble construction kit.
Goggles have a lot of promise. I'm excited about the concepts Brave is bringing forth.
The idea is amazing. Just the "no pinterest" and "copycats removal" examples have me super excited.
https://search.brave.com/goggles/discover?goggles_id=https%3...
https://search.brave.com/goggles?q=long+tail&source=web&gogg...
"...don't touch it...."
I would love a "Product Reviews" goggle that helps find legitimate reviews or recommendations without having to narrow to somewhere specific like reddit.
This is great feedback, thank you! We already have a few ideas on how to make Goggles more powerful in the future, and the things you listed sound very interesting. Would you mind elaborating a bit on how you would see yourself using these features in a Goggle? (e.g. ads count, linking to high-quality content, etc.)
$downrank=%adcount%
www.amazon.com*tag=$inlinks,downrank=2
news.ycombinator.com$inlinks,boost=1I use the uBlacklist addon for that
However, I will not recommend it to my father/mother. I think it is too geared towards techies, and full of traps to confuse people less involved in online tech culture. All this talk of crypto/wallets/coins/ads and the associated buttons and icons everywhere will just confuse the hell out of them. You need to spend a bunch of time disabling stuff or setting things up that they will just not do. Kind of a shame, because the out-of-the-box experience is really good for privacy and ad blocking, etc...
When I was teaching programming students were bewildered, befuddled, and confused by the incorrect, out-of-version, and incomplete technical information which litters the web.
I have often wondered about creating a curated collection of known-useful information for students to search.
This could be the start of something good.
This name is surprisingly close to their largest competitor, interesting choice.
It'd be nice if Goggles could be "additive" and you could subscribe to them; I'd love an "anti stackexchange spam" one.
https://chrome.google.com/webstore/detail/hypersearch/feojag...
It's 100% open source and client-side https://github.com/abhinavsharma/hypersearch
I wrote a post explaining why we think this is a more pragramtic approach here https://abhinavsharma.com/blog/google-alternatives
Goggles is also open-source, the main difference of the approach is that Goggles collaborates with Brave Search index while your approach uses any search index as source of URLs.
That difference is fundamental, a client can use a host search engine to do query expansions to build a large result set and then apply the filters and boosts that the user defines. However, that recall set is going to be in the order of hundreds URLs (more will take either too long or the client will be blocked); and I assume it would be challenging to apply thousands of filters at once. The smaller the result set, the smaller is the effect or benefit of the user-defined rules.
Goggles, because it collaborates with the host search engine — as of today only Brave search — can apply the filters and boosts to a recall set of tens of thousands of URLs. So the net effect of such rules is much larger.
Goggles is a bit more complicated, but it's for a reason.
Disclaimer: I work at Brave.
I love Brave as a search engine but these statistics are misleading. They fail to account for the fact that there were only 147 million users in 1998, but today we have 5 billion users.
As with most of Mozilla's problems, it seems the impediment is the sweet sweet Google billions rolling in every year without any regard to the work they're putting in.
>$boost=XX—is used to alter the ranking of specific results by XX (e.g. $boost=1 would not alter the ranking, while $boost=2 would make a result two times more important).
But then we have this in the example:
>! Generic boosting
>rust$boost=1,site=rs
So this line is a no-op, because it uses boost=1?
In this case it is needed because there is a generic '$discard' rule in the Goggle (which means: discard any result that does not match any other instruction from the Goggle; you can see this as a 'default action' applied to results if they are not caught by any other instruction).
Using a 'boost=1' allows you to keep some sites, that you don't necessarily want to boost more than their "natural ranking".
We have a bit more info about that in the "Getting Started" guide here: https://github.com/brave/goggles-quickstart/blob/main/gettin...
I hope that helps!
How can we accomplish this?
I'm not sure if it's possible or feasible.
But i want the discussion to happen. Maybe someone will eventually find a clever way to solve this or maybe it's just a matter of time until our cpus/bandwith/storage is good enough for this to work
---
[1]: https://github.com/brave/goggles-quickstart/blob/main/gettin...
The idea of having a separate Goggles aggregator was also mentioned yesterday; imagine a site when you can click a few check-boxes of Goggles you like, then get a new link with the combination of all of them (you could then submit this link in Brave Search to use the aggregated Goggle).
In any case, applying multiple Goggles at once is something we have discussed internally, we're not yet planning on adding this but feedback like yours is super valuable for us to decide on the next steps.
[1] https://github.com/brave/goggles-quickstart/blob/main/faq.md...
The Python programmer community, the Monty Python fans and owners of actual Python snakes probably have mutually useful lists of sites.
I don't really want to exclude many of the copycats/spammers, I just want to replace them with the original or my preferred source for that info, which I don't think this does?
Imagine being able to select a bit of code in your editor, and getting an explanation of what that code does based off the content of all of the accepted answers on Stack Overflow.
It is true that there are currently less examples of negative boosts, but it is supported by Goggles thanks to the 'downrank' option.
For example:
$downrank,site=example.com
If you want to be stricter, you can also discard sites completely of course:
$discard,site=example.com
I hope that helps,
This is the kind of feature that seems obvious in hindsight.
There are still some features I miss from FireFox (forget site, select all, etc.) but otherwise Brave has worked pretty well for me!
I mean you are gonna get the official Toyota website on top anyway, for obvious reasons. Showing it as an "ad" to users, just to milk some extra money from advertisers like this makes the entire thing feel like a scam.
I get a "204 No content" response.
I don't really have any good use cases for this but it'd be funny, to me at least.
Does anyone have any insight on the performance cost to the search engine here if lots of users had large rules and filters files? Does this add a significant cost per search?
Democracy dies in darkness, a line recently adopted by the Washington Post as their slogan, warns us that unless people are informed with facts and truth, no true democracy is possible. Those who benefit from darkness have always tried to control media in order to control and manipulate public opinion with propaganda. Until recently, propaganda has been the exclusive domain of nation-states or state-sponsored actors through mass media [19]. With the mass popularization of the Web in the last two decades and the subsequent privatization of it by big platforms like Google, YouTube and Facebook, the paradigm has changed. Propaganda is no longer a tool of an elite, but it has been commoditized to the extent that it is as accessible as advertisement, becoming a weapon that too many actors have access to. One must appreciate the irony that those most vocal about the risks of propaganda are those who controlled it in the past. Nevertheless, the risk of fake-news—a neologism created to mitigate cognitive dissonance—cannot be ignored [5, 6, 30, 33, 36]. It is dangerous for a society if people living in it cannot distinguish between facts, opinions and outright misinformation. Although this danger has always existed, today the situation is dire if only because quantitative becomes qualitative and although all information is theoretically available, in practical terms it is not.
Doesn't this seem a bit overblown for what is ultimately just a tool to make soft white lists for search results? I feel like Brave has this habit of framing mundane software as some sort of weapon in a grand ideological war. You're just letting people filter Pinterest out of their search results. Chill.
IMO, I don't think it's going to be a good thing for users if chromium ends up being the last thing standing.
So it's doubly punny.
No, it was an accident.
Also, seeing the initial examples from their beta (e.g. boost tech content or boost left-wing news sources) makes me a bit weary about its usefulness (if most people, rather than creating their own goggles end up using some prepackaged ones).
I mean, its at least transparent, people are aware of it or have to opt in.
It feels like a weird compromise in terms of misinformation, cultural division, etc... but letting people choose which kool-aid they're drinking instead of letting the "totally not hand tweaked for edge cases in favor of the creator of "the algorithm"" to decide. Out of the handful of "ideas" around wrangling the the trashfire that is the modern internet, this seems like the most sane the best fit solution so far.
In terms of it being an actual search tool for finding information, answers, documentation, references that are actually relevant or useful it sounds insanely useful.
Narrow down searches for anything + "datasheet" to manufacturer websites and a handful of non paywalled datasheet catalogs - fuck yeah!
I forsee lots of angry website owners who run fluff content or bury reposted useful info under mile thick layers of ads.
I built a tool to recursively scrape RSS feeds from web pages linked to twitter bios. You pass in a "root" trusted user, and look in their bio and every "followee"'s bio for a website, then look for anything "rss" "feed" "atom" or "xml"-y on the link itself or in the domain's sitemap.xml.
Surprisingly very useful. There's a decent amount of value in twitter's content, but arguably much more value in the followee network of "smart people", and the websites "linked out" from their profiles and tweets.
Reddit, similar but in a different way, filters itself into variously useful, well-moderated communities. Top X posts of subreddits A, B, C is a great heuristic for getting 90% of the value out of reddit with very little of the toxicity.
You needn't limit yourself to r/all and the twitter equivalent! :)
Does it put me in a bubble? Yeah. But it's a bubble of my own design, and it's a pretty nice place from my POV, that reflects my real world interests, hobbies, etc.
I see goggles as a parallel system for "I'm looking for something specific" search.
I think Goggles might turn out to be a mild retardant to toxic bubble formation, because toxic content is rocket fuel for the kind of "engagement" metrics that google/youtube/twitter/fb have spent the last decade min-maxing.
IMO, existing "search" players (Google, FB, Twitter) are culpable for toxic echo chamber formation, only to the extent that they push toxic content "by default". By allowing users to fine-tune their weights, Brave+Goggles is attempting to dislodge this norm by introducing user feedback into the equation. Like you, I think it's a good idea.
Some people will always opt-in to toxic bubbles, but I think that's more of a human/society problem and not one for a search player to solve.
Everyone has context bubbles already. You have work, outside of work, your hobbies both online(IE HN) and offline. You influence these bubbles composition. Hobbies in particular are entirely self selected.
But you are aware of which context you are communing in and are free to move between them and outside them into the general public sphere.
I see this as more of the same. As you say, "its at least transparent, people are aware of it or have to opt in"
Interestingly by the comments here, a big driver for these self selected bubbles are so that people can avoid the advertising they don't wish to see. IE Pinterest and other SEO spam. If "self selected bubbles" are bad, then "targeted relevant advertising" is worse.
The solution put forth by Goggles is to make such biases explicit and transparent, and to make them switchable. No one can honestly claim to be objective if they are using the "MAGA" Goggle. That said, people are particularly good and wearing biased goggles and believing they aren't wearing any goggles at all, or that they are wearing "the one true" goggle.
Anyway, read the white paper, which discusses your point. https://brave.com/static-assets/files/goggles.pdf
There's no objective measure of what site is 'good', and what site has been 'SEO-gamed' to rise to the top of <arbitrary search engine's algorithm>. The measures are subjective, and people complain when changes to the algorithm aimed at punishing black-hat-SEO also punish their website (Rightly or wrongly).
Not to mention the problem of using black-hat-SEO to punish a competitor (By creating scummy links to them, that make it look like they are trying to game the search engine.)
... 100,000 words later ...
I don't think it will be very hard get right.
A very very significant example is the primary language group exposed on the site.. for example, Cantonese ? I personally support the rights of "minority" languages like Gaelic and Welsh in the European setting, Tamil for South Asia, things like that.. make it so..
"Use our brave search and escape the leftist google agenda!" or such.
The obvious questions of morality aside, my perception of those stories is that most of the victims are hopelessly addicted to the feeling being righteous and correct and part of something bigger to the point that it takes over their whole lives. They end up a husk of a person all for a fake cause. But what is interesting is that after you take it away they generally don't find something else to latch onto, they slowly go back to having hobbies and normal conversations and normal relationships with their family, friends, and coworkers.
This is of course all anecdata, but if Goggles are another tool for giving people their loved ones back I think that's worth weighing as part of the equation.
The article also mentions that Goggles will not stop polarization, it suffices to not exacerbate it.
No technology/system on any period of time has been able to suppress it, censorship included.
Disclaimer: I work at Brave search
By the way, cool feature :)
Allowing the users to choose their own filters will allow advertisers to actually read the market based on the the sites that are whitelisted in the filters instead of shotgunning ads at any website that claims to be relevant to the target demographic they can actually see the popular ones that users choose based on these filters.
its still targeted advertising, but abstracted one layer away from the actual user so that the targeted ads don't have to be as fucking creepy with all the data they're gathering on people. With the customer choosing what sites they want to see results from, the advertisers can stop wasting money on ad revenue for click farms that everyone hates.
its a better deal for advertisers, and provides a better end result for the user and some degree of transparency.
It can't make people accept inconvenient/uncomfortable facts, it won't solve any political problems. You can lead a horse to water but can't make it drink, you can point out any number of problems to a person but you cannot make them care.
edit- relevant to solso's comment about an active choice
The active choice democratizes the ad market allowing users to choose, instead of the passive route of allowing an algorithm to coerce the market.
How is it less creepy if the advertisers still end up with all the same data? Whether they snoop on my browsing history or snoop on my search rankings doesn't make any material difference to me. The problem is building a profile on me through snooping.