Mwmbl: Free, open-source and non-profit search engine
mwmbl.org
mwmbl.org
From:
MWMBL
[Search on mwmbl...]
Welcome to mwmbl, the free, open-source and non-profit search engine.
You can start searching by using the search bar above!
Find more on
[Github] [Wiki]
To: MWMBL
[Search on mwmbl...]
A free, open-source and non-profit search engine.
[Github] [Wiki]Also, your own posting appears to be missing from the index: https://mwmbl.org/?q=mwmbl+ycombinator
(and, yes, another vote for changing the domain name; you can have a quirky project name, but if I can't remember the cat-walking-on-keyboard domain, I'm not going to use it)
This is a very interesting idea that other search engines have tried before. Actually, the Brave search engine is built over Cliqz[6] that implemented this same idea but *without* the user's consent.
Copy pasting from an old comment I made about this "human web" crawler idea:
Both PeARS[1] and Cliqz[2] tried to do that. Both got direct support from Mozilla[3][4] but it looks like neither really kicked off.
PeARS was meant to be installed voluntarily by users who would then choose to share their indexes only to those they personally trusted, so the idea is very privacy conscious but also very hard to scale.
Cliqz, on the other hand, apparently tried to work around that issue by having their add-on bundled by default in some Firefox installations[5] which was obviously very controversial because of its privacy and user consent implications.
I still think the idea has potential, though, even if it's in a more limited scope.
[1] https://github.com/PeARSearch/PeARS-orchard
[2] https://cliqz.com/en/whycliqz/human-web
[3] https://blog.mozilla.org/press-uk/2016/06/22/mozilla-gives-3...
[4] https://blog.mozilla.org/press-uk/2016/08/23/mozilla-makes-s...
[5] https://www.zdnet.com/article/firefox-tests-cliqz-engine-whi...
[6] https://www.theregister.com/2021/03/03/brave_buys_a_search_e...
One's normal phone and laptop is actually only a fraction of uses, and even one's normal device isn't just one thing that needs to be done one time in grade school and then set for life. It's a dozen different things, and they are all perpetually rotating, and most people are not highly optimized with profiles they actually export and import.
This idea is great but it's going absolutely nowhere without a better understanding of actual humans.
I'm an actual human, I use alternative search engines, I don't memorize their full names, and the only thing perpertually rotating is the planet
There are actually browsers also built in to 4 TVs, also in the rokus and google TVs attached to those same TVs, also in the Xbox and ps3. But I won't even count any of those. I have actually used them, but I'll give you those for free since I don't actually use those browsers very much.
Also that just reminded me that all of the old devices are fairly regularly getting reinstalled with some new version of a linux or bsd distro fresh every time I pick one back up, so, no configured profiles.
The windows partition on my main machine is frequently reinstalled since I experiment with trying to use either a partition or Frameworks custom usbc module or a regular usbc external drive, or just a partition on a bigger faster external drive. That's one physical device but a few different OS's, and most of those OS's besides my main daily driver get moved around and reinstalled a lot so they are always new and unconfigured., yet, I still need to use them, and that means I use a browser to search from within them.
My kobo, and 3 or so other eink readers. Which, again, occasionally gets reinstalled, so even the one device needs to be set up more than once.
The only reason I don't have to set up a new phone every 6 months is because I value a headphone jack more than most everyone else. So if you would say my usage pattern is an outlier, I would say, 1 so what? Outliers exist and could even be argued to outnumber the center peak of the bell curve, and 2 some of my outlier usage pattern goes the opposite way, like using the same phone for 5 years.
And then of course I use many machines which are not mine. And this is not even counting that my work used to involve some amount of user it support where I would use a users desk or a hot desk at a customer site, I just mean my own personal normal activity is on many other machines besides my own, including relatives, friends, & public machines.
I had to type "google" (back when I used google primarily, and it wasn't already everyone's default) countless times, even though it was the home page on my own main machine.
This question didn't really even deserve the dignity of any answer it is so obtuse.
But even then compared to that effort remembering a new word is trivial
> So if you would say my usage pattern is an outlier, I would say, 1 so what?
I'd say it's not relevant to this conversation where you barge in with an uber-confident "100% wrong" when it's only "1 person" wrong
Show HN: I'm building a non-profit search engine - https://news.ycombinator.com/item?id=29690877 - Dec 2021 (199 comments)
Much is at stake in this arena.
ChatGPT's recent huge success in performing a specific tasks previously within the domain of Google by doing something other than they are is a good example of this.
Google used to be good at it but it's now utterly befuddled by specificity and returns such garbage that I had given up.
But the form of "I'm doing this, I'm seeing this and I'm wondering if X is possible" chatgpt is solid on that - basically a personal stack overflow
The biggest problem used to be when seemingly the whole internet was satisfied with an answer that is extremely wrong and broken when you do it.
Chatgpt can work though this without getting into a weird markov cycle maybe half the time which is great.
Patterns like "Hey I tried that. It still doesn't work, can you give me another option"
This not at all trivial but quite possible[1], but ChatGPT will in 100% of the time either hallucinate APIs, disregard the instructions to not use Hadoop or give otherwise plausible but incorrect-looking answers.
The trick is that it isn't doable by simply finding the correct dependencies and API calls, you need extract and override filesystem classes from the Hadoop project to cut those ties.
I don't know if "think hard" does anything but it seems to work and if I was the one making chatgpt I'd certainly have configurable keywords like that to tweak the generation settings - mostly so I could skate by on cheaper queries 90+% of the time and then have a fix when they fail
It would be very easy to improve Google's search result quality by removing their promoted results, and then penalizing websites with ads and adtech.
I strongly believe the smartest interfaces have the right fidelity to empower the user to effectively control the tool.
These parameters need to have the right dimensionality, faceting and perimeters to be expressive in this way.
I know you've got your own semi famous search engine and I express these ideas with that known
wget https://data.commoncrawl.org/crawl-data/CC-MAIN-2018-17/segments/1524125937193.1/warc/CC-MAIN-20180420081400-20180420101400-00000.warc.gz
wget https://data.commoncrawl.org/crawl-data/CC-MAIN-2023-23/segments/1685224643388.45/warc/CC-MAIN-20230527223515-20230528013515-00483.warc.gzAs for the rest, in order to perform well, the indexer needs to be built specifically tailored to the what the search engine is doing. Often you're scrounging for places to cram in individual bits to encode some additional piece of information about the term.
If a DBMS tries to support every use case, a search engine index does the opposite, it supports a singular use case and cuts every corner imaginable and then some to make that happen with as much resource frugality as possible.
But when it does find something it is very quick! So I'll give it a go.
Other random examples: search for "2023" and the very first link is "2023 Pomeroy College Basketball Ratings". Search for "iphone", and the 5th link is a page about iPhone 6s that was last updated in 2021. Typos don't work: "haker news" has only one result, a hungarian press article.
https://www.romeartlover.it/Vasi60.htm https://www.maquettes-historiques.net/P19b.html
But the crawler seems to be lacking quite a bit. For my first search (current work problem) "rust json diff" it only found 6 links, only one of which was a rust crate. Unfortunate.
Second Search: "black sabbath sleeping village lyrics" only gave 2 results, only one of which was correct.
Also the repo is missing the SearXNG[1] search engine.
Either would be sad, because the world needs more open source search engines.
Thanks for your encouragement. Would love to have a chat some time.
Bye bye.
You do not need "scripts" to turn the text string I'll supply into a list of candidate links. How can you not understand this basic accessibility foundation?
Particularly I am having a great time reading the crawler extension source-code: https://github.com/mwmbl/crawler-extension
This website requires you to support/enable scripts."
JSON results, no Javascript
https://api.mwmbl.org/?search=search+the+web+without+javascr...
Dal ati! We really need open source alternatives to Google.
In general hash map table index designs don't tend to be very efficient. If you use a skip list or something similar, you can calculate the intersection between sets in sublinear time.
[1] https://nlp.stanford.edu/IR-book/html/htmledition/faster-pos...
>Tangerine Business Savings Account: This account offers a high interest rate of 2.65% to 3.25% on your balance, no monthly fees, no minimum balance requirement, unlimited transactions, free e-transfers, and access to over 3,000 ATMs.
>Wise Business Account: This account offers low-cost international payments in over 50 currencies, no monthly fees, no minimum balance requirement, free local transfers, free debit card, and access to over 10 million ATMs.
>BMO eBusiness Plan: This account offers no monthly fees, no minimum balance requirement, unlimited transactions, free e-transfers, free cheque deposits, and access to over 3,500 ATMs.
>RBC Digital Choice Business Account: This account offers no monthly fees for the first three months ($5 per month thereafter), unlimited electronic transactions, 10 free debit transactions per month ($1.25 each thereafter), free e-transfers, free cheque deposits, and access to over 4,200 ATMs.
The banks? In this case. Because if you do the manual searching, you will “manually” go to each bank site, go to accounts, business section and read, a good search engine will do that for me, no middle man (aka some 3rd party sites) and summarize it based on my query, a bad search engine however, will look into a 3rd party website that already created a list, recommended some based on affiliate links, boosted itself in the results by playing the SEO keywords game.
Why go with an unpronounceable name?
I mean, great that it was made, but I can't even tell people I'm using... mwumble? But it's spelled em-doubleyou-em-bee-el dot org.
> How do you pronounce "mwmbl"?
> Like "mumble". I live in Mumbles, which is spelt "Mwmbwls" in Welsh. But the intended meaning is "to mumble", as in "don't search, just mwmbl!"
UUho knows, maybe the name can uuork after all!
The marketing claim above is so far from universal truth. Choosing a "clever" versus a straightforward brand name often depends on the brand strategy, target audience, and market conditions.
But, sorry, for this example, I personally think the current brand name is atrocious.
Ah Welsh, the golden standard of phonetic spelling and easy pronunciation!
In other words, one needs a manual just to learn how to pronounce the name of that thing.
Off to a great start, aren't we?
>An explanation is at the very bottom of the github Readme
AKA the first place anyone visiting the website would look at. NOT.
Did writing "pronounced Mumble" on the landing page hurt puppies or something?
>"don't search, just mumble!"
I think half of the country is already doing that when it comes to fact-checking, and they certainly don't need a search engine for that.
And here I was thinking GIMP was a horrible name.
Make it so you can use it in a sentence to replace “Just google it”
There are no other outcomes if they don't already understand why everyone is telling them this is unusable.
> How do you pronounce "mwmbl"? Like "mumble". I live in Mumbles, which is spelt "Mwmbwls" in Welsh. But the intended meaning is "to mumble", as in "don't search, just mwmbl!"
mwmbl is a shortening of the welsh writing of https://en.wikipedia.org/wiki/Mumbles. Only tricky part is knowing that the w is pronounced as a u. Maybe it would be slightly easier if one followed the fad of leaving out vowels, but guessing a vowel and having a tricky vowel does not seem much different.
Is that really tricky? W is basically pronounced like U in English already[1]. It just looks funny when you exchange the two.
[1] e.g. say this sentence "uorld uar tuo uas the uorst"
> e.g. say this sentence "uorld uar tuo uas the uorst"
This doesn't work with the English pronunciations of the letter u from words like "uninteresting" or "mumble". It mostly seems to work with the pronunciation of "you", which does not naturally fit those letter placements.
Not knowing the proper linguistic terms, I'd consider "w" to be a modulation of a another sounding vowel by closing your lips and pressing your tongue a bit down to make room. Without a vowel to modulate, there is no sound, and so "mwmbl" is a bit of a question mark.
But most words require prior knowledge to pronounce correctly, especially in as messy a language as English.
It was 100% abandoned and I think that's a mistake. It'd be nice to explore some of those ideas again
If I want to learn how to do an INNER JOIN in MariaDB, this is the authoritative source: https://mariadb.com/kb/en/join-syntax/
The problem being that INNER JOIN isn't particularly important to that page using most IR measures of importance, it's also primarily in a <code>-block which is typically further de-prioritized. To learn that this is an important link, you need to look outside of mariadb.com.
What if instead of crawling the php generation of database rows with a bunch of cruft, the administrator published some kind of schema with scraping and querying rules and you could alternatively make a single call to capture all of the data in a sematic schema.
You can still do all the stuff you're talking about but it could make search more coherent.
An entry for that humans and an entry for the computers.
You can't trust everybody like this sure, but say imdb, discogs, wikipedia, all of which provide database dumps anyways (eg: https://datasets.imdbws.com/). That's what I'm advocating for revisiting. Lots of legit sites such as universities, newspapers, public records offices...
You could even have a search toggle "screened sources" or whatever for the ones that make the cut