Marginalia: DIY search engine that focuses on non-commercial content
search.marginalia.nu
search.marginalia.nu
If anything it's running faster now. All you've done is warm up the caches and given the JVM a chance to optimize the hottest code.
(real talk the SSDs are running pretty near 100% utilization though)
Do you happen to have a writeup somewhere of your tech stack?
But besides that, there's still a lot left to be desired when it comes to how it actually works. Not everything is easy to glean from the code alond.
[1] https://sparkjava.com/ I don't use springboot or anything like that, besides Spark I'm not using frameworks.
This is $5000 worth of consumer hardware, give or take.
But AWS and its competitors don’t have an offering even close to comparable to what you can get in a commodity server. A 1U server with one or two CPU sockets and 8-12 hot-swap NVMe bays is easy to buy and not terribly expensive, and you can easily fill it with 100+ TiB of storage with several hundred Gbps of storage bandwidth and more IOPS than you are likely able to use. EC2 has no comparable offering at any price.
(A Micron 9400 drive supposedly has 7GBps / 56Gbps of usable bandwidth. 10 of them gives 560Gbps, and a modern machine with lots of PCIe 5.0 lanes may actually be able to use a lot of this. As far as I can tell, you literally cannot pay AWS for anywhere near this much bandwidth, but you can buy it for $20k or so.)
True but this also depends on design decisions AWS made with regards to those volumes.
Indeed it could be that the volume is internally (at the hypervisor level) redundant (maybe with something like ZFS or other proprietary RAID), but there's no way to know.
Furthermore, AWS doesn't allow you to really keep a tab or reservation on the physical machine your VM is on - every time a VM is powered up, it gets assigned a random host machine. If there is a hardware failure they advise you to reboot the instance so it gets rescheduled on another machine, so even though technically your data may still be on that physical host machine, you have no way to get it back.
AWS' intent with these seems to be to act as transient cache/scratchpad so they don't seem to offer much durability or recovery strategies for those volumes. Their hypervisor seems to treat them as disposable which is a fair design decision considering the planned use-case, but it means you can't/shouldn't use it for any persistent data.
Being in control of your own hardware (or at the very least, renting physical hardware from a provider as opposed to a VM like in AWS) will indeed allow you to get reliable direct-attach storage.
https://instances.vantage.sh/aws/ec2/im4gn.8xlarge?region=us...
I can buy a rather nicer 1U machine with substantially better local storage for something like half the 1-year reserved annual cost of this thing.
If you buy your own servers, you can mix and match CPUs and storage, and you can get a lot of NVMe storage capacity and bandwidth, and cloud providers don’t seem to have comparable products.
--edit--
I forgot to add the following: that's 32k if you run the system 24/7. Usually it's up for a few hours per month, so you end up paying maybe 2k for the whole year.
How do you figure a $5k annual cloud spend is cheaper than ~150€ per month?
~150€ in cloud costs is cheaper than $5k cost to buy the hardware the guy has in his living room.
But with dedicated servers, are we really talking cloud?
My point was rather underlining the absurdity of using cloud for everything.
Herzner is just an example of a dedicated server provider. There are others, some in the same price range, others a bit more.
As an aside to my point, it is often cheaper and more flexible to use dedicated servers than you buy and collocate your own hardware.
(No affiliation.)
The other common cause of issues is things like crypto which they don't want in their network at all.
This will sound like I am downplaying what people have exprerienced and/or being apologetic on their behalf but that is not my intention. I am just a small time customer of theirs. I've had 1 or 2 dedicated servers with them for many many years now upgrading and migrating as necessary. (It used to be that if you waited for a year or two and upgraded you'd get a better server for cheaper. Those days are gone.)
I've only dealt with support over email where they have been both capable and helpful, but what I needed was just plugging in a hardware kvm switch (free for a few hours - i never had to pay) or replacing a failing hard drive (they do this with zero friction). Perhaps I am lenient on the tech support staff. After all they are my people. I've been to a few datacenters and have huge respect for what they do.
On the presales side they seem to reply with a matter of fact tone with no flexibility. They are a German company after all.
I'm a bit wary I'd get lumped in with the crypto gang. A lot of what I'm doing with the search engine is fairly out there in terms of pushing the hardware in unusual ways.
It would also suck if there ever was a problem. The full state of the search engine is about 1 Tb of data. It's not easy to just start up somewhere else if it vanished.
Most pages with images do lazy loading so I'm not hit with 30 images all at once. They're also webp and cached via cloudflare, softens the blow quite a lot.
FWIW I'm going commando with no ECC ram too.
One thread FTW. :)
If that happened at the SaaS company I worked at previously, it would be a bloodbath. The churn would be huge. And our customer's customers would be churning from them. If that happened at a particularly inopportune time, like while we'd been raising money or something, it could potentially endanger the company.
(I'd like to stress again this is not a criticism of HN/dang, but just to illustrate a set of requirements where huge AWS spends do make sense.)
A good example is the obsession with "fast" web frameworks on your application servers, completely ignoring the fact that your database will be the first to give up even in most "heavy" web frameworks' default configurations without any optimization efforts.
Every time I deploy a service it goes down for anything between 30 seconds and 5 minutes. When I switch indices, the entire search engine is down for a day or more. Since the entire project is essentially non-commercial, I think this is fine. I don't need five nines.
If reliability was extremely important, scales would tilt differently, maybe cloud would be a good option. A lot of it is for CYA's sake as well. If I mess up with my server, that's both my problem and my responsibility. If a cloud provider messes up, then that's a SLA violation and maybe damages are due.
Much more complex systems do not perform as consistently as simple ones, and they are exponentially harder to debug, introspect and optimize at the end of the day.
Yep, this is a good example of warping the Feeling Lucky pattern into a really neat little discovery tool.
IMO it would even be cool if the site was this part first, oh and hey it's also a search engine.
(While I'm random-ing: The Arch Wiki is in there? Seriously? Just for that, I propose that it either be skinned to max vaporwave, or host a webcam pointed at a Manjaro machine, or both...I'll be waiting over here, downloading 4.1 GB of marginalia for my AUR build of PCManFM)
One big difference then from now is that you basically need a PhD in the Canvas API (or WebGL or whatever) to accomplish something a 5 year old could do in Flash. Web design was a lot more accessible. You didn't have to worry about responsive designs and fluid layouts. You could just position:absolute everything and that was kinda fine.
It's trivial to have a "weird" position:absolute design with a break for mobile that switches to a more fluid layout. Desktop users can have their "weird" layout but I can still read the page on mobile and you can readily crawl and index it.
People moved away from design tools like DreamWeaver that helped make "weird" stuff and instead installed WordPress or some CSS/JavaScript framework that just bakes in all the "boring" fluid layouts.
You're not necessarily wrong about Flash in terms of design or creation but your search engine wouldn't be terribly practical if everyone was still using Flash for everything. Flash allowed content packed inside SWFs but also allowed fetching external resources. You wouldn't be able to index any of that unless your crawler executed the Flash and/or inspected all the URL references for external resources.
Flash created an inaccessible deep web just like today's JavaScript website-is-an-application "sites".
Don't get me wrong, I love the old web with quirky table-based layouts, "unofficial" fansites, and personal homepages hosted on forgotten university servers in a closet. There was a vibrancy that's largely missing from today's web.
I think a big change has been tools have become more geared for boring than the creative and people treat content on the web as a side hustle. Google et all haven't helped by favoring recency over other relevance factors.
Reddit even has some kinda-similar subs.
(There's also explore2.marginalia.nu which is not even limited to websites with a screenshot)
You could pretty trivially shard the index by `hash(domain) % numShards`. There's no support for this because I literally only have this single server, but it wouldn't be much work.
That said, it's gotten way better at finding stuff with the last few releases.
I also think that having "a google", one central search engine, is inherently a bad thing for the health of the Internet. It drives a lot of this search engine spam epidemic we're seeing.
A broader and (IMO) more interesting problem is Internet discovery.
I bet one could make a facinating ranking algo that groups sites by subject then sort them by nr of links to others in that group.
So the perfect SEO would be to have a blogroll at the top of the left menu with every related website in it.
i.e. 3 stores sell the same item. Nr 1 is the one linking to the other 2. Extra points for linking to that specific product page.
Even besides the point that the websites they indexed were a lot less adversarial, they put a lot of emphasis on indexing academia, and were outspoken against what came to be their present mixed motives[1].
https://search.marginalia.nu/search?query=spanish+rice+recip...
[1] https://github.com/MarginaliaSearch/MarginaliaSearch/blob/ma...
Made me chuckle a bit.
I use it mostly for tech/programming/FOSS stuff. Especially for programming topics it can be good for filtering out all the ‘w3schools’ type of blog spam that just floods Google’s results.
Marginalia Search has received an NLNet grant - https://news.ycombinator.com/item?id=34945541 - Feb 2023 (17 comments)
A Theoretical Justification (2021) - https://news.ycombinator.com/item?id=32586273 - Aug 2022 (22 comments)
The Evolution of Marginalia's Crawling - https://news.ycombinator.com/item?id=32565052 - Aug 2022 (22 comments)
Botspam apocalypse - https://news.ycombinator.com/item?id=32339314 - Aug 2022 (346 comments)
Marginalia Goes Open Source - https://news.ycombinator.com/item?id=31536626 - May 2022 (72 comments)
Uncertain Future for Marginalia Search - https://news.ycombinator.com/item?id=31200319 - April 2022 (37 comments)
Marginalia Search: 1 Year - https://news.ycombinator.com/item?id=30823481 - March 2022 (29 comments)
Show HN: Marginalia – Exploration Mode - https://news.ycombinator.com/item?id=30047455 - Jan 2022 (53 comments)
A search engine that favors text-heavy sites and punishes modern web design - https://news.ycombinator.com/item?id=28550764 - Sept 2021 (717 comments)
(just as a reminder, these lists are only to satiate curious readers - there's no reproach for reposting! Reposts are fine on HN after a year or so: https://news.ycombinator.com/newsfaq.html)
Is there a way to donate money?
I kept looking for a "Donate" button :-P
Thank you!
Later my boss asked me to look at this web thing that he had heard about. I fired up telnet and eventually found an on ramp to CERN. To me it looked rather like everything else but I'm not exactly a rocket scientist!
I've got some ideas in the pipe, but haven't had the time to give them enough polish that I'm happy with them.
This is an early draft: https://imgur.com/a/vMVO7CK
The draft looks nice. The text colour is a bit hard to distinguish from the surrounding background, and I don't have any eye conditions.
Now, the bad news is there is an associated discussion going on over in the "DDG integrates GPT" thread where the intersection of "I want to pay" and "I don't want to GPT anything" is damn near nil :-(
Honestly, this makes me really happy. I would prefer that my traffic be driven by curated search engines, even if I get less traffic.
Also, I use this. I think it's great.
"https://frontendmasters.com/courses/complete-react-v5/gettin... Getting Started with Pure React - Complete Intro to React, v5 | Frontend Masters The "Getting Started with Pure React" Lesson is part of the full, Complete Intro to React, v5 course featured in this preview video."
It's in part a measure to limit the scope of the project (the entire thing runs off a single PC), but it's also hard to build a good language model for a language you don't speak, and I only speak English and Swedish. But if the project grows, gets more hardware, and contributors that speak other languages, then maybe this will change in the future.
If I'm searching for roman coins I certainly don't want to find commercial sites selling them (I know what those are), or even the well-known online national collections or auction archives... I'd like to be able to find the specialist sites built by collectors (and maybe academics) that are non-commercial and way more interesting.
In the early days of the internet some specialist content/pages were organized into "web rings" each linking to each other, but nowadays we're mostly relying on search to discover new pages, and it seems a lot of the hobbyist content is way harder to find, assuming it's even out there.
#1: http://www.romancoins.info/Content.html
#2-4: were not very good
#5: https://www.forumancientcoins.com/dougsmith/voc1.html
#6: https://www.cngcoins.com/Greek+and+Roman+Coins.aspx
#7: https://www.crystalinks.com/romecoins.html
If you search for specifically the 'as' it may be eaten as a stop word :-/
BTW #1, 5, 6 are all good sites, but those are very mainstream - those will be top links in Google as well. #6 is purely commercial - an auction house. #5 is a coin dealer's commercial site, but has good collector resources (discussion board, Wiki, collectors galleries) as well.
Some examples of other non-commercial roman coin hobbyist sites (that will also rank fairly highly with Google) are:
augustuscoins.com wildwinds.com beastcoins.com www.notinric.lechstepniewski.info https://www.nummus-bibleii.com/
I'm at work right now, so these are just some examples off the top of my head. I can give more examples later if it's useful. Some of these site will include links to other collector/hobbyist sites.
constantinethegreatcoins does show up for 'imp constantinvs' though.
I've got 128 Gb RAM and more would be better. I run a test instance on 32 Gb though.
list of Italian generals
list of CPU architectures
list of positive rights
etc...
Return no relevant results. Perhaps it's not giving enough weight to the 'list' aspect?
The query processing is fairly crude. For better or worse, it doesn't do much special processing. Which means you basically need a website that repeatedly says "list of CPU architectures" to rank well.
Most of the pages that contain such a title are also actual lists. The index de-prioritizes documents that are mostly lists or tabular data, as they're rarely very often false positives as they often contain repeated words.
Even if the word 'list of ...' only appears once, it wouldn't be filtered out, right?
Out of 100 million pages, it seems like there could easily be a few hundred thousand with lists.
search for just “3d box” or something like that.
Search for "draw 3d box" or "draw a cube" and it starts giving results.
I get tricked so often into clicking a news snippet offered by Google only to then land on a site which not only presents me a paywall, but also does want me to accept their cookie policy before they present me the paywall.
It makes me angry every time anew.
How do you discover relevant new domains?
It can rank websites even if they aren't indexed, based on who is linking to them.
Vanilla PageRank can't do this very well. Domains that aren't indexed don't have (known) outgoing links, in the periphery of the rank. There's a some tricks to get these to not mess up the algorithm completely, but they basically all rank poorly. That's even without considering all the well known tricks for manipulating vanilla pagerank. The modified version seems very robust with regards to both problems.
[1] https://memex.marginalia.nu/log/73-new-approach-to-ranking.g...
Not exactly this, but close enough: https://memex.marginalia.nu/links/bookmarks.gmi
I've changed the crawler design a couple of times, but the principle for growing the set of sites to be crawled is to look for sites that are (in some sense) adjacent to domains that were found to be good.
Internet arguably doesn't even have a size. You can construct a website that's like n.example.com/m which links to '(n+1).example.com/m' and 'n.example.com/(m+1)', for each m and n between 0 and 1e308.
Dunno about the others, but my crawler has a set depth it will crawl. It'll BFS for like 1000-10000 documents depending on some factors.
2) Yes. Everything is in-house.
Do you build a word index by document and find documents that match all words in the query?)
Yeah. It's actually got three indices;
* One is a forward index with `document id -> document metadata`
* One is a priority term index with `term -> document id`.
* One is a full index with `term -> (document, term metadata)`
They're all based on static b-trees.
https://search.marginalia.nu/site/www.thran.uk
https://search.marginalia.nu/site/wmw.thran.uk
Only this is possible as long as the index knows about the domain. Yours are, but if not, anyone can shoot me an email or something and I can poke them into the database.
The limitation for known domains is in place to avoid abuse.
Universal broadcast does not work (beneficially for society) in an industry built to monetize reach.
Everyone is entitled to their opinions, but voices are not equal in utility or worth.
"2) who decides which voices do matter?"
This is always the problem, isn't it? I don't have an answer for you.
We could get away from it only if we figure out an answer to the second question but I suspect we'll never get to an answer.
Only in that reach will continue to be monetized regardless of its impact on society.