Refresh this page as many times as you want, and I'll show you the living Internet: https://search.marginalia.nu/explore/random
Refresh this page as many times as you want, and I'll show you the living Internet: https://search.marginalia.nu/explore/random
One should subscribe to it out of support, hope, and to send a signal if nothing else. But it’s actually considerably better on most searches, perhaps similar to a mid-2000s Google, except with mild structure added that isn’t ads.
(You can still !g like ddg if you feel you absolutely must.)
(might make your fans spin up)
My local farm has a website where they list the stuff they have available. Meanwhile their actual scheduling and detail updates are on their Instagram because of course it is.
Who on earth would expect a local farm shop to be on par with Amazon when it comes to inventory and availability data online?
Twenty years ago you just called them for the information, and it's way better for them to broadcast it than have a hundred 1:1 conversations.
As kids are now raised on smartphones instead of the family desktop, I think they need MORE of this, not less, for at least the very important skill of typing. I wonder how many 12 year olds in america can type using the "standard" method, instead of hunt and peck.
I don't want computing to be something only known by the children of turbo nerds. I want young adults to be able to solve their own problems with computers, ie build some spreadsheets for home finance or even just be able to graph the data from one of their science classes.
As you can't develop software on phones and tablets, very few people are tinkering with software. The Pi and iOS app craze brought a momentary change, but it seems to have gone back to how it was—and worse.
Kids of today are mostly out of their depth when put in front of a computer of any description if it is beyond basic website usage. Complex program? Forget it. Decent typing skills? Forget it. Networking know-how they'd have picked up from doing LAN gaming with consoles or PCs? No chance. Change a drive? LOL.
For the handful of kids that game on PCs, they're generally not very clued up and they're just copying builds they've seen on YouTube to the word. It's a sad state of affairs.
And yes—of course, there are the kids or us turbo nerds, because of course there are, but they are so few and far between.
This is frustrating (among other reasons) because Instagram has become much more aggressive in not allowing you to even see their content without logging in. Sometimes you can see the gallery but not an individual posts, sometimes no individual posts, sometimes you can't see anything at all.
I’ve written a couple search engines. Have you tried making one with beautiful soup?
You can calculate anchor tag density across the DOM tree and prune branches that exceed a certain threshold to remove navigational elements with reasonable accuracy if that is a problem.
It's not going to be perfect, but even Google messes this up every once in a while. I wouldn't consider it a major hurdle.
edit: https://git.marginalia.nu/marginalia/marginalia.nu !!!
private static final double PRUNE_THRESHOLD = .5;
public void prune(Document document) {
PruningVisitor pruningVisitor = new PruningVisitor();
document.traverse(pruningVisitor);
pruningVisitor.data.forEach((node, data) -> {
if (data.depth <= 1) {
return;
}
if (data.signalNodeSize == 0) node.remove();
else if (data.noiseNodeSize > 0
&& data.signalRate() < PRUNE_THRESHOLD
&& data.treeSize > 2) {
node.remove();
}
});
}
private static class PruningVisitor implements NodeVisitor {
private final Map<Node, NodeData> data = new HashMap<>();
private final NodeData dummy = new NodeData(Integer.MAX_VALUE, 1, 0);
@Override
public void head(Node node, int depth) {}
@Override
public void tail(Node node, int depth) {
final NodeData dataForNode;
if (node instanceof TextNode tn) {
dataForNode = new NodeData(depth, tn.text().length(), 0);
}
else if (isSignal(node)) {
dataForNode = new NodeData(depth, 0,0);
for (var childNode : node.childNodes()) {
dataForNode.add(data.getOrDefault(childNode, dummy));
}
}
else {
dataForNode = new NodeData(depth, 0,0);
for (var childNode : node.childNodes()) {
dataForNode.addAsNoise(data.getOrDefault(childNode, dummy));
}
}
data.put(node, dataForNode);
}
public boolean isSignal(Node node) {
if (node instanceof Element e) {
if ("a".equalsIgnoreCase(e.tagName()))
return false;
if ("nav".equalsIgnoreCase(e.tagName()))
return false;
if ("footer".equalsIgnoreCase(e.tagName()))
return false;
if ("header".equalsIgnoreCase(e.tagName()))
return false;
}
return true;
}
}
private static class NodeData {
int signalNodeSize = 0;
int noiseNodeSize = 0;
int treeSize = 1;
int depth = 0;
public void NodeData(int depth) {}
private NodeData(int depth, int signalNodeSize, int noiseNodeSize) {
this.depth = depth;
this.signalNodeSize = signalNodeSize;
this.noiseNodeSize = noiseNodeSize;
}
public void add(NodeData other) {
signalNodeSize += other.signalNodeSize;
noiseNodeSize += other.noiseNodeSize;
treeSize += other.treeSize;
}
public void addAsNoise(NodeData other) {
noiseNodeSize += other.noiseNodeSize + other.signalNodeSize;
treeSize += other.treeSize;
}
public double signalRate() {
return signalNodeSize / (double)(signalNodeSize + noiseNodeSize);
}
}
It renders the text of this link (at present): https://news.ycombinator.com/item?id=32594821Into this search-engine friendly text:
The hard part is understanding which parts are the content versus navigation or promotions of other content. I’ve written a couple search engines. Have you tried making one with beautiful soup? Why does it matter? You love seafood, so just literally run grep on the entire page and if it contains the word then include it as a correct. In reality, you will miss a lot of real seafood pages because they don't really need to mention "seafood" and context matters, so what? Chances are that that one website where person randomly added "I love seafood" to the top of the page will be the only page that you've ever wanted to see anyway. There's too much data for you to go through in entire life in any case, so why worry about it as long as you can get something that's good enough? You will never get best data, if it was possible, google would be giving you best data already. How do I know? Well, looking up my real name shows where I grew up, what school I went to, graduated, and even which exam I scored 100 on... And even some places I used to work for in the past, and while that part is going to make most people paranoid, I wish ALL results were as detailed as this one, but there's little you can do. No I use JSoup for my search engine. You can calculate anchor tag density across the DOM tree and prune branches that exceed a certain threshold to remove navigational elements with reasonable accuracy if that is a problem. It's not going to be perfect, but even Google messes this up every once in a while. I wouldn't consider it a major hurdle. I don't presume the source is available... unbelievably cool project that I'm sure a lot of people have imagined themselves doing.
Dunno, not only are people sending me money to develop my search engine, not enough to live off but still, I also get emails and tweets from people who say they love it almost on a weekly basis.
I think attempting to be as comprehensive (or more) than Google is a trap. The better move is to fly under them. Be cheaper and better at something. Recipes is a great example of something Google is just miserable at, that is easy to do much better. There's plenty of such niches.
You love seafood, so just literally run grep on the entire page and if it contains the word then include it as a correct.
In reality, you will miss a lot of real seafood pages because they don't really need to mention "seafood" and context matters, so what? Chances are that that one website where person randomly added "I love seafood" to the top of the page will be the only page that you've ever wanted to see anyway.
There's too much data for you to go through in entire life in any case, so why worry about it as long as you can get something that's good enough? You will never get best data, if it was possible, google would be giving you best data already.
How do I know? Well, looking up my real name shows where I grew up, what school I went to, graduated, and even which exam I scored 100 on... And even some places I used to work for in the past, and while that part is going to make most people paranoid, I wish ALL results were as detailed as this one, but there's little you can do.
If it is a white list, then why have a search engine rather than an old school curated Yahoo directory?
They can do everything except the one thing that would actually hurt the search engine spammers right in the coin purse: Penalize websites for having ads.
- Sergey Brin and Lawrence Page, The Anatomy of a Large-Scale Hypertextual Web Search Engine
Oh ... this is such a good idea. I'm like tempted to try it and see what happens.
Then it was we can measure that and make money.
Now it’s just we can make money.
Whether that part is sinister or not, we know that we have a good number of bad actors, and from search engine results we can be sure that they have not developed a workable Byzantine fault tolerance mechanism to filter out the bad actors. Those who scream the loudest get put on a stage.
Whitelists that I wrote by hand also don't introduce new unexpected entries by the way :)
This instead could be more like RSS where as your crawler gets new sites, you get updates on new things, and you could filter in your crawler or in RSS client directly, doesn't matter.
How can we call everything a walled garden when many of them are free to get in and interconnect with each other?
As if you cannot look up address range of your own country then crawl your whole country for websites that may be hosted by people living locally.
As if you cannot do the same with a foreign country that interests you.
Maybe you could even find a list that only shows residential IP's so you're sure to be only finding webservers ran by individuals and not corporations.
And if somehow "port scanning" by trying to send a http request to a residential IP is illegal in your dystopian country, you can always start by scraping the site that you're interested in, there will always be at least one more link to another domain somewhere.
For large scale servers python is shit, but that doesn't mean that you cannot spend few weekends writing your own python crawler for your needs, which is so easy that you don't need to be a programmer to do it, and if you really care about this at all, a bit of a startup hurdle won't make you immediately disinterested.
And if it really does, there's always options like https://yacy.net/
You should see these things more like real life. If you wanted to know more about your own neighbourhood, what better way is there than to go outside and walk around your neighbourhood and see things with your own eyes?
Maybe that's just my opinion, but status quo is noone's but your own fault, because I never had this problem.
You could crawl forums and find deep technical discussions. Not anymore. And if a term was ever part of any news cycles, you get walls of Google selected propaganda.
Second, the quantity of intentionally fake noise has grown even faster - the spam problem that you have to solve is much harder than 30 years ago, any naive approach will simply fail to notice the needle in the haystack.
This question reveals a failure to understand the equipment, labor, and bandwidth costs of running a search engine.
It's completely unnecessary to make that estimate, a nonsense proposition since any two implementations are two orders of magnitude in cost apart, and a question that should never be asked of someone who hasn't done it.
Which is weird, because if you are who I think you are, you've done this in a trivial way, focusing on tiny sites.
And who knows? Maybe you're about to tell me that you've indexed several tens of thousands of pages yourself, that nobody's helping you, that it runs on two computers, and that it's Not That Difficult (tm).
Of course, then someone compares that engine to a practical search engine that also encompasses modern sites, and therefore needs to run tooled browsers to cope with their AJAX nonsense, and has to hit them every hour to be up to date.
And then you look at the disk cost.
Microsoft spends about $6 billion a year on Bing.
Duck Duck Go has more than 200 staff and raised $170+ million before their first profitable quarter
I think it's very easy for someone to put a homebrew HTML chess game on the phone store and then turn around and insist they know what it takes to run EA
It indexes not tens of thousands of pages, but has a peak capacity of about 100 million documents. I can crawl over a billion documents per month.
I don't really see anyone suggesting competing with Google or Bing off a PC in your garage, but it is absolutely and demonstrably feasible to build complementary services without any budget at all.
It doesn't require huge numbers of developers, it doesn't require a small country's allotment of bandwidth, and it doesn't require data-centers full of prohibitively expensive hardware.
This is much larger than expected.
This is basically the reason my team and I are building an alternative set of YouTube recommendations. You can check them out here:
I was just tired of YouTube steering me back to the same old small niche of videos, many times giving me repeat recommendations for stuff I'd already seen. Our algorithm is designed to surface smaller channels and find more obscure content.
(casual observation: Try matching titles without spaces, I did 'thisoldtony' and got nothing, but 'this old tony' matched. )
Thanks!
When I search for technical information 2 out of 3 times I get a website that I must pay to view content.
The internet is clearly going in a bad direction and most average joe users are suffering and will likely suffer more in the future.
Between datasheets and old cringey fanfic of mine, there are more and more resources that I am aware of that absolutely still exist on the internet, with reasonable robots.txt, but can't be coaxed out of google even with exact snippets.
It used to ensure most searches would have a few blog results, a wiki link, some large corps, some small corps, but that’s fallen apart.
I know this for the wrong reasons. I used to publish pages for my bank’s phone numbers because… I’d just publish their phone numbers.
While this is kinda a bad idea, now searches will give you 10 links to the bank’s own website and they make it difficult to find a number because they don’t want you to call them.
If the car was in an accident and the aircon doesn't work anymore, it means the gas loop is leaking. You can try to refill it but depending on the size of the leak it's going to work for a few hours to maybe a couple of days. You should evacuate the loop and do a vacuum test. If it is leaking, refilling the system with some added dye can show you where it is leaking. The Schrader valves are the usual suspects but as the car has been in an accident it could be anywhere. Adding refrigerant to a leaking system is just blowing away money that could be used to actually fix the aircon properly.
If you can’t find a sticker (or if that sticker says R-12, it still may have been converted), unscrew the cap on the service port and match it up to the type of port used by each refrigerant.
If it came off then I'd suggest calling a dealer parts department with your VIN and they should be able to get the information.
I usually get thousands and thousands of cloned websites that were likely set up in bulk using a template. They copy-paste just enough text to produce a search engine hit, while the real website it came from may not even be in the search results no matter how many pages of results I click through.
And then there are the elaborate clones of Github content, Stack Overflow, and various other technical help websites, all designed to make it look like all of those discussions are happening on the clone rather than the original. Some of them include a link back to the original, some don't. I get why some of those websites are ok with their content being openly reused (not that spammers care anyway), but in practice it destroys discoverability of their own service and wastes people's time.
Pinterest has spread through Google Images like a virus, they're plastered all over the results for searches that clearly aren't from boards made by real Pinterest users. I doubt it's a 3rd party spamming Pinterest because the only entity who actually benefits from it in practice is Pinterest itself. They've changed their onboarding pattern a lot over the years, but at one point it was virtually impossible to click through to the original website at all before the account creation popup blocked everything else.
Putting Pinterest at the top of image search results is effectively nothing more than a funnel to onboard more users for Pinterest, they rarely, if ever, have any relevance. I can't imagine why Google hasn't knocked them out of the results entirely at this point.
Whatever they're doing to combat actively hostile spam websites is either failing or they simply don't care anymore. The end result could not be more obvious.
The engineers were so preoccupied with whether or not they could, they didn't stop to think if they should.
If you want to break out of the dead, corporate internet, that's exactly what the GP built marginalia.nu to do.
> On the contrary… The search engines themselves…
With you 100% except for the opening rebuttal. What do you think /caused/ search engines to devolve like this if not digital marketing?
I pay cash for kagi.com, and recommend it.
Engineers should try their “lens” approach. I’d pay more for trusted curated lenses, and hope that’s in their model. The site above could offer a curated list of valid sites, and then I’d find them in the one engine too. (See Similar Projects on Marginalia’s About page.)
I also pay for Neeva, but they’re clearly trying to have their advertorial cake and eat it too. Still, it’s a better resource than Google when seeking an actual product.
I worry that solving digital marketing’s ‘unreasonable effectivness’ requires more than just ability to subscribe to content without ads, it should be possible to buy products without marketing budget built into the cost. Lower cost products would outcompete those spending money on ads, so all else being equal, enabling products to compete without marketing budget is the only solution I see. I don’t think a Neeva solves this by itself, though it’s likely a necessary component.
To me it falls into the category of IntelliJ products, where it makes my life and productivity so much better that the price is a no-brainer
Our (collective) disinclination towards paying for things on the Internet is what has led to the "everything must be monetized via ads" local maxima we're now stuck in.
If you care enough about this state of affairs, and can afford to do so (most people here can), then please consider paying for parts of the Internet that are important to you, like a search engine.
For better or worse, the direction of a paid product is usually fairly well defined, as long as they've taken time to understand their customers.
You pay for search one way or another, I'd rather be direct about it.
Sennheisers headphone division got eaten by its own success: the sennheiser hd 650 is so durable and has such a great soiund quality, that people just aren’t switching away from that 20 year old headphone.
In case the link you mentioned isn’t about that mid range headphone, it’s probably about the beyerdynamics T1.
Besides talking about audio equipment: I hate the sites on google who update the dates of their articles although the content wasn’t changed. Happens way to often. I googled the release date of BOTW2 a few days ago, and 3/4 of the search results were blatant seo spam where the initial article was about something else, and then the headline and date were changed in order to get more traffic from google.
Imo you are setting the wrong priorities for a search engine.
But an outdated article doesn't guarantee this to be true. For example, maybe the manufacturer released an updated version that is a better value proposition and kept the old model around to have a more budget friendly option. Or perhaps another company purchased the manufacture and demand they cut costs. Or maybe this model is new and people haven't yet learned that there is a specific part that frequently fails after a few years of use. An old article can't speak to these hypotheticals. It doesn't mean the article is wrong. It just means that the article is less informed than if it were written today giving the exact same recommendation.
>the sennheiser hd 650 is so durable and has such a great soiund quality, that people just aren’t switching away from that 20 year old headphone.
But this is only something that can be truly known after those 20 years.
>I hate the sites on google who update the dates of their articles although the content wasn’t changed.
I agree, and while this is a related issue, it isn't really the same problem. It is a failure in Google's anti-SEO features. They don't need to trust the date on the article. They could compare cached versions of the page to see what changed besides the date.
Nor does a new article on a "review" site monetized with affiliate links guarantee it. So who do you trust more? Older but honest review from an expert or a new affilate driven review? Kagi choses the former as likelier to be more valuable to the user in this case.
If I saw a photograph of all the cellphones from 2004, I can pick out the best cellphone, this doesn’t mean that the best cellphone from 2004 is still the best cellphone.
It’s implied that “best” usually means what is “best” for what people need today as those needs evolve drastically over time and especially with tech products.
As a data scientist, just being able to block Towards Data Science and other garbage DS content churned out by amateurs to get their resumes boosted is well, well worth it. It's ridiculous how much top ranking content on Google is flat out technically incorrect, or at least clearly misunderstanding the subject.
"... it costs us about $1 to process 80 searches. ... An average Kagi beta user is actually searching about 30 times a day. At USD $10/month, the price does not even cover our cost for average use, and we are basically betting that average use will go down a bit with time because during beta people may be searching more than normal due to testing etc. Our goal is to find the minimum price at which we can sustain the business. If it turns out that we have more room we will decrease it. But it can also be that we may need to increase it."
I went and looked up my Google search history for yesterday - it's 40 searches. I'd expect it to be above average, but still... if it's $10 per 80 queries, it feels like $10 is likely to be too low to be sustainable. And while I personally don't mind paying more, I wonder how many people will - and what it'll mean for the service long term, if they just can't attract enough people to make it worthwhile.
I wish dumping the top million also dumped anything with “Top N” in the title of the page…
https://you.com/search?q=best+laptops https://www.google.com/search?q=best%20laptops
[1] For instance, this link - which I discovered just ten minutes ago. I know for a fact that I have never submitted poetry to the Porkopolis website (motto: "Considering the pig, a single-minded bestiary") but it's always a pleasure to discover other people putting my words to good use! - http://www.porkopolis.org/pig_poet/rik-roots/
Getting outside one's comfort zone and putting in the time to find something good/interesting/new is highly underrated. But it is work. And many a corporate empire has been built by making a mediocre or sufficient experience the most convenient thing.
Urban Spoon was an amazing resource for us road warrior types. I found many fantastic places > 1/4 mile off the interstate. Nowadays, I ask employees at worksites for their opinions. If they recommend a box chain, I ask someone else.
While I like using Marginalia to find these websites, I don’t think it’s a demonstration of how “alive” the internet is but more like a lens into what the internet used to be, like walking around an archaeological dig site.
Access to the real, genuine Internet people and places will be made invisible; protected by the gargantuan SEO-fed lipid-berg of AI-generated content, keeping the social media peasantry ever-corralled in the cattle pen where they shall be kept happy and fed by their keepers.
Is September 2022 the final Eternal September?
This is a critical distinction, because the former is a problem like "the water is too wet", like you can't really fix that. You can build new digital infrastructure though. That's a solvable problem.
It is extremely effective in the first way, but extremely ineffective in the latter.
Influencers are some of the most popular people on the planet for young people.
Maybe if you know the exact url or specific keywords, but generally not now. Google has turned into ad placement the same level ask jeeves and their ilk were. It's atrocious for surfacing anything other than click bait. Duckduckgo is better, but not by much imo.
Isn't that the point? The general populace isn't getting off of social media and Google, thus, dead internet theory continues to compound itself..
*The weird internet's death has been highly exaggerated.
Have you tried searching lately? It feels like it is becoming increasingly difficult to find actual articles with useful information in a sea of SEO trash.
You take the red pill, you stay in Wonderland, and I show you how deep the rabbit-hole goes.
A colleague showed me a website the other day from 2013 that was an absolute jewel in terms of knowledge. I am sure more recent sites like that exist, but I am afraid finding those with google are almost zero.
[1] http://ilpubs.stanford.edu:8090/422/1/1999-66.pdf (ch. 6)
Right now it's absolutely amazing if you have like a broad topic you want websites about, but kind of weak when you want something more specific.
While yeah, marginalia finds interesting stuff, I've not been able to find anything useful that I've tried searching for with it so far.
It's not about it being not existent. It's about it being too small a percentage. And will algorithmic generation and rampant re-posting of news content 1000s of times on different outlets, this is probably true...
8 billion people are able to manually create much fewer content than thousands upon thousands of automated generation scripts and bots...