Techdirt has been deleted from Bing and DuckDuckGo [fixed]
techdirt.com
techdirt.com
Update: Still investigating, but have made some progress -- Determined that on desktop there was a link to Techdirt up continuously via our About module (when you search for "Techdirt"). And now the traditional web link is back up as well (for desktop and mobile): https://duckduckgo.com/?q=techdirt&ia=web.
Over the past fifteen years, search has become "universal" with dozens of indexes and modules throughout the page. Also it is important to understand people click/engage with things on the page roughly half as much each position down, so by the time you get to the bottom it is like 100 times less than the top. This means that things on top, increasingly non-web links, have become more and more important.
In this context, as mentioned in the other comment, the largest modules on mobile and desktop we power ourselves, that is, local and knowledge graph. AI will be the same. We do use Bing as the primary source for traditional web links, but not for all and even when we do it sometimes looks different in various ways. In something like this, we can re-insert this link if that is what is needed.
One potentially relevant line from the article:
>I love that first one. Microsoft, a company with a $2.5 trillion market cap, “may not have enough resources” to crawl and index Techdirt? Cool.
I have to imagine DDG's valuation wouldn't quite hit 2.5 trillion if it went public. It's not absurd that they aren't able to fully duplicate those efforts (or that they don't need to in order to have a decent product).
They fixed the problem so it shows ability to. Maybe the issue is “detect what parts we should do ourselves”? How do we detect censorship beyond blog posts like this? Can we diff various indices on a regular basis?
The whole point of using Bing as the base is so that you don't have to make and keep an index of everything yourself, so DDG wouldn't have another exhaustive index to compare to (and if they would, they wouldn't need Bing). They index some things to add onto Bing, but they don't try to capture everything.
Furthermore, any specific site missing from the index does not necessarily mean a flaw - okay, you would diff some indices, and find out a long list of what Bing has excluded... and then what? For every 'fixable censorship' case (accidental or intentional) there will be thousands of spam or malicious sites which should be excluded, as nowadays the key part of search is not finding everything but throwing out the results which want to be found but shouldn't. Again, DDG wants to piggyback on the effort that Bing is doing to filter the index, and if they want to second-guess all Bing's filtering (as opposed to just making a fix for this specific case) then they have to replicate and improve on all of the (huge and expensive!) Bing's filtering effort; which goes counter to the reason for using Bing which is to avoid all this expense.
> Bing is our largest source of traditional web links.
I think "traditional web links" is the main (and I think should be the only) product of a search engine, and it seems like you rely on Bing quite significantly to serve search results.
- maps
- directions
- local business listings
- about boxes
- basic facts
- conversions
- news
- videos
- images
- shopping
- sports
- stocks
- weather
- flights
- entertainment (e.g. cast)
- etc.
> We do use Bing as the primary source for traditional web links, but not for all and even when we do it sometimes looks different in various ways
I don't know that it's reasonable to demand that they provide you an exact percentage.
Personally I'd like to know if that means "we have a manual list we re-added". I think what myself and probably many people are most interested in is "do you rely on bing for < 95% of word match searches (no knowledge graph)"
> when we do it sometimes looks different in various ways
IMO this is also vague - I assume this means they're knowledge graph augmented? I really want to know if they have their own indexer (vs a manual list + caching bing results) for text matching.
What we really want to know is if for links alone - what extent do you index? Let's ignore the knowledge graph, let's say we just want a simple word match.
I'm surprised you didn't have techdirt's homepage indexed at all. "techdirt" did not return "techdirt.com". To me this means you aren't indexing the home pages for popular websites at all - let alone their content - and rely effectively exclusively on bing with maybe some "caching".
Talking about links alone - exact word matches with the domain name (or page content other than your knowledge graph), what do you index?
What specifically do you do other than adding sites manually to a list and I assume "caching bings results"? To me this would still be 99.999999% ~= 100% reliant on bing for search results.
This isn't meant to be an attack in anyway, I'm just looking to clarify what's going on.
This is context-dependent. If you're looking for a link, then you're looking for a link; but more often than not, you're actually looking for answers.
Say you're looking up an exchange rate or the value of a stock; converting units; checking the weather; or just want to look up an actor's photo. Why would you click and wait for a (likely, slow & ad-infested) page to load, if the search engine can provide factually correct data without any extra steps? How often do you read past the first paragraph on Wikipedia?
If you want to dive in / verify any of that information, it's still all just a click away. But having to only do one thing instead of two is almost the definition of technological progress.
I want room for both.
There are times I want the search engine to stop "helping" (sic) and just give me a straight search based on my terms.
There are other times where I'm looking for something specific but I don't remember exactly what it is, so I want help correcting/ narrowing down my search.
Bing is their initial and primary source of data for web links. Not the only one either, but the primary one. They maintain their own index from it.
That’s not “a mirror” by any stretch of the imagination.
What appears to have happened here is some form of replication bug where in a way yet to be determined a removal from the original data source was unintentionally replicated.
Data synchronization is hard when you’re not directly mirroring data. This kind of bug crops up
Not really. That’s the point. Data replication is hard, especially when you have to consistently update from the same data source.
Bugs happen.
What is their other data source? And how did it come to suffer from the same problem?
I keep seeing insinuations that DDG has some other secret source than Bing, but it's always hand waved and never explicitly named. Incidents like this seem to strongly imply that only Bing really matters at the end of the day.
It’s not an insinuation, it’s a straight fact direct from the founder.
That you don’t understand data replication problems—or know the inner workings of a private company’s codebase & data funnel—doesn’t make it some conspiracy.
It just means you’re uninformed.
I'm glad to be corrected. What is their other data source?
Personally, I don't even mind if using DDG is the same thing as searching bing so long as it actually works, but they still can't get a simple search like "office -microsoft" or "headphones -best" right. I even thought they acknowledged the problem and were looking into it at one point.
All the stuff they claim to do themselves is <1% of whatever value DDG provides. Lucky for them it’s just enough to be able to use terms like “largely” instead of “entirely”, “other sources” instead of “only source”, etc.
They would fast become outdated however, but that’s a different problem. One solvable with another source.
If some tech kerfuffle is happening, there's a good chance someone involved will see it. Its like a wider ranged subreddit, or a more orderly Twitter.
https://news.ycombinator.com/item?id=36899187
https://news.ycombinator.com/item?id=36899072
Edit: results appear on DDG, for me, apparently.
“This shouldn’t happen, we are working out why it did” is probably the only official response possible right now.
Elsewhere in the comments he’s addressing misconceptions & outright falsehoods about how DDG works.
(We definitely don't want fake/bait stuff here either so I share your feeling on that - this just isn't a good way to express it, at least not on HN.)
<https://news.ycombinator.com/item?id=35669888>
You might take his suggestions to heart.
I don't really like to be cheeky/sarcastic, but I don't really like people telling me how to behave, I do me, you do you, especially if you are an admin
It’s either people are compatible with a community or they are not, asking to change means people have no personality and you can just shape their behaviour at will, i guess that’s true for tech people, but not the rest of the world
We're just trying to have an internet forum that doesn't suck—it's nothing personal! The problem we face is that large internet forums end up sucking by default - that's where the arrow of entropy points - so if we want to avoid that outcome, we need to spend a lot of energy to deflect it, and we need users to help with that.
There's nothing intrinsically wrong about posting unsubstantive/flamebait comments on the internet - there are other sites where that is common and expected. We're simply trying to play a different game here, and for that we have different rules*. Just like baseball rules say you can't tackle the pitcher, and football rules don't let you whack the ball with a stick, HN rules say not to be snarky, not to post flamewar comments, and so on (https://news.ycombinator.com/newsguidelines.html).
Btw, what you said about "changing others" is correct in a way - I post like this as a moderator because I'm trying to persuade users who are breaking the rules to change their behavior. (Not make them change - just persuade them that it's in their interest.) But you're wrong that this isn't doable—it does work sometimes. You decide whether or not it works in your case, of course, but HN has a long history of users switching to using the site as intended, once they understood why it was in their own interest to do so.
Here's why it's in your interest: HN is only worth visiting in the first place because we have these rules and most users follow them. Flamewar, for example, can be exciting for a while, but in the end it's just repetitive and boring. The idea of HN is to satisfy curiosity, and for that we have to avoid the snarky/flippant/sensational/indignant/etc. types of internet posting, all of which destroy the site for that purpose. It's in your interest to participate in the intended way, because then you're contributing to the reason why it's worth visiting at all. I'm sure you wouldn't leave fires burning in a campsite that you enjoy, or (less dramatically) drop trash in a pleasant park. The way we see it, asking people to play the intended game here and stick to its rules is sort of like that.
* this is also why I often ask people "Please don't do this here". It's not "please don't do this" or "it's wrong to do this" or "you're bad for doing this" - it's just "please don't do this here."
I'd reported a problem with the "lite" interface about two months ago. It was not only acknowledged, but fixed, in just over half an hour:
Long story short - at some people in the past someone proxy-mirrored all content of SaaSHub. You'd open a page like "someshady-proxy-mirror.com/duckduckgo-alternatives", and it will mirror saashub.com/duckduckgo-alternatives. The same happened for all all pages. Soon after that, Bing (and DDG respectively) dropped all SaaSHub links from the index and kept the proxy-mirrored content! WTF.
After a week-or-two of "fighting" I managed to block the proxy-mirrors; however, I never got SaaSHub back in Bing's index.
I'm a single founder and feel helpless with this Bing/DGG issue. Any help would be appreciated. Thanks.
At this moment "techdirt" only returns the wikipedia article, twitter, then mostly unrelated mentions of it. Surely DDG would have seen at least their homepage before?
> Most of our search result pages feature one or more Instant Answers. To deliver Instant Answers on specific topics, DuckDuckGo leverages many sources, including specialized sources like Sportradar and crowd-sourced sites like Wikipedia. We also maintain our own crawler (DuckDuckBot) and many indexes to support our results. Of course, we have more traditional links and images in our search results too, which we largely source from Bing. Our focus is synthesizing all these sources to create a superior search experience.
https://duckduckgo.com/duckduckgo-help-pages/results/sources...
They eventually got it fixed but only after a few other people finally noticed and the owner contacted someone and then waited a week or so. But it was like that for years. You search a perfectly good search term that should pull up one of the articles on that site first, and all you got are all kinds of other 2nd and 3rd hand references like email archive posts and articles on other sites, but scroll down as many pages as you want and the actual site would never come up on ddg.
That's when I got interested in Kagi.
As to their businesses proposition privacy was central from day 1: In 2011, Weinberg, then the company’s sole employee, took out an ad on a billboard in San Francisco that declared, “Google tracks you. We don’t.” That branding—Google, but private—has served the company well in the years since.
“The only way to compete with Google is not to try to compete on search results,” says Brad Burnham, a partner at Union Square Ventures, which gave DuckDuckGo its first and only Series A funding in 2011. https://www.wired.com/story/duckduckgo-quest-prove-online-pr...
Edit: Since you edited, no one is arguing that their whole spiel isn’t about privacy. What I’m saying is that they refuse to admit that they’re a privacy wrapper around Bing which would make their value proposition quite weak. They puff their chest and pretend to be a big standalone player in the market but they aren’t. You aren’t at the whims of DDG as a user, but at those of Bing/Microsoft. If you’re cool with that, carry on, but I’m not going to let yegg continue to make vague statements uncontested. If they’re so sketchy and dishonest about their role in deciding the links, what else are they dishonest about? Some of you let a lot slide because they’re the underdog.
PS: As to your edit, they have specifically mentioned using Bing in the past alongside other engines. They are clearly running their own search process and widgets while also using a lot of data from Bing. I can see why you might consider it deceptive, though I really don’t.
If they're manually adding a entry for "insert site here", maybe some "caching" of bing's results for the biggest most common searches then IMO that's still 99.999999% ~= 100%
Now, they might be using low level API access to Bing’s index or something else. But I don’t see how it can be a simple wrapper.
A couple of weeks ago, I was debating with someone about what "LMR" stood for in the context of cable specifications, such as LMR-240, LMR-400 and so on. I thought it meant "Land Mobile Radio" while the other person disagreed that it stood for anything. A Google search on LMR coax cable acronym returned a helpful info blurb stating that LMR stood for "Last Minute Resistance" as a means of fending off sexual assault.
Needless to say there was no way to tell exactly what site Google had copied that definition from, and no useful way to provide feedback to them. Sometimes there's a "Feedback" link, this time there wasn't. Sometimes the feedback link is present but only offers the option of reporting illegal activity. That option wasn't present either.
For whatever reason, Google clearly does not give a flying fuck at a rolling donut about search quality anymore. With the right leadership, Bing could own that entire line of business, in a manner reminiscent of IE's original dominance over Netscape. I'm not holding my breath, but at this point I'm cheering for anyone who can offer Google some competition.
I rarely use search engines anymore. I'll bet the same is true for many people.
In fact when I used regular Google, the results are good enough for me to deduce that on my own.
So I consider GP's search skills inadequate. I mean it's not exactly wrong to desire a tool that handholds you and feeds you the answer; but if you are willing to do a little bit of deduction Google is fine.
I would assume the best actual source is the USPTO trademark registration database, which has this entry: https://tmsearch.uspto.gov/bin/showfield?f=doc&state=4805:pw...
It appears to indeed mean nothing, but of course, it took me a whole two minutes or so to find an actually authoritative source. Everyone knows no real user will ever do that and just wants a search engine or chatbot to dictate reality to them.
People are building workarounds in real life due to how bad Google's results have gotten.
I tried the example you cited. The "blurb" is called a snippet; the snippet comes from the web page itself. One of the pages had that actual text, which is why it appeared. Why it had it on a page that's primarily about coaxial cable isn't clear, but we'll look into how to improve.
As for sending feedback, each link in the results has a little three dot icon next to it that brings up our "About This Results" panel, and you can send feedback that way.
Also, to belatedly introduce myself, I'm the public liaison for search at Google. It's a position we have within the actual search engineering team to help us gather feedback to improve search quality. Feel free for you or anyone comfortable sharing examples of unhelpful results to flag me about it:
https://twitter.com/searchliaison https://mastodon.social/@searchliaison
Surely you don't expect us to believe this ever gets read, much less acted upon? Don't insult the intelligence of your audience.
I've been assuming that this was just a bug with DDG and not a deliberate attempt to cripple their search engine. I've seen it with Google results too, but I could buy that Google would break them, since making websites harder to find encourages people to pay for prominently placed ads, but what would DDG's motivation be?
If you type 'word1 word2 word3', where word3 is less common than word1 and word2, a lot of the time, it will act as if word3 simply wasn't in the query.
And when DDG runs out of "web links" it just fills the rest of result pages with local results that are out-of-place and useless. Like, getting "Visit Paris" sites after searching for rare computer parts.
It seems there's stuff going on behind the scenes. DDG got taken over by some vested interest, perhaps. Or Bing doing stuff and DDG never branching out of it and just blindly getting results from it.
Either way, it's starting to rival Google in uselessness and I'm likely to stop using it fairly soon.
Kagi seems to work pretty well.
The thing with search engines is: you try one, scan results for 30 seconds, don't find a useful result, curse under your breath and hastily try another search engine. It's a normal programmer flow, which sadly almost completely eliminates the possibility to provide actionable feedback to the search engine's maintainers.
I gather you're from the DDG team? If so, I'd say that the writing is on the wall that Bing is no longer a good backend for DDG. You guys should start branching out because from where I'm standing, many programmers are 10-15 annoyances away from switching.
Again, my apologies for not providing examples.
But I have found the LLM powered Bing, using Skype (!!) as the user interface really compelling
I have played around myself with getting LLMs to summerise and evaluate text for me, and it seemed that an automated way would be very cool.
Then Microsoft implemented it.
Wow. It is very very good
So DDG is still good. Very good. And privacy, a very strong selling point
But wading through click bait websites looking for a website that is not just straining for eyeballs: Bing's LLM based machine is a dogsend!
I would pay for a service like that from DDG. I already pay real money for an OpenAI API key.
Sell me yours. I would much rather pay my money to you
number of Americans without a credit card
panic nova git automation
microvision (the video game console)
time in chicago (seems to be fixed now)
mozilla open directory (seems to be fixed now)
Now that you've solicited this feedback, I'll make a point of saving futile searches for the next time you pop up on HN.I've too noticed result quality in the main search engines I use (DDG, Brave, Phind) go way down in recent times, I wonder what happened...
This also happens to me and I have found reports on Reddit of it happening to many others.
Sometimes, I cannot even find the original result that I had clicked on, when I return to the search results page.
This drives me crazy.
free for 100 searches. $5/month for 300 searches $10/month for 1,000 searches
I'm not at the point where I feel DDG and Bing are so bad I want to start paying to get better search results. I'd be interested to see how many people are there though.
But I too started getting disgusted by "everything is a subscription" but I might jump the train if I like Kagi enough.
Because apparently the internet companies can't figure out an ad model that's not extremely toxic and does not trample on every single privacy rule the world has (and the 100x more that the world still doesn't have but should not be broken anyway because they should be a moral / ethical no-brainer but alas, go tell that to the "money above everything" types).
But finally, after decades, all the VCs funding internet companies that planned to capture the market with network effects and then start charging, are showing their true colors. I am glad. It makes them more honest and gives the users better information to act on.
Us the nerds just practiced endless bikeshedding and the corporations took over everything in the meantime. Oh, let's not forget the people who kept inventing LISP dialects BECAUSE THAT'S EXACTLY WHAT THE WORLD NEEDED!
1. https://news.ycombinator.com/item?id=32360874
- We have on the order of a million lines of search code at this point and have a lot of talented people working them. That code does a myriad of things across many indexes.
- As an example, mobile searches are the largest category of searches, and local searches are the largest category of searches within mobile. We don't get any local search module content from Bing.
- Similarly, on desktop, knowledge graph / Wikipedia-type answers come up the most and we don't get any of that module content from Bing either.
- Bing is our largest source of traditional web links, which have become less and less relevant/engaged with over time as more and more modules are in search results and put on top of traditional links (and people interact with things on top of the search results page about two orders of magnitude vs. things on the bottom).
- When Bing has dropped things out of the traditional web index, we have put them back, and we've been working with them so this happens less and less. In fact, there hasn't been hardly any reports of this in the past month or so, which is why I've asked for other examples in the comments.
Once this particular issue is sorted out I'd would also be cool to see some sort of post-mortem report. Seeing why this sort of thing happens and a standard process for fixing it would be beneficial to all involved I think. Complex systems are complex and shit happens. But if this happened to a site far smaller than TechDirt (they're not even that big AFAIK) I don't think there would be any avenue to cure the situation for them.
/2¢
I'd be interested to know because it gets to the heart of the matter and why confusion seems prevalent. If we ignore the modules and local search, what differentiates DDG and Bing when returning bread and butter traditional web links?
> At DuckDuckGo, we've been rolling out search updates that down-rank sites associated with Russian disinformation.
- Gabriel Weinberg via Twitter, March 10, 2022. Archived: https://archive.is/SLGYb
I'll note that there is a April 17th update tweet referenced in that thread. That is unrelated and it was about a rumor that DGG was purging certain media sites. Nothing to do about censoring Russian sites. Archive of that: https://archive.ph/I2iUp
I realized I previously explained how our news rankings work very poorly on Twitter that got grossly misinterpreted, so I subsequently put out a clarification in this help page with a much clearer (and detailed) explanation of how our news rankings actually work: https://duckduckgo.com/duckduckgo-help-pages/results/news-ra...
From that page: "When we apply our own ranking signals we do so in a strictly non-political manner, meaning we don’t evaluate or otherwise take into account any potential political bias or leanings of websites in our search result rankings." That is, we did not/do not have a disinformation/"truth" detector, nor did we go looking for any Russian narrative (or any other narrative for that matter). Instead, we just have essentially a spam detector that had detected some spam from Russian state sites. That's it.
For example, if the top three results for a news story come from CNN, MSNBC, Fox News, the next search would display MSNBC, Fox News, CNN.
Anyways, 4+ year user of DDG and still love it!
https://daverupert.com/2023/01/shadow-banned-by-duckduckgo-a...
https://www.jessesquires.com/blog/2022/03/25/my-website-disa...
Guess I need to start playing more with things like Algolia or Kagi.
Techdirt has been an upstanding place for journalism forever, and getting censored this way is ridiculous.
They build their own index: https://help.kagi.com/kagi/why-kagi/kagi-vs-competition.html
I’d love it if instead of ‘magic’ we just had a search engine that let you be the magician with a better query language and ux filters.
I am totally not affiliated with kagi but since I started using it, I haven't gotten pissed at my search engine for sucking ass.
OH! Another feature that actually works on Kagi: Date filtering! Never again will you set that date filter to the last two weeks and receive results that were posted in 2005.
This is completely unfounded. Please provide any evidence that this is intentional or related to censorship.
The only problem is the pricing model. I don't do overage fees because they make me anxious. I like knowing there is a concrete ceiling on what something will cost me. 300 a month also feels far too low and $10/mo, whether fair or not, feels far too high for a search engine.
This is why I've been working on a behavioural change: stop using search engines and start going right to the websites I know and trust.
Meanwhile the DDG founder calmly stating it’s being looked into, is definitely a bug, and explains that this should not generally happen… and being ignored.
It's not some random unknown blog with zero traffic.
https://news.ycombinator.com/item?id=36898807
Meanwhile, thanks for illustrating my comment extraordinarily effectively, I guess?
He claims when sites get dropped from bing they are automatically dropped from ddg and require manual efforts to get them up.
There is no process either. You need to connect with someone on the inside. No form, phone number or any reasonable way of contacting them. If you don't share a daycare timeslot with him or can get your submission upvoted here good luck
You’re arguing in bad faith, apparently because you came to a conclusion first & are scrambling to justify it.
Why are you shilling and telling me to look at the comments instead of linking directly to the comment that would address my point?
Just because delistings can happen because of Bing doesn’t mean that’s what happened here.
Founder says they get info from many sources. Talks about local and factual information, flights, images, videos, etc as important categories while calling normal websites part of legacy web. Fails to mention they get this content from only bing. Tries to explain people are searching for less legacy while people believe they mostly use a search engine for legacy web.
I can't see a bug explanation being anywhere near the truth when every site delisted from bing is automatically delisted in duckduckgo and requires from some action to save it.
How do you explain the other sites?
Because god know a you’re not the only one that’s managed to turn “I don’t know why” into “I’m sure this is what is happening”…
So it's a black box, and you kind of just hope for the best.
[1]: https://reclaimthenet.org/microsofts-bing-censors-tank-man
[2]: https://www.washingtonexaminer.com/news/duckduckgo-slammed-f...
There's a reason why mainstream search engines like DuckDuckGo perform this kind of blacklisting. Because 99.99% of the time this blacklisting only renders invisible what is worthless anyways. But that 0.01% of the time when it incorrectly blacklists a site HN is outraged.
I can't help but think I've encountered this type of content before...but where.
I also wouldn't exactly call this kind of response "outrage" but something close to befuddled amusement. "Whoops! Looks like we accidentally blocked those sites that criticize us! Yikes." If in err, then what an error! If by intent, then who would believe otherwise?
I suppose some cranky, overly emotional weirdos might get in a tizzy, but this is pretty bog standard stuff. It's just fun to gawk at the sweaty fat man on the tricycle trying to backpedal across the motorway.
That’s a feature, not a bug, from my point of view.
Kagi is definately worth the $10 fee. First, it allows me to block websites that I find irrelevant or annoying in my areas of interest, such as Pinterest or Quora. It also promises an environment free of advertisements and tracking. I particularly appreciate the excellent customer support, where actual human beings respond to your emails, not automated responses. Additionally, it offers 'custom lenses', which are essentially search templates. For instance, using this feature I can refine my search to target only educational sites, filter for PDFs, limit the time to the "last 48 hours", and search for specific subjects, like machine learning, along with my query.
My only gripe with Kagi is its stringent limit on searches, set at around 10,000 per month. I've bumped against this ceiling a few times. Despite this limitation, the wealth of features and the quality of the search results make Kagi a worthwhile investment for me.
There is no defensible reason to disable it on a content heavy site like yours.
Google has 100k+ results for the same prompt.
I find it even harder to believe that when he confirmed the problem, he then didn't do anything about it for months.
This feels incredibly performative.
Hey Mike: instead of bing chat, try https://www.bing.com/webmaster/tools
I still believe DDG has a place, in being a more privacy-focused aggregator to Bing and a few other sources, but this did shatter the previous image of DDG having a large degree of autonomy for me.