“YouTube-dl” and “Pirate Bay” back on DDG
fosstodon.org
fosstodon.org
https://github.com/ytdl-org/youtube-dl/graphs/commit-activit...
By itself, using the ANDROID API instead of the WEB API^2 does nothing to affect download speeds. I can block yt-dlp's POST request indicating the client and API name and this has no effect on download speed.
A website that forces users to run Javascript in order to get faster download speeds. This is not a new idea.
1. https://www.youtube.com/s/player/{player_version}/player_ias...
2. From the YouTube video page JSON:
ANDROID API KEY
AIzaSyA8eiZmM1FaDVjRy-df2KTyQ_vz_yYM39w
WEB API KEY
AIzaSyAO_FJ2SlqU8Q4STEHLGCilw_Y9_11qcW8
One of the reasons yt-dl and yt-dlp are so slow is because they do (too) many do other things besides running the Javascript, modifying n and sig, and spitting out a fast download URL. Before they can run the JS, the video page needs to be retrieved, but waiting for Python to start up and do this is slow. A YouTube video page can be retrieved much faster using netcat, outside of Python. YouTube video can be downloaded much faster using an HTTP client like tnftp, directly, outside of Python. YouTube video can be converted much faster using ffmpeg, directly, outside of Python. And so on. These programs start instantaniously when compared to the slow start up time of Python. The Python startup latency is unbearable.
At the very least yt-dl and yt-dlp should accept an already downloaded YouTube video page as input instead of forcing the user to use Python (or Python calling another program) to download the page. The recalculation of "n" and "sig" do not have to occur within seconds of retrieving the video page. There is no need to require the user to use Python for downloading webpages. The values in the YouTube video page are good for a substantial amount of time.
Both scripts have an option to output the new n and sig values in a download URL or JSON containing the download URLs, so the user can use an HTTP client directly, outside of Python. But the user still has to use Python to download each video page. Using an HTTP client directly would be faster.
Using yt-dl and yt-dlp just to output optimal download URLs feels like overkill.
Absolutely nobody thinks optimizing the meta/API processing of yt-dlp & co is worth it. This is exactly why we have high level programming languages that make all of this much easier, instead of trying to write HTML and JS parsing in plain C. Keep in mind these tools support dozens or hundreds or websites, not just YouTube.
If you think rewriting yt-dlp in C is worth it, go right ahead, but you're not going to make it significantly faster; you're just going to make maintenance a much bigger pain for yourself. Pick the right tool for the job. Python is absolutely (one of) the right tools for this job.
(For the record: I use C and Python on a daily basis, and will use whatever language is appropriate for each situation.)
This opinion assumes it is a single job. I see multiple jobs. The number "options" provided by yt-dl(p) gives us a clue.
There is nothing wrong with preferring to use larger, more complicated, "multi-purpose" utilities. There will always be plenty to choose from.
However the idea of using smaller, less complicated, single purpose utilities is not "nonsense". It makes sense in many cases and some users may prefer it.
The statements I make about speed are from day-to-day experience not conjecture.
Plus I can request multiple video pages over a single TCP connection with netcat.
For example, in a single TCP connection, with a list of 30 videoIds, I can get initial sig and n values for 413 videoIds. No wait time for netcat to decompress or startup. Using netcat is quite fast. Then I have utilties written in C to extract URLs from stdin. As such, all I need is a utility to update the sig and n values in the download URLs to make them fast ones instead of throttled.
How long would it take to get sig and n values for 30 videoIds let alone 413 with yt-dl or yt-dlp over a single TCP connection. The startup time for yt-dl/yt-dlp plus the time waiting for the downloading of each video page makes it much, much slower. These are "do-everything" scripts that are as a result quite inflexible.
yt-dl and yt-dlp use Python to run the Javascript functions that modify the initial "n" and "sig" values to make non-throttled download URLs, but a faster language could be used to interpret and run the snippet of Javascript. For example, V8 is written in C++, not Python.
One of the reasons yt-dl and yt-dlp are so slow, IMO, is because they do (too) many other things besides running the Javascript, modifying n and sig, and spitting out a fast download URL. Before they can run the JS, the video page needs to be retrieved. I have self-created utilities for downloading webpages wth TCP clients that are sgnificantly faster and more flexible than yt-dl/yt-dlp. I use these small dedicated programs every day. I am used to the speed. Waiting for yt-dl/yt-dlp to start up in order to download web pages is annoying. The latency is unbearable.
At the very least yt-dl and yt-dlp should accept an already downloaded YouTube video page as input instead of forcing the user to use Python (or Python calling another program) to download the page. The recalculation of "n" and "sig" do not have to occur within seconds of retrieving the video page. There is no need to require the user to use Python for downloading webpages. The values in the YouTube video page are good for a substantial amount of time.
Both yt-dl and yt-dlp have an option to output the new n and sig values in a download URL or as JSON containing the download URLs, so the user can use an HTTP client directly to download video, wthout need to launch yt-dl/yt-dlp. But the user still has to use yt-dl/yt-dlp to download each video page. Using a TCP client directly would be faster.
Using yt-dl and yt-dlp just to output optimal download URLs feels like overkill. It would be nice to have a program that just focuses on running the inecessary JS in base.js in order to output optimal download URLs. Then the user can use whatever programs she wants, directly, for downloading HTML/JSON, extracting URLs, downloading video files, converting video files, etc. Downloading video from most websites is generally easy. I never need a program like yt-dl/yt-dlp. It is only YouTube that plays games with users trying to get them to enable Javascript and be tracked. One need only look at the size of the extractors in yt-dl/yt-dlp as evidence. The extractor for YouTube is 3x the size of the next largest, and over 10-20x the size of most of the others. I want a small utility that just focuses on YouTube. A simpler solution instead of a massive project.
--external-downloader aria2c
--external-downloader-args "--continue --max-concurrent-downloads=3 --max-connection-per-server=3 --split 3 --min-split-size 1M"
(possibly in your config file)Also, neither --format best nor --format bestvideo chooses the best encoding in all cases; they use bitrate as a heuristic for quality, and a less efficient codec can have higher bitrate but worse quality, resolution, or framerate. The workaround for this is specifying --format with an enumeration of every combination of codec, resolution, and framerate in preferred order, which goes like this:
--format "(bestvideo[vcodec^=av01][height>=4320][fps>30]/bestvideo[vcodec^=vp9.2][height>=4320] ...
Here's a full example (hmm... they're using it with yt-dlp, which I thought had fixed this?):https://github.com/TheFrenchGhosty/TheFrenchGhostys-Ultimate...
I think there's a bit of variation in the exact order among the config files found online. If you're goals are archival, consider also retrieving metadata, thumbnail, and subtitles in all languages; I also have in my config the options:
--verbose
--download-archive ./ytdl-archive.txt
--cookies ./ytdl-cookies.txt
--merge-output-format mkv
--add-metadata --all-subs --embed-subs
--write-info-json --write-thumbnail
--no-overwrites --continue
--force-ipv4
(the only remaining workaround for age-restricted videos is to give it cookies extracted from a browser with a real Google account logged in)Very misleading phrasing: you would download in 10 hours a ~10 hours long video, which (of course) could have been downloaded in a fraction of the time.
The throttling has the user download at a speed similar to that required to viewing the video.
Typical video bitrates for bog standard 1080p30 will be in the 1MB/s range, so the throttling is around 20x slower than real time.
format 18 - 640x360 AVC1, V+A : ~50kB/s
format 22 - 1280x720 AVC1, V+A : ~60kB/s
format 137 - 1920x1080 AVC1, V : ~128kB/s
format 400 - 2560x1440 av01, V : ~450kB/s
format 401 - 3840x2160 av01, V : ~850kB/s
Are you sure you are not reporting mega/bits/? When you mention the «player reported data rate», are you sure that is not the "connection speed"? 10MB/s means downloading 36GB/h (one CD per minute, dozens gigabytes per hour)...I use it to download videos into a Plex library.
function ytp() { youtube-dl --get-id "$1" | xargs -I '{}' -P 200 youtube-dl -i --embed-thumbnail --add-metadata -f 'bestaudio[ext=m4a]' -o '%(title)s.%(ext)s' 'https://youtube.com/watch?v={}'; }https://github.com/firedm/ shows no public repos and of course https://github.com/firedm/FireDM gets a 404
https://dereferer.me/?https%3A//www.jwz.org/hacks/youtubedow...
It not only completely prevents stuff like this, it profoundly increases your privacy on the web by preventing sites from tracking which domain you came from. There is no good reason any site needs to know that. I am surprised that Mozilla hasn't simply made this the default setting for all users.
I was under the impression it was, I doubt I'm the only one, so thanks for drawing attention to it.
Thanks! Just checked and it is at zero for the default. Reviewed https://wiki.mozilla.org/Security/Referrer and learned something new.
https://old.reddit.com/r/WaybackMachine/comments/kzvzxl/fail...
> it happened to me and I figured out it started when I disabled website.referrersEnabled
So if you completely disable sending the referer header, it breaks. This would probably also happen if you set `network.http.sendRefererHeader` to 0 or 1.
But that's not what I suggested! I only suggested disabling sending the header externally, to other sites, when the host domain (*.example.org) changes. In fact, someone in the thread you linked says that doing this instead of disabling it completely fixed the issue for them:
> ok I left network.http.sendRefererHeader on 2 (had it on 1), and changing network.http.referer.XOriginPolicy to 2 it works
Why?
Just copy the link to not send any referrer information.
> Connection reset by peer
in a script that I have that downloads podcasts and immediately transcodes to low bitrate opus using ffmpeg.
It's a real bummer...
But in my experience Yandex just isn't as good at searching things in English, as DDG/Bing. At least, I always preferred them for most use-cases except from some really specific ones. So it's a pity that this thing with DDG happened. I guess, I'll still continue to use it for now, because I don't want to hurt my productivity other that thing (and switching to a search engine I would have to adapt to most certainly would do that). But we'll see.
If capitalism caused censorship you would expect the least capitalist places to have the least censorship but the opposite is true.
Additionally, we don't live in a stationary society, so, whatever capitalism did or did not cause in the past might not apply as-is today.
For example, someone close to me expressed support for the "Proud Boys." This person in particular has been duped by evangelical movements in the past that all have one curious detail in common: they prohibit masturbation. Half to make fun of him, half to help him, I wanted to share a link with him to the Proud Boy's official website because he wouldn't believe me when I said they ban jackin' it. Facebook Messenger (in private DMs) refused to send a message including the link. "Send failed. Operation could not be completed." I tried an archive.org archive of it, same thing.
At least TPB adds value to peoples lives.
This isn't a rhetorical question of course. There are major arguments for both. But I think a pretty good chunk of DDG users were more often after the library index than what ideally would be a librarian recommendation, but in reality is more like a stereotypical used car salesman style recommendation.
Besides, if someone accidentally searches for shite, he can revise his search terms. Just like how you might reword something if a listener was confused.
This "curation" just doesn't need to exist. But that's a moral argument and those applying censorship aren't moral, so won't be partial to it.
Both youtube-dl and the pirate bay are available on google.
I wonder what Yandex is like these days…
There was some drama in 2010 over ED being censored in Australia, but it looks like Google has since quietly delisted it completely.
The entire job of a search engine is to curate results from a vast internet and boil it down to what is probably interesting for you. You also don't want all kind of "shock sites" on top if you search for "gun homicide" or "ISIS beheadings". Most likely you want some background information on these things, not "shock sites". That's the service Google and such provide.
I commented because learning of such special-casing of a site by Google [a] was quite memorable for me at the time; since I also see plenty of other threads around us debating search engine curation, what follows is
I came across some Reddit threads from a certain time in 2014 which noted that the top Google results for some subjects were ED pages that mentioned them. Perhaps ED's downranking was in response to that. Whatever the cause, this is actually a more severe downranking than that famously applied to thepiratebay.org [b]; for example, the query thepiratebay org foo returns only results from the actual site, [c] but the same query for ED returns no results from ED. [d]
I did find two more cases where Google does return a (single) page from ED. The sole query without any operators that does so is encyclopediadramatica online (and its punctuation/whitespace-equivalent variants), for which the Main Page is the first result and the only ED result. [e]
The other case is when searching for a phrase within quotes that occurs nowhere else in Google's index other than on ED. For example, the expected only result of the query "Encyclopedia Dramatica help pages" is the ED page containing that text. [f]
So, to be more precise, ED is severely downranked rather than delisted, albeit with the result that no one can find it on Google without already knowing its URL (or a hapax legomenon quoted from its pages).
-----
[a][b][c][d][e][f] If I may use a simplistic model of Google that determines a score for each potential search result (i.e., 'distinct' crawled URL) by starting with the same initial score for all results and applying a sequence of steps to each result that each change its score, the following is a speculative explanation for all these behaviours:
- Following the meat of the algorithm is the downrank-specific-sites step, which decreases the score of each page of a downranked site by some amount (affects both TPB and ED: TPB scores are decreased moderately to implement [b]; ED scores are decreased massively, explaining [a]).
- Then comes the query-contains-site-URL step, which greatly increases the scores of all pages of that site (TPB results are now higher than non-TPB results, explaining [c]; ED scores were decreased so much that they are still lower than all other results, explaining [d]).
- Then comes the query-is-site-URL-exactly step, which makes the score for that site's base URL (which in the case of ED redirects to the Main Page) higher than any other score (explaining [e]).
- Search operators are last, and thus the highest-priority (explaining [f] and the site: operator).
These steps have held up in general where applicable for all the queries that I've tried.
1. They will eventually require payment.
2. The service provides a lot of customisation options. Up and down ranking domains, for example, and setting up special filters. This kind of customisation can be built over many years and should carry with the user between browsers and devices.
Personally, I'm less worried about privacy and more about receiving accurate and unbiased search results. Of course, you can always use a VPN and use a burner email when you create the account.
Also curious to know the extent to which Google de-lists things?
Long term, perhaps a decentralised search engine could get around de-listing and provide a more reliable and rigorous search experience.
https://twitter.com/yegg/status/1501716484761997318?s=20&t=9...
Before someone chimes in, yes I understand the humanitarian perspective. That's not my point. My point is that DDG is not neutral, and is politically biased.
Ultimately what I'm getting at is: There's no market for DDG. Use Google for biased searches, and use other search engines that are not biased(which excludes DDG) for unbiased searches.
What are the use cases for DDG?
Maybe i was only the few that cared more about the filter bubble angle and less the "we're selling bottled privacy™" angle [1] and am not interested now that `yegg` has clearly reneged upon the former
[0] https://techcrunch.com/2011/06/20/duckduckgo-to-google-bing-...
[1] https://pictobar.tumblr.com/post/63785124046/the-banality-of...
When I want unbiased searches I've been using Kagi but more are popping up. they approach search differently so it's useful when google feels to "sanitized" for certain searches
It's not to me, when was the last time you used it? I haven't used Google in around 4 years now and I'm getting on just fine, majority of searches answered on first page.
If anything I now toggle between Duckduck and Ecosia as Ecosia still isn't 100% up to scratch (frequent 500 errors, slow, results are bad) but I like the idea of my searches planting trees.
I'd expect that all sites wanting to draw traffic would attempt to grab the reins of the search engine to point toward themselves, and the result would be search results ordered by rein-grabbing power.
Not that centralized search engines are immune to this; they're almost as vulnerable (seeing as sponsored search results exist) but the maintainer at least has to balance that with the utility of the search engine overall, to prevent the search engine from falling out of favour.
With a decentralized engine, parties that have deeply invested in manipulating the results will still want the engine to be popular too, but I'm not sure how you resolve the prisoners dilemma there as a whole.
> I'd expect that all sites wanting to draw traffic would attempt to grab the reins of the search engine to point toward themselves, and the result would be search results ordered by rein-grabbing power.
I would venture to say that a combination of allow-lists and block-lists from trusted parties, ranked using some kind of distributed web-of-trust system would work reasonably well.
The basic idea is when you rate something, your client also looks up in the DHT other people who have rated the same content with similar ratings. Your client then pulls the latest ratings collections from those people, and computes the cosign distance between your ratings and their ratings (over the intersection of content that both of you have rated). Periodically, your client signs and publishes an updated ratings document, where the rating for other raters is the cosign distance. The cosign distance, the size of the ratings intersection set, and maybe some other factors go into deciding which raters get published out in your ratings update.
When you query for the rating for a given piece of content, your client grabs the list of ratings for that content from the DHT. It then pulls the latest ratings published by those raters, computes cosign distance, and then does something similar to Djikstra's shortest-path algorithm to recursively search the DHT using these cosign distances as weights. In general, the DHT wouldn't have many signatures stored under the content's hash, but by recursively following the graph of other raters, your client hopefully finds other raters that rate things similarly to you and have rated this content. The path weight to a given rater is the product of cosign distances, and so by using a priority queue for querying, you get something close to a breadth-first search of the ratings graph. Once your client has accumulated enough weight of ratings for the given content, it stops and shows you the weighted average of the ratings (and maybe the weighted std. dev. is displayed as a confidence score to power users who have enabled it).
Presumably, the UI for the ratings system maps 0 to 5 stars to 0.0 to 1.0 (probably not linearly, more likely the client locally keeps a histogram of the user's ratings and then maps the star rating back to a percentile rating), and the "spam" button rates the content as -1.0.
The tricks come down to the metrics used for how the DHT decides priorities for cache eviction of the per-content ratings and also the per-rater ratings. You don't want spammers or other censors to be able to easily force cache eviction. Getting cache eviction metrics right is the key to having the system scale well while also preventing spammers/censors from evicting the most useful sets of ratings.
Microsoft (and any other big company) has many competing interests other than just being helpful to users.
Are the main ones i) reducing competition and ii) managing their reputation?
If so, the case for a decentralised search engine got stronger.
I don’t think anyone wants to use a search engine that never delists anything. Ransomeware, Markov chain junk, plagiarism. A search engine that never delists anything is useless.
The problem is when delisting is used against the end-user’s interests.
They use Bing as an index, and it was Bing who de-listed it.
Great engine with some really nice features
Seems to have an uncaught exception trying to use the beacon API (which I have disabled).
You.com looks like junk.
Braves search is still kind of meh, but I appreciate that they’re trying.
wow that sounds so incredibly different from the alternatives
[1] https://www.wired.com/2006/05/att-whistle-blowers-evidence/
If this decision was because of legal pressure by Google, I don't see how that got resolved in a matter of days. Which means it wasn't because of Google, but rather a poor decision made by DDG management. How can people trust their product now?
The removals were a rarity. It's not as if you can't add a `!g` bang query to redirect to Google if you can't find something. And DDG is rampant with all sorts of stuff that shouldn't be there, so I don't think they're hellbent on censorship.
You don’t always know what you don’t know.
If I don’t know about YouTube-dl and I search “download YouTube command line”, how am I going to know that ddg is hiding the best result from me?
This definitely isn’t a small annoyance kind of problem, in my eyes. It’s a deal breaker. I’ll never use ddg again. If they’re going to censor like Google does, then I’m going to use Google because it generally has better results. Ddg needs to offer something beyond what Google offers to make up for their bad results, and they aren’t doing that.
If you're searching on google for medical terms it builds a profile on you, and you should be concerned that they could be selling that information to insurers, directly or indirectly.
Is this really going on? Maybe not, but is it worth it to ignore the risk? Can you just use DDG so it isn't a question?
(except in reality it's less simple than "they know your occupation", it's a huge cloud of data points that an ai makes correlations and assosciations that no human actually knows. It means the searches would also be weighted by indirect things like, not only your occupation but your assosciations. Say you don't have the medical occupation, but your computer makes a lot of medical searches, because your roomate in your college dorm has a medical major etc. And right now, the ais are still pretty stupid and absolutely making a lot of obvious unsafe conclusions, but they also do get more and more spooky every day.)
But this doesn't make it any better. If they did a perfectly accurate job of profiling you, that is not better than doing an inaccurate job.
That much insight is like being married to someone, where they intimately know all your biases and motivations, know all your buttons, know how to manipulate you, know how to weigh any opinions you might express against their knowledge of where you got every idea you ever had,
except it isn't a marriage, and they aren't subject to all that same vulnerability to you, and they aren't even a human but a corporation, and they have this intimate knowledge of everyone not just one spouse or sibling or best friend.
Edit: I was wrong.
TPB is the first result too. IIRC at one point searching for TPB only returned proxy sites, but that doesn’t appear to be the case now.
https://help.duckduckgo.com/duckduckgo-help-pages/results/so...
"We also of course have more traditional links in the search results, which we also source from multiple partners, though most commonly from Bing (and none from Google)."
I've tried comparing Google, Bing and DDG on a private window before, and I didn't find Bing and DDG more similar than Bing and Google. Searching for monkey: https://news.ycombinator.com/item?id=27598329
Also, throughout this whole situation I always got the Github page as the first result for "youtube-dl".
Why? Their plan is to be an amazing SE, but only for the comparatively small amount of people that are willing to pay. They aren’t VC funded but boostrapped, so they don’t need to ruin a good thing for dumb returns.
It is highly likely that $5/month, for example, would attract millions of users, and $25 million/month would certainly keep the lights on.
Also, you can’t forget that every search actually costs them money, as every single search incurs the fee from Google and Microsoft for their APIs. I think there was an idea around having a $5 trial plan with only 10 searches a day or something. Unlimited searches for $5 would be "We are losing money on every customer, but we are making it up in volume"
Are they really querying Google and Microsoft for every search? This seems highly inefficient. Then Kagi is basically just a fancy wrapper. I thought their ambition was to build out Teclis and TinyGem to serve more and more content and rely less and less on other search providers. At the very least, I don't think they need to hit Google every time someone queries "games." This is hopefully queried once per period and stored in Teclis/cached.
To be blunt, if this is just a fancy wrapper/aggregator then there is no reason to use this over you.com.
And I mean, it should be easy to compare. For me, kagi is substantially better than any other SE I tried, so I’ll pay for it when it releases, but just test it yourself.
In other words it does not matter if we could have $100,000/month with another price point if it would cost us $150,000 to do so (every search has a fixed cost that does not go down with scale).
Our current price point includes a tiny margin that would allows us to break even at around 50,000 users. That is what we are optimizing for.
Our profit margin is already razor thin, I do not think it can be further optimized. Meaning we can not further reduce price to get more users without operating at a loss.
It sounds like you have a high marginal cost. As in, you pay every time someone searches for something. Can this not be reduced over time with greater efficiency of scale?
FYI I use Kagi and like it. I’ve replaced Google :)
And it is only about one cent to search the entire web in 300ms with everything else Kagi does, that has to be impressive ? :)
[0] https://techcrunch.com/2011/06/20/duckduckgo-to-google-bing-...
or
Duck Duck Go
Also if anyone else is slightly hung over and searches DDG in DDG and gets super confused for three seconds because DDG is a rapper apparently just know you aren't alone. I'm right here with you, whoever you are.
> ... [W]e are not "purging" YouTube-dl or The Pirate Bay and they both have actually been continuously available in our results if you search for them by name (which most people do). Our site: operator (which hardly anyone uses) is having issues which we are looking into.
(Note that "site:" in his comment is how you restrict DDG searches to a specific domain e.g. "site:example.com")
I noticed problems with their site: operator too earlier today, and still now as well. In my case, when I used it in a search I saw that the word “site” itself was also bolded in some results. So it looks like it is using the operator itself also as a search term, which it shouldn’t.
I find it surprising that hardly anyone uses it though.
Same! I use it for searching reddit all the time.
> For example, searching for “site:thepiratebay.org” is supposed to return all results DuckDuckGo has indexed for The Pirate Bay’s main domain name. In this case, there are none.
> This whole-site removal isn’t limited to The Pirate Bay either. When we do similar searches for ["site:1337x.to", "site:NYAA.se", "site:Fmovies.to", site:"Lookmovie.io"], and ["site:123moviesfree.net"], no results appear.
https://torrentfreak.com/duckduckgo-removes-pirate-sites-and...