ByteDance is abusing the free video downloading service Cobalt for mass scraping
twitter.com
twitter.com
> it's very unlikely to be someone else because pricing is astronomical. you also have to "contact sales" to get access to anything outside of a free trial. no one would pay that much for a block of ips with terrible reputation
You don't have to contact sales if you are a Chinese-speaking customer. And pricing is fine. ByteDance has a different brand for their cloud services in China: https://www.volcengine.com/ [1]. But of course the underlying infrastructure are all the same.
This is very likely done by a Chinese customer using ByteDance's cloud service.
[1] Well, Alibaba Cloud did this too, and ByteDance is copying Alibaba 1:1 (who in turn is copying AWS) so I'm not surprised. But at least Alibaba named their international brand "Alibaba Cloud" and their CN one "AliCloud", similar enough.
Donations are a monetary incentive
> while heavily breaking the TOS and potentially ignoring copyright laws
Cobalt also breaks the TOS and ignores copyright laws, personally I don't think that matters but having a double standard when one company does it "It's ok when they do it" and when one you don't like does it you try to use copyright laws and TOS as a weapon just makes me think it really isn't about TOS or copyright is it.
Also just gives YouTube ammunition to impose stricter protection against smaller violators like cobalt, like self running yt-dlp
“bytedance's scraper was specifically built to go around cloudflare & other web security solutions, which is just genuinely evil”
So I would say it’s a fair comparison.
Then they either didn't set up CF correctly or they just use the mode in most headless browsers that bypasses default CF protection when CF is not in attack mode.
can't say the same for bytedance, which is designed to exploit users with various ads
If your opinion changes because the owner is different, even though the service stays the same, that's hypocritical.
I wrote to them to please stop (I think the address was in the user agent or something), they replied sorry and actually stopped.
Not sure why all these crawlers can't pace themselves.
faster, bigger, MOAR
sometimes it’s hard to have nice things
Maybe this is some ByteDance engineers getting really desperate and resorting to abusing every youtube proxy service they can because apparently they do have a residential proxy network which doesn't cut it anymore?
Unless it's just a cost-optimization measure (residential proxy traffic is relatively pricey).
Still, notice how most of the low effort avenues are slowly being cut off one by one. I will use non youtube example. Not that long ago, I was able to rip blurays using off the shelf external bluray writer, but new firmware on currently sold drives remove that ability.
Now, Google typically won't be ( and isn't ) everyone's hardware provider, but there are ways they could degrade 'non-sanctioned' experience in browser they can ( and do ) control.
Granted, in Firefox ( and other non-google browsers ) it may not be as simple, but future there is not as straightfoward either given Mozilla's trajectory and financial dependence ( and moves ).
In short, I agree with you but note that initially it was genuinely trivial to download youtube videos. This has changed over the years.
That's already the case for some (anecdotally increasing ) number of videos.
I'm not doubting the OP. But why is ByteDance doing this? What does that company get out of scraping YouTube?
https://decrypt.co/284353/tiktok-maker-powerful-ai-video-gen...
earlier today i noticed very elevated traffic to cobalt api that looked a lot like ddos. it turned out to be bytedance!
we can't tell what videos they were downloading or where the original request comes from as it's built to go around all limiters, but there's still a pattern
first request: json post with content url & settings from a residential proxy
second request: tunnel with pseudo microsoft edge on windows user agent & youtube origin/referer, from byteplus ip
third request: same tunnel with aria2 user agent & no referer, also from byteplus ip
cobalt is a media downloader, mostly known for supporting youtube even at worst times. cobalt's tunnel is either a proxy stream or ffmpeg live render
considering all of this, i can safely assume that bytedance was scraping youtube videos by abusing our private api
with release of v10 we implemented cloudflare turnstile, but later disabled it due to access issues by a chunk of our users
enabling it back brought the server load to normal levels and stopped bytedance from choking our servers cuz they didn't account for this (yet)
before resorting to turnstile, i attempted using other cloudflare services, but none of them seemed to help much
my theory is that bytedance's scraper was specifically built to go around cloudflare & other web security solutions, which is just genuinely evil
this incident caused a few minutes of api unavailability, but taught me that cobalt (and probably anything else) can no longer exist without active bot/scraping protection
im really glad that cloudflare turnstile exists because i don't know what i'd do without it here
byteplus AS that was spamming requests is 150436 and last seen ip range was 207.166.160.0/21
the amount of unique users on cloudflare analytics rapidly increased by 2.25 times and didn't go down since, while web analytics (plausible) show no increase whatsoever
> im really glad that cloudflare turnstile exists because i don't know what i'd do without it here
Why not just blackhole the byteplus ASN?
I've posted a canary token[1] URL as a Mastodon post, to check how scrape-resistant Mastodon actually is (it is not resistant at all), and have been getting quite a few hits from the ByteDance spider recently.
Last hit is from 47.128.114.151, Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com)
Edit: added missing footnote.
That's expected, no? It's a social network that is explicitly designed to be as open as possible, as it's using ActivityPub. To be "scraping resisting" would be to go against the very goal of Mastodon.
If you look at the technical side of things, you're absolutely right. If you look at the social side, however, there's a lot of talk on there about opting out of scraping, scrapers being bad, not wanting to be part of AI training and so on. Naming-and-shaming people who have been caught scraping is a routine practice.
I think that many Mastodonians believe that defederating from scraper-friendly instances and blocking scraper-like requests on their own protects them, this was a way to show that this very much isn't true.