DMCA Notices Took Down 20,517 GitHub Projects Last Year
torrentfreak.com
torrentfreak.com
Gamely, I accepted the challenge. Browser dev mode to sniff things out was very difficult - they put a lot of work into jamming that up. Turns out it is trivially easy with Fiddler.
What I learned in the process was interesting, trying on a few video files. What these sites do is slice up their video files into thousands of chunks, and then rename the files as innocuous file extensions (like *.jpg, *.html, *.csv, etc.) and stick them on various hosting services where you can do a GET on the file for free with no authentication. They spread those thousands of files for each video file all over the internet. Usually across multiple accounts / repos on a service like GitHub but not just GitHub - any place they can manage. A lot I looked at while testing had files on GitHub though.
Then, once the thousands of files are distributed, they stitch them together via an *.m3u8 playlist file which the in-browser player can handle buffering to make it all seamless.
It's also quite nice how they've neatly open-sourced and atomized all the little tools that go into making the service work such that one could, in theory, build other services on top of these tools. The quality of the documentation is outstanding, too.
Astonishing is the right word. It's like "neat sparkling new JS-based startup" meets "mid-2000s underground piracy", and I can't get enough of it.
Unfortunately I'm unable to access twitter threads now so ymmv - I think it still works.
The TL;DR is that the playlist metadata file (mpd or m3u8) references multiple fragments, and arbitrary file offsets within those fragments. Slice your video into fragments (using standard video processing tools), embed them into PNGs using standard twitter-compatible file format polyglot tricks, and upload them to twitter.
I'm sure there's a certain amount of laziness that also comes with how keen each site is on enforcement. The most recent one I looked at, all the files were on the same site, in the same folder.
That's an odd assertion. Some of the smartest and most sophisticated people I've ever interacted with were very actively involved in the piracy scene. These types love a challenge, and that's why they get involved in piracy. In some cases they aren't even all that into the content itself.
I don't unilaterally hate all IP (other than patents) btw, so this is not an endorsement of piracy itself, just a frank acknowledgement that a lot of these people are very smart.
These streaming sites are laden with ads, and some even have offer subscriptions, so there's a profit motivation for the people that run them.
I'm not talking about the people who upload torrents (there's not much in a profit motive there, since that stuff is hard to monetize), or people who crack software. That sort of stuff. Interviews digging into their motivations have been going around for a long time now.
If you're hacking on running a pirate streaming site to make money off it, if you have skills enough to stay ahead of the people constantly trying to take your content down, at some point you're going to realize you can make a lot more money using those skills working for a legit company.
My take on the profit motive for malware authors is not the same; a malware dev can make a lot of money if the group manages to hit the right victims. The ones who are state sponsored probably make decent money too.
If you're interested in some of the tools used that essentially re-unites from endless sources, auto determines bitrate, allows you to interject interstitial ads based on geography/user tokens etc, check out things like bitmovin. No relation, I've just used it. Then cedexis for the latency/error-rate magic that makes your cdns compete against one another. There's also jwplayer which a lot of video platforms like Univision and Fox use.
https://bitmovin.com/ https://github.com/bitmovin
Another interesting thing is how the DRM is handled, the video chunks that are all muxed/reunited/demuxed, etc on the fly- https://developer.bitmovin.com/encoding/docs/how-to-protect-...
I started out in video stuff back in RTSP days (2005ish), took a LONG break in that area (2016ish) and the systems today are so cool. IIRC back then everything was rtsp.
I actually have no idea beyond what you just mentioned how the pirate side does things so that's neat. I wonder what tools they're using.
https://jwplayer.com/ https://www.cedexis.com/ .. which I guess was bought by citrix so I have no idea whats up with that now.
I've heard.
If read-only access is all that is needed, SSLKEYLOGFILE might help. https://my.f5.com/manage/s/article/K50557518
But they have some script in there as well that checks for if you're in dev mode somehow, and as well on top of that misbehaves if you're in private mode.
At that point I got tired of fighting it, and realized none of this Javascript nonsense was going to work against Fiddler or other proxy-based solution.
I did try streaming one time with that method, besides the high latency (~5seconds), it was CPU intensive for both the server and the client(browser), so not so sure how viable this would be.
If anyone is wondering the other piece of this, which is re-unifying all the thousands of video fragments, the game is over once you have the URI to the playlist file. You feed that URI into VLC's "Convert/Save" functionality and it handles it like a champ. No need to script out anything to grab the pieces and work FFMPEG magic.
At the end of the day, it is a 2 step process I could document for my friend who is more or less a layman. A skilled user (decent with Excel), but not a power user (can't script).
You left out an important part which is that they're proxying these requests through their own servers to avoid the source being found.
This way they could be placing the files on servers that actually require some sort of authentication such as a login page and once that's done the sky is their limit as you could even be hosting files on Discord and it would be very difficult to find the source.
User -> Pirate proxy server -> Github (or some other site)
I sometimes wonder if I'm missing some relevant change in law or case law, as the sections seem easy to read as somebody who isn't a lawyer.
In a sane world, civil suits between firms should really only go ahead if the case's result is ambiguous, and the courts in this case would almost certainly side with the takedown request - and moreover would be pissed at Github for wasting the court's time on it.
It's not that users are irrationally responding to DMCA requests when they don't need to, its that as a service provider, GitHub defaults to a stance where a claimant's DMCA is automatically processed if a repo author doesn't reply within 1 working day. This effectively means that for DMCA requests of all types on GitHub, "no response" is synonymous with "I'm guilty of infringement, plaese take down the content".
A DMCA claim response in GitHub is admissible as testimony should the case ever make it to court, and therefore carries the penalty of perjury (both parties are informed of this before opening/responding to the claim in the GitHub UI). So unless the repo author is absolutely certain that they're not hosting DRM circumvention software, they have little-to-no recourse but to allow the takedown request to go ahead.
All of the safe harbor provisions and takedown measures have to do with hosting copyright infringing material. They make no mention of needing to do anything if you're hosting DRM circumvention technology.
By your logic, I could submit DMCA takedowns over libel and the service provider would have to take it down despite it being obvious that libel is not part of the safe harbor provisions.
It also criminalizes the act of circumventing an access control, whether or not
there is actual infringement of copyright itself.[1]
I couldnt substantiate this claim (circumvention of access to non-copyrighted material) by going through 17 U.S.C. 1201.[1] https://en.wikipedia.org/wiki/Digital_Millennium_Copyright_A....
If I pick a lock for a safe I bought, it’s not illegal. But if I do that with software it’s illegal? Fuck that!
If you try to report spam/malware on the other hand, good luck. GitHub only lets you open 2 tickets every few days and each has to be for a specific repo.
I've tried to tell them how I was able to easily find spam/malware by using the search and they shut me down.
These are covert intentionally malicious binaries that attackers were hosting on GitHub and trying to get victims to download and run.
Also, lots of empty repositories with descriptions like "download leaked onlyfans of her here: <url with tons of malware>".
I'm not going to waste my time trying to report that anymore though.
- Nilay Patel [0]
Does DMCA help to make this happen less often?
As for attribution... why in the world can't we just prosecute people for fraud? Fraud has been illegal literally the whole time. There is no reason it should be resolved with copyright or contracts.
Being able to obtain, manipulate, and reshare bits of culture like videos is a huge source of creativity. Even in the commercial creative world, being able to download a trailer, advertisement, or really any video with a particular effect and take it frame-by-frame to dissect and understand it is a powerful learning tool.
And it's valuable to be able to archive works that are likely to be torn down or that need to be accessed offline. As an example of the latter, I'm currently a high school software teacher, I often download videos demonstrating this or that and share the MP4s with my students since YouTube is blocked.
When I was in the university, some of my roommates were in the film department. Their was this understanding that you could use a certain number of seconds of a copyrighted video in your project.
But in practice today, it is hit or miss. I have seen a TikTok stitching Harry Potter to make a parody taken down as violating copyright.
While there exists other videos on TikTok doing the same thing for longer duration.
I would say if there is anything you actually like on Youtube you're a fool if you don't download it.
Aren't most people here pretty big on "right to be forgotten"?
> No. When you receive material under a Creative Commons license, you may not place additional terms and conditions on the reuse of the work. This includes using effective technological measures (ETMs) that would restrict a licensee’s ability to exercise the licensed rights.
> A technological measure is considered an ETM if circumventing it carries penalties under laws fulfilling obligations under Article 11 of the WIPO Copyright Treaty adopted on December 20, 1996, or similar international agreements. Generally, this means that the anti-circumvention laws of various jurisdictions would cover attempts to break it.
> For example, if you remix a CC-licensed song, and you wish to share it on a music site that places digital copy-restriction on all uploaded files, you may not do this without express permission from the licensor. However, if you upload that same file to your own site or any other site that does not apply DRM to the file, and a listener chooses to stream it through an app that applies DRM, you have not violated the license.
It seems you would have to link to a download from the description? Hiding the derived work some place else seems to get in the way of reuse.
If the material is already distributed elsewhere and you post it as~is on youtube it seems the license is satisfied?
For example, this - https://github.com/orgs/community/discussions/105156
Edit: Turns out the repo got taken down