The Scene: Pirates Ripping Content from Amazon and Netflix
torrentfreak.com
torrentfreak.com
I prefer to use the third option as few as possible, but I find myself having to use it more and more often as the years go by - geolocked content, huge swathes of content being removed on the whims of providers (The Office comes to mind), and 'originals' that are beyond even the so-bad-it's-good category that cannot even begin to fill the void left by the good content leaving.
I cannot ever see myself subscribing to a plethora of services - the best case is subscribing to a service when it's absolutely necessary (HBO/Hotstar) and then unsubscribing when you're done, which is not at all conducive to fulfilling long-term revenue plans of such services.
So what is the endgame of such services? Is it simply to build a huge catalog of content that you'll mindlessly devour whenever you're bored? Or something else?
Similar gripe, finding out that only the sequel or the last movie in a trilogy is available: why Netflix, why?
I wonder how many content providers see the popularity of old shows and think they can pull in the audience, without realising that a big part of the appeal is having all the shows in one place with one subscription?
Couldn't agree more. Besides a very few "tentpole" features, and a small number of standup comedy specials, that Netflix Originals logo up the top left may as well mean "do not watch". If I could exclude them at the UI level, I would. I pay to see top-range feature films and high quality productions, not an endless bargain bin of third-rate filler.
I suspect the "endgame" is that after all the studios have tried out their little walled garden plan and most have failed, they will band together and finally create some kind of unified system, same as the RIAA has done. It'll probably take a good few years, though.
This is only possible in a world where nonsensical bullshit is rewarded with a big pile of money.
In a world we'd all want there would be ONE SERVICE on earth that has EVERYTHING. As long as they don't get that I'll be a pirate.
At least the big companies don't suffer from the loss of my money as a little band would do. That's why I still buy vinyl and visit concerts but I have no choice to do that with movies/series I'm afraid.
To the people who always argue for the "free market": Is this what you meant by "competition is good for the consumer"?
I used to watch a lot of Monty Python sketches on YouTube, and in the past few months they've been removed. Last time I looked it was impossible to find the super popular Parrot Sketch!
They're not on Archive.org, no decent results on alternative sites such as DailyMotion, just disappeared.
That's absolutely shocking to me, a huge piece of popular culture is unreachable unless you specifically go pay for it or are lucky to find a collection on a torrent site.
Even content that is constantly marketer directly to me by ESPN is US only (ESPN plus) despite my Canadian geo. It’s annoying.
After Fox, Disney et al pulled content from Netflix, coupled with regional restrictions, there’s literally nothing to watch on Netflix in Sweden except Netflix’s own original content.
Hulu isn’t available outside of US.
HBO is just as bad as Netflix outside the US.
Even iTunes Store is subject to regional restrictions (but the selection yhere is still much larger)
Isn't piracy illegal though? I assume it will get deplatformed
Piracy decreased because the alternative became more convenient. It's easier to pay a few bucks on a monthly basis than it is to download the torrents. But, as those fees are increasing alongside the number of services that you need to pay for, piracy is bound to fill in the gaps. It's illegal, but the less people could afford the alternative, the more difficult it is going to be to stop piracy. A domain or two could be shut down, but if TPB is of any use now, it's to remind us that it's not easy to stop a highly-motivated group of people. I'd argue that The Web is still hectic enough to accommodate such a group of individuals.
Piracy was significantly reduced by Spotify and Netflix because it was more cost and effort-effective to just pay a monthly subscription to get most if not all content you need.
And then copyright holders decided it wasn't good enough and crippled Netflix without providing any alternatives (as always). Watch them blame pirates again in the nearest future.
I am not so sure about that. It will always exist, but it is getting increasingly difficult to break DRM. So we may have long periods where the rips are inferior compared to the source (i.e., transcodes via capping). All the new 4K stuff is already quite difficult due to SGX/other hardware enclaves. So it's not a question of just time/dedication to reverse engineer a blob.
Also, hardware does get hacked too.
When life gives you lemons you JTAG those keys out the enclave if you have to. Don't worry, life always finds a way.
I'm not so optimistic. That is precisely the threat model. That is not to say that it is impossible. Side channel exfiltration from stuff like SGX has been shown, but that's a different game really.
And because DRM is fundamentally broken (i.e., there is always the analogue hole at the end of the spectrum) it will never prevent piracy on its own.
Do you have real examples of where this is the case? Outside of a vanishing minority of titles, DRM-stripped releases appear in a timely manner just as they have always done.
Pirates are going to prison for providing a better service at lower cost, the core essential of capitalism.
Intellectual Property is a strange concept, but there is something going on, after all movies and TV shows have enormous budgets and they depend on revenue, which is still basically made from individual cinema goers and TV channels (ads), and subscribers (no ads).
The never expiring copyright is the problem, and the very strict fair use (eg sampling, and fanfics are illegal unless proper licences have been attained).
This leads to relatively few complex and kind of efficient for-profit content producers, and a lot of very small not-for-profit enthusiasts (who are probably hyper-efficient in some dimensions, but since their competence is naturally limited, they will not be that efficient in all the relevant/required domains). So we have Hollywood, a few independent filmmakers doing the small budget creative stuff, the regional markets with their respective woods (eg India, China, Japan), plus usually countries have some funding for local productions, but that's it.
And maybe even this would be more than enough weren't almost all of the content locked up in archives. (And only available maybe on torrent, or maybe on a random on-demand pay-per-view site.)
Of course it works marvelously, for example there are shows that otherwise are inaccessible to many. (Animes come to mind.)
If there were some kind of pay as you wish backchannel probably a lot of people would happily give more than zero to content creators.
There is Hulu in Japan, but with a different catalogue afaict.
Edit: I'm referring to the content ripping groups. Apart from the technical and legal challenges depicted in the article, they need to seed their files which (must be) very costly. Also, content is released consistently and extremely quickly.
The common answer is that the motivation is non-monetary: Fun, community service, group pride etc. But I find that hard to believe entirely. The product is just too polished, subjectively resembeing in quality a revenue generating operation rather than a(n illegal) side hobby.
It’s not about money… it’s about sending a message.
Not really surprising when you look at how much non-monetary development goes into the open source community.
The reason is that I'm trying to collect data for learning languages, especially Chinese. Kind of like VoiceTube [1] and Youglish [2] are doing for English. So far it's only for personal use, not a revenue-taking company.
I tried investigating a Netflix stream, but the subtitles weren't plain text! They were distributed as PNG files [3].
Does anyone have more information about this? Do the pirates really OCR the subtitles for their MKV files? (I doubt it). Is there another way to get the plain text? Contact me directly if you're afraid to comment publicly.
[1] https://www.voicetube.com/
[3] https://www.slideshare.net/RohitPuri23/timed-text-at-netflix...
You can definitely send subtitles as raw text and metadata – it's what they do on iOS devices IIRC, since you can change the subtitle appearance on the client.
Also, their ingest works with TTML (a text-based format): https://partnerhelp.netflixstudios.com/hc/en-us/articles/215... – this is a fairly widely used standard.
Yes. At least as far as I have seen, which is not scene level but in private tracker discussion forums.
OCR-ing digital text is pretty accurate afaik.
Viki has a lot of subtitled Chinese shows. They generally don't offer Chinese subtitles for the Chinese shows though.
The show itself often will, but those subtitles are just part of the video, not even an overlay.
https://www.youtube.com/playlist?list=PLhIooD7mFhphhT5nDdhK0...
Subtitles sites (opensubtitles, podnapisi, etc.) have similar data and simple alignment algorithms will work for unaligned subtitles.
> Do the pirates really OCR the subtitles for their MKV files?
OCR is really simple for text. No reason why they wouldn't do it. There were free programs decades ago for OCR when you were ripping DVDs and they worked mostly on character patterns (you would start with an empty database and then just tag the pattern, after a minute you'd have the whole subtitle OCR-ed).
People/place/town names align stuff pretty easily. The fact that a sequence of lines matches some other sequence of lines makes finding the optimal alignment very efficient too.
I'm guessing that just making a set of words for each subtitle line, counting the common words and picking a criteria of deciding if line maps to other line is more than enough. Subtitles are much easier than free-flowing documents because they are time constrained.
At least the subtitles on there are already aligned English & Chinese, but they do require OCR. Message me directly if you want my OpenCV script that can pull the yellow text out.
Tesseract gave me this. The W0「d n! Gnd has bP〔‥mE ‥任Sh 嬰孩降生 道咸肉員 Peace has Come for O… Km… 峒 ﹏m US 靜安 來自‵我君手 `
I'm sorry to continue insisting, but seriously, Tesseract gave me terrible results. Identical lines don't even give the same OCR result, so trying to run partial matches on Google Translate-quality word equivalents isn't going to cut it. The OCR can't even give me the right number of characters! In the end I'm trying to do it manually, but it's really time-consuming.
[1] https://www.youtube.com/channel/UC8_emPVKZvOwJnoBVpU0ytg/vid...
I'll try to make a simple one in the next couple of hours.
Although, the task you're asking me to do is different from subtitle alignment you could do on .srt files.
What you can do with leptonica is extract the symbols (letters and ligatures) and manually tag them. You could also extract the lines containing letters and then process them through tesseract.
我們傾倒耶穌腳前
This is my output from tesseract (chi_tra) (for https://www.youtube.com/watch?v=RzIh2pamcwU)
I have just extracted the U component (YUV colorspace), binarized the image and tesseract just loves it. The players in the background are completely removed.
I would probably use some background removal or some morphological operations from leptonica to make stuff more robust (instead of binarization).
Of course, if tesseract does something incorrectly you can do a quick connected components in leptonica, extract the symbols and do your own OCR (manually). For chinese it's a bit of an issue because there are bunch of symbols, but for english you will be done tagging all the unique symbols in no time.
tesseract 3.04.01 leptonica-1.74.1 libgif 4.2.3 : libjpeg 9c : libpng 1.6.34 : libtiff 4.0.9 : zlib 1.2.11 : libwebp 1.0.0 : libopenjp2 2.3.0
If you could email me your scripts, that would be extremely useful. And then I'll send you some subtitles where alignment is important. I just used this channel as an example of where I'd tried Tesseract and given up.
That kind of error ratio isn't in the realm of "correcting common errors". Surely there must be better software, but I can't find anything that runs locally and can read these subtitles. I've resorted to typing them out by hand, which is super slow because I don't yet speak Chinese, so I have to use handwriting recognition to enter each character.
[1] https://www.youtube.com/channel/UC8_emPVKZvOwJnoBVpU0ytg/vid...
It's literally self-deprecating behaviour.
On two of the three 4K monitors I own, I have no way of watching any films or TV shows at that resolution whatsoever. There is no mainstream content at all that can be consumed on Linux at 4K through any legitimate service that I know of. This is in spite of me subscribing to a ‘4K’ Netflix service, owning multiple movies in 4K in various forms etc. The only device that supports 4K legitimately is my Apple TV, which I bought simply to stream Netflix at native resolution. Now that I have it, I often find myself renting films on it, but the only options there are for watching on my laptop are piracy. And honestly, piracy is higher quality and a better UI (being able to download in advance, the press play with no worries about buffering).
- Linux Desktop - 21:9 monitor
? If you guys are cracking one 10th of a percentile I would be surprised. I also use netflix on linux desktop (regular monitor though) but I cannot pretend that they should have to care about such a tiny portion of their overall userbase.
P.S. on Netflix, I should verify on Amazon Prime
Edit: after posting I recalled a plugin for Chrome/Firefox that forced 1080p playback on Netflix, the github page has the details https://github.com/truedread/netflix-1080p
Huh? Are they recompressing it? How else would it shrink? If they're recompressing it, why don't they just use a capture card and avoid the hassle of reverse engineering stuff?
Video engineer currently implementing MPEG Common Encryption here:
Actually, there is.
MPEG CENC (Common Encryption) uses AES CTR and requires conveying one initialization vector for each audio/video frame. Those are generally 16-byte long.
To make the matter (a little) worse, MPEG CENC adds signalization to the file, to specify which parts of the streams are encrypted or not: indeed, for remux-without-decrypt reasons, each video frame must have its headers in cleartext.
In fine, you get some dozen of bytes of overhead per audio/video frame ; do the math.
Of course, this overhead remains negligible and could by no means explain a 30% size reduction when decrypting.
Creates an obvious problem though.
Aside from the quality loss due to re-encoding - which should be identical to the loss encountered when capturing digital video and audio output, assuming the capture is done correctly - I fail to see anything "obvious".
Edit: I just saw parents reply. It took a while to type this out on my phone :/
I suspect what they are referring to there is the process of remuxing the file and removing extra audio streams, ie. a BD50 BluRay of "They Shall Not Grow Old" comes in at 20GB with two audio streams, DTS-HD and Dolby Digital. The remux of that BluRay comes in at 16GB and has an identical bit-rate on the video stream, but the Dolby Digital audio stream has been removed, as well as the menu graphics (though this would not apply to a WEB stream).
Everything else probably 95% you see in the wild are transcoded to lower bitrate and stripped out of all the extra files in the container (sub/multiple audio stream/highbitrate audio) to smalled files - because thats what most people want:
Small Video easily playable on a regular screen.