As I understand it, they use HTTPS to prevent spying and data-mining by unscrupulous ISPs. It doesn't affect their DRM at all, which would work just as well over plain HTTP.
As I understand it, they use HTTPS to prevent spying and data-mining by unscrupulous ISPs. It doesn't affect their DRM at all, which would work just as well over plain HTTP.
Because secure should be the default where we can do it.
If I'm watching a show, I don't necessarily want every single router along the way knowing what show i'm watching. Either because of advertising profiles or just because I don't want people knowing that I'm binge watching a specific show, or to go to an extreme, perhaps because what I'm watching isn't legal where I live.
If I'm watching a documentary or news show, I don't want the possibility that someone could modify my stream and change parts of it. Misinformation is a big deal, and if we have a simple solution to get rid of whole attack vectors (especially before they become common), why not use it?
And TLS gets rid of other attacks as well. token stealing is a lot harder if EVERYTHING is TLS encrypted vs if only some things are. It can also secure against vulnerabilities in the software that reads the streams. They no longer have to harden their systems as much against potentially hostile data, they can be behind well tested TLS termination that only allow verified data through to the "backend".
Everything should be encrypted in transit, zero exceptions. Just because you don't think it's important to hide doesn't mean that others don't. Just because you think your parsing is secure doesn't mean that another layer isn't useful for defense in depth. We have the technology. It's free, cheap to run, and doesn't hurt the user experience in any significant way, so why not use it?
Unless Netflix has also started padding their compressed video segments to a deterministic size, TLS doesn't really help with the "I don't want people to know what video I'm watching" problem, as the encrypted media segments are first compressed using a variable bitrate encoding that causes a unique pattern in their resulting chunk sizes; and so, with a little database of signatures--apparently even just at the TCP level, without caring at all about the HTTP and TLS layers... which surprised even me, and I'm quite cynical about this stuff--you can pretty rapidly fingerprint what movie someone is watching from a passive traffic dump.
https://dl.acm.org/doi/10.1145/3029806.3029821
> Identifying HTTPS-Protected Netflix Videos in Real-Time (2017)
> After more than a year of research and development, Netflix recently upgraded their infrastructure to provide HTTPS encryption of video streams in order to protect the privacy of their viewers. Despite this upgrade, we demonstrate that it is possible to accurately identify Netflix videos from passive traffic capture in real-time with very limited hardware requirements. Specifically, we developed a system that can report the Netflix video being delivered by a TCP connection using only the information provided by TCP/IP headers.
> To support our analysis, we created a fingerprint database comprised of 42,027 Netflix videos. Given this collection of fingerprints, we show that our system can differentiate between videos with greater than 99.99% accuracy. Moreover, when tested against 200 random 20-minute video streams, our system identified 99.5% of the videos with the majority of the identifications occurring less than two and a half minutes into the video stream.
Note that this exact same kind of attack has been applied successfully to other places where a pattern of sizes can differentiate usage, including: Google Maps tiles (which are either variably-compressed fixed-dimension images or variably-dense fixed-region vectors), web pages (which are different sizes already but also link to images and other resources that are random sizes, which helps create the fingerprint), and non-personalized type-head search queries (where the result for each letter brings up some unknown set of titles and URLs; as you type each character the sequence of sizes for your intermediate results can expose the query, potentially even without needing any model of the likely keywords).
Maps: https://ioactive.com/ssl-traffic-analysis-on-google-maps/
Search: https://eprint.iacr.org/2014/959.pdf
(edit: I just realized it might be interesting for me to note that I am the lead developer of a new VPN protocol called Orchid and have "attempting to solve as many of these issues as generically and yet with as little required overhead as I can" on my todo list; to be very clear: I don't handle any of this stuff yet at all, but should get to it in the next few months... while the company behind Orchid that is paying me to work on the protocol talks a lot about itself already, I'm really intending to do my own "launch" of the product once I solve some of the stability issues and finish my traffic analysis mitigations.)
But in the end it still doesn't mean TLS isn't worth it for all the other benefits (not that you implied that).
The issue then is that no matter what you do with the compression algorithms, the user is downloading these chunks and you can see the pattern of packets in one direction setting up requests and the packets in the other direction replying with the chunks and you can fingerprint what movie someone is watching. It isn't some kind of encoding issue, as the data is all encrypted: it is the entire concept of taking fixed length segments of a movie that will compress to some non-determinstic size. If you take the first two minutes of every Star Wars movie, divide each up into 10 second segments, and then compress those segments, the sequence of sizes of the segments will be pretty unique.
What is so great about this particular paper is that you don't even really need to analyze the TLS layer and try to pay close attention to really figure out the request/responses: they just fingerprint the TCP flow and that's sufficient, which in retrospect doesn't surprise me as what you are really looking for is some kind of rate of requests to responses for the chunks over the course of those first few minutes of watching the video, and don't really need to know for sure where the boundaries are: you have a long enough sequence and a small enough catalog (there aren't tens of millions or billions of videos on Netflix) to get a really strong fingerprint using just the relative rates.
To fix this you really need to either inject extra random traffic (such as extra packets to the server that break up the request rates) that adds so much noise that you can't figure out the signal "in time"--if it takes longer to fingerprint a movie than the length of a typical movie, that's "good enough"--or you need to destroy signal (which is a better description of what we do if the video segments are all the same size: at that point all movies are by definition the same sequence over and over again; if you then pad the length of every movie to the same 4 hour runtime and force the user to download all of the padded black video frames, you essentially 100% solve the problem ;P).
That said, as noted by someone else on this thread, there has been some work done figuring out what people are saying by analyzing encrypted speech packets; but I imagine that kind of technique would be almost impossible to pull off with these segments on the order of multiple seconds long, and including the video in the stream would seem to make that a total non-starter.
Every Netflix client that is watching say 1080p Rick & Morty s1e2 fetches an index of the bits that make up that video, and then fetches the relevant chunks of data.
So if you just fetch that index, no need to "watch" the video, you know OK 1390450 bytes is the 1st chunk of data making up that video, and then 2046917 bytes is the 2nd chunk. When you see that somebody fetched 1390450 bytes of encrypted data from Netflix, they're probably watching s1e2 of Rick & Morty in 1080p. When they fetch 2046917 bytes of encrypted data that confirms it.
So it's "non-novel" but they don't need to (relatively expensively) actually watch the videos they want to match. It's likely any sophisticated adversary has an index of all the common videos from sites like Netflix, the Play Store, Disney+ and so on.
There is definitely more a client could do to make this harder, but there'd be a sharp trade-off where to make it impossible for passive eavesdroppers to know what you saw costs both you and the "broadcaster" a lot of bandwidth, and bandwidth ain't free.
I'd like to see Netflix at least gently stroll in that direction, maybe use some TLS padding to make it a bit harder to do this trick, but I don't expect them to set off at a sprint and ruin their profitability in pursuit of a non-goal.
Off the top of my head: if you took a live streaming approach along the lines of live TV, or indeed Google Stadia/GeForce Now/OnLive, then you could presumably make the traffic patterns invariant over what was being watched. You'd need to ensure the bitrate didn't drop during low-detail scenes, I suppose, similar to what you mentioned re. Skype below.
That approach would of course be far more intensive on server resources, but I'm sure a similar level of privacy could be achieved without such inefficiency.
for a single user it's possible. for a network with gigabytes of data, it's basically impossible. it's not just a "tcpdump" and compare against fingerprint.
btw. it's also outdated.
also eme basically solves this on the audo and video side. (maybe not the best solution tough)
the attack inside the pdf does basically mention the old technology which netflix used and not dash+eme inside video element (MPEG-CENC). i.e. widevine which is kinda funny because the browser drm is highly controversial, but in this case serves as additional privacy since video metadata and content get's encrypted, no matter which transport layer is used.
EME encrypts the video content, yes, but I don't see that this changes things. Netflix uses HTTPS (TLS) encryption over the top of the EME encryption, as this thread's title states.
> even metadata is encrypted.
I don't believe so. If you're delivering EME-encrypted blobs over insecure HTTP, an ISP will be able to see which blobs you are requesting, simply by their URL.
Aside: I recall reading that Netflix's CDN servers ('OCAs') store EME-encrypted blobs, so only the HTTPS encryption burdens the delivery server's CPU. Unsurprising, of course.
> without eme you can basically extract the video via tcpdump
Six years ago, sure. Today, no, as the stream is sent over HTTPS.
> in this case serves as additional privacy since video metadata and content get's encrypted, no matter which transport layer is used.
Perhaps unencrypted streams are still used to support legacy devices, but Netflix are committed to maximal use of HTTPS.
I don't see how EME's additional layer of encryption changes anything as far as privacy and unscrupulous ISPs are concerned.
Also you don't need a precise measurement. Just measuring a random 10k users in 1M network daily will give you more precise viewership measurement than anything Netflix publishes.
edit - said Skype initially but not sure it was specifically skype
Unfortunately for video CBR is generally thought to be too expensive to be acceptable. 50kbps CBR Opus wastes 22Mbytes of data in an hour if you're silent compared to a good VBR codec. A small price to pay for improved security. But a CBR Netflix encoding might move several gigabytes per hour of extra data compared to a visually similar VBR encoding, and that's going to really hurt both Netflix and quota restricted customers.
The TLS 1.3 on-wire format is weird because it needs to pass rusted-in-place TLS 1.2 only middleboxes. One of the convenient side effects is that you can add one or more bytes of padding "free". Any 1499 byte packet can be transformed into a 1500 byte packet just by adding one byte of padding whereas many naive formats couldn't do that.
So this means you could "right size" packets just before sending, so that you still send the same number of packets but they're always full, reducing the ability to distinguish their contents by length.
Signal does something more sophisticated to defeat this sort of analysis for its GIF support. The local client does two or more overlapping GETs. So maybe it needs a 14230 byte GIF but it fetches 8430 and 9401 bytes of data then throws some of it away to re-assemble the 14230 bytes needed, an adversary is left with too many possibilities. But it's probably a bit much to expect Netflix subscribers to pay extra for bandwidth just to defeat an adversary who wants to know how much of Tiger King they've seen.
There was a skit I saw a couple decades ago where a person was showing off his radar detector, then in order to combat that, the police had developed a radar detector reflector, so then he had made a radar detector reflector protector or something like that, then they made a protector detector, and so on. It was a couple minutes of explaining his best efforts to counter the police counters to his counters and on and on. But, you know, funny.
The trace busta busta busta.
Creating fast.com always seemed like a pretty brilliant move by Netflix to me.
Almost nothing supports ESNI yet. Chrome does not have it yet. Firefox does but it very difficult to enable, there is a config flag but it does nothing on its own unless you also enable DNS over HTTP in Firefox.
OpenSSL has no support for ESNI yet either.
ESNI also never tells the user if it is working or not yet, making downgrades fairly easy.
ESNI is a long way from being deployed, let alone useful.
Users can test their ESNI support online here: https://www.cloudflare.com/ssl/encrypted-sni/
But, I encourage you to try to get it working in _your_ firefox, where that site says you are in fact using ESNI. It is trickier than just enabling a flag.
If you think this means people are using it, sure. Firefox has ~9% of market share, and a very small percent (maybe 0.1%) of those users have correctly configured this, and when those users visit a page hosted by cloudflare they will may be use ESNI. That is in fact more than zero users. It is a very very small number today, but hopefully it will grow in the future.
And while that page you link tests ESNI, I cannot verify that any other domain supports ESNI. I am unaware of any tooling that lets me test that today.
Only if you use their built-in DoH resolver. The public keys for ESNI are distributed by DNS records and as I understand it there's a bunch of work to be done for Firefox to retrieve these records from classic DNS servers. That work has been classified P5 (we won't do it, but might accept patches)[0].
1. Integrity is redundant with authentication, so really you could say they're ensuring confidentiality + authentication. You can't authenticate a thing without implicitly obtaining assurance of integrity. It's a strictly stronger property.
2. Confidentiality is (usually) insecure and unreliable without authentication. Without authentication you have no PKI for a key exchange to symmetric encryption, so you can't even do TLS in the first place. And if you don't have a carefully applied MAC or a native AEAD mode, your symmetric mode isn't that secure either.
So really what you're asking reduces to the question of why they need the most sophisticated TLS scheme for encrypting their streams. If they want the most secure TLS scheme for confidentiality, TLS 1.3 is the way to do it. They explained one particular facet of why this is the case, re: perfect forward secrecy.
I can't speak for all of Netflix's motivations, but one of the biggest for OTT platforms is preventing MitM of your streams by middle-proxies - including ISPs.
That MitM prevention encompasses data mining/user privacy, ad injection/replacement, and maintaining control over stream quality (Quality of Experience).
There used to be [and still is, in some places] a lot of transcode-to-lower-bitrate-and-cache-inside-our-network behaviour, especially amongst mobile carriers. Reduced stream quality would reflect on Netflix, not the ISP who might be doing this transparently. A user switching up to 1080p or expecting 4K HDR (as appropriate) wants to get that, and Netflix [or others] want to deliver that experience as intended.
The reason we see a LOT less of this now is due to how easy (and thus, widespsread) TLS became.
There are only a few tens of thousands of shows on Netflix, so it ought to be easy to identify.
One way to defeat it, that is already happening to some extent (though not for this reason), is buffering. Maintain a large enough local buffer, and fill it at a constant bitrate, rather than runrate, and you lose this attack vector, no?
Adapting the research to other non-CBR streams doesn't sound far-fetched.
However, I strongly suspect that Netflix's video and audio are sent in different streams. Occasionally video for some titles is missing but audio gets through, confusing our daughter to no end. So while you can't infer individual syllables from the audio stream (as you would from VoIP), the audio streams should have varying enough size and chunking characteristics that allow to identify them.
Interesting data point on the separate AV streams. I can't say I've noticed that myself.
Presumably it depends on the device, but I vaguely recall testing this out on a Chromecast, and I found Netflix to buffer up for several minutes. Some other streaming providers, such as Rakuten, only buffered around 15 seconds.
> I strongly suspect that Netflix's video and audio are sent in different streams
I think this might be right. They presumably want to support lots of different combinations of audio codecs and video codecs, so keeping them separate would make sense. Pure guesswork on my part though.
The encrypted media extensions (EME) are only available on secure contexts, which means the page needs to be HTTPS. Segments are typically fetched with standard XHR/fetch, which means they need to use HTTPS on secure origins as well. You really don't want to have to deal with mixed content.
Unrelated to DRM, Apple's ATS on iOS kinda requires developers to only use HTTPS origins in their apps.
At that point, you might as well just do HTTPS everywhere.
I didn't know that, but you're right. [0] There was a time Netflix didn't use HTTPS. [1] I guess they must not have supported HTML5 at that time, and must have been using Silverlight for in-browser viewing.
[0] https://www.w3.org/TR/encrypted-media/#privacy-secureorigin
[1] https://arstechnica.com/information-technology/2015/04/it-wa...
Joking aside, Netflix's web page runs on a vast variety of devices from phones to smart TV's which don't have the same security profiles. You don't want someone to be able to inject a packet to your Netflix stream to pwn your TV.
Does it work at all though? I does prevent paying customers from getting higher quality streams. But aren't all the shows still widely available in torrents at whatever quality you want?
I definitely have had to enter payment details on Netflix.com before
Pervasive Monitoring Is an Attack
> From the field test, we are confident that TLS 1.3 provides us a better streaming experience.
so they actually measured an improvement in performance metrics.
Or do you just not think that more security is a good thing generally and should be done for its own sake?
I'm all for HTTPS everywhere, but for Netflix and YouTube, that means encrypting exabytes of data. That's going to cost them quite a bit. Seems reasonable to go for more detail than It's good practice in general.
> Are you hinting at something here
No, as I put in my first comment, I think we know the specific reasons (untrustworthy ISPs), it's just interesting they didn't mention them.
[0] https://whydoesaptnotusehttps.com/
[1] https://developer.valvesoftware.com/wiki/SteamPipe#Server_ad...
It seems like the benefits of transmitting your streaming securely (e.g. defeating interloping ISPs) more than exceed the costs, hence why they're doing it.
When people whine about the overheads of HTTPS, they're generally talking nonsense, as the overheads are typically marginal. Video streaming providers are the exception.
https://arstechnica.com/information-technology/2015/04/it-wa...
apt used to be 100% HTTP based. I believe they've changed course recently though. Unsure about Steam.
Steam were clear that the reason they adopted insecure HTTP (previously they used a proprietary protocol) was to enable web caching, by ISPs and other organisations. I can imagine that being very beneficial at a commercial LAN party, for instance.
Related reading: https://wiki.debian.org/SecureApt
I run an out-of-band network tap and have seen only TLS 1.3 traffic for some time now.