Why does APT not use HTTPS?
whydoesaptnotusehttps.com
whydoesaptnotusehttps.com
I have a few problems with this. The short summary of these claims is “APT checks signatures, therefore downloads for APT don’t need to be HTTPS”.
The whole argument relies on the idea that APT is the only client that will ever download content from these hosts. This is however not true. Packages can be manually downloaded from packages.debian.org and they reference the same insecure mirrors. At the very least Debian should make sure that there are a few HTTPS mirrors that they use for the direct download links.
Furthermore Debian also provides ISO downloads over the same HTTP mirrors, which are also not automatically checked. While they can theoretically be checked with PGP signatures it is wishful thinking to assume everyone will do that.
Finally the chapter about CAs and TLS is - sorry - baseless fearmongering. Yeah, there are problems with CAs, but deducing from that that “HTTPS provides little-to-no protection against a targeted attack on your distribution’s mirror network” is, to put it mildly, nonsense. Compromising a CA is not trivial and due to CT it’s almost certain that such an attempt will be uncovered later. The CA ecosystem has improved a lot in recent years, please update your views accordingly.
Hint: The Netherlands is world leader in surveillance of its own citizens and inhabitants.
It's one thing if you sign a bad key for Google.com, publish it in CT logs and then put it up on the public internet - it's quite another if you sign a bad key for midsizecompany.com, keep it out of the CT system and use it only in a targeted attack against non-technical individuals who are unlikely to examine a certificate or use things like Certificate Watch.
With that said, I still believe serving it over HTTPS would be a substantial improvement. Perhaps pin the cert or at least the CA out of the box to prevent such attacks.
There really aren't "so many ways to detect this" - there's about 3: the user examines the certificate, CT logs catch it later and detection in browsers of major changes on the most high profile sites. Anything falling outside of those will almost certainly go unnoticed.
The SCT proves that this particular log saw a signed document with these specific contents, at this specific moment.
Chrome (for a long while now), Safari (announced for early 2019) and Firefox (announced but a bit vague on when) check SCTs for publicly trusted certificates.
The browser can look at the SCT and verify that:
* It was signed by a log this browser trusts
* It matches the contents of the leaf certificate (DNS names, dates, keys, etcetera: Distinguished Encoding means there is only one correct way to write any certificate so there can't be any ambiguity)
* It has an acceptable timestamp (not too old, in some cases not too new)
It can also contemplate the set of SCTs and decide if they meet further criteria e.g. Google requires at least one Google log and at least one non-Google log.
If any of these is wrong, the site doesn't work and an appropriate error message occurs, no user effort is needed or useful here.
We know this works because Google managed to do it to themselves by accident once already, blocking Chrome access to a new Google site for - I think it was several hours - because their dedicated in-house certificate group screwed up and didn't log a new certificate.
There is more to do to defend the system completely:
1. You could compel a log operator to emit an SCT, but then not actually log your certificate. Certificate Transparency can detect this, a browser would need to remember SCTs it has seen and periodically ask a log for a proof which allows it to verify that the log really has included this certificate.
2. You could go further and compel the log to bifurcate, showing some clients a history with your certificate in, and others a parallel history without that certificate. This can only be detected using what is called "gossip" in which observers of the log have a way to discuss their knowledge of the log state and find out systematically if there are inconsistencies.
Once both these things are in place, there's basically no way around just admitting what do you did. Which of course doesn't mean any negative consequences for you, but it does make _deniability_ (much desired by outfits like the NSA and Mossad) hard to achieve if that's something you care about.
(Also, remember most of the apt development had already happened way before free ssl certs became a thing. While saying "Why don't then just use certbot/LetEncrypt is an easy criticism, give them credit for having actually build a GPG sig secured distributed software delivery system years before LetEncrypt existed...)
DNSlytics, DomainTools, W3Advisor and others offer this.
From a breif review, I see two potential issues:
a) the encrypted sni record contains a digest of the public key structure. This digest is transmitted in the clear (as it must be at this phase of the protocol), so a determined attacker could create a database of values for the top N mirror sites.
b) in order to be useful, the private key for the public keys would need to be shared across all servers supporting that hostname. That's not a big deal for a normal deployment, but it's not great for a volunteer mirrors system -- lots of diverse organizations own and operate the individual mirrors and we need to count on all of those to keep it secure. Also, it adds an extra layer of key management, which is an organizational and operational burden.
It's also pointless without DPRIVE. If people can see all your DNS lookups they can guess exactly what you're up to. That's why that Firefox build did both eSNI and DoH
It’s being added, but early stages. Internet draft is here:
APT's use of plain text HTTP (even with GPG) is vulnerable to several attacks outlined in this paper: https://isis.poly.edu/~jcappos/papers/cappos_mirror_ccs_08.p....
Yes, this paper is old, but APT is still vulnerable to most of these attacks. I would advise anyone wanting to use APT to do so only with TLS.
That's a reasonable complaint. I think it would make sense for the individual packages to be signed as well (and checked during install). This way you'd get a warning if you install an untrusted package regardless of the source. I'm not sure why it doesn't work that way.
>Furthermore Debian also provides ISO downloads over the same HTTP mirrors, which are also not automatically checked. While they can theoretically be checked with PGP signatures it is wishful thinking to assume everyone will do that.
I do do that. If you care about such an attack vector why wouldn't you? And if you don't, why should Debian care for you? There are plenty of mirrors for Debian installers, I often get the mine from bittorent, trusting a PGP sig makes much more sense than relying on HTTPS for that IMO.
>Finally the chapter about CAs and TLS is - sorry - baseless fearmongering.
I don't think it's baseless (we have a long list of shady CAs and I'm sure may government agencies can easily generate forged certificates) but it's rather off-topic. Their main argument is that the trust model of HTTPs doesn't make sense for APT, if that's true then whether or not HTTPS is potentially hackable is irrelevant.
Those who don’t go out of their way to defend themselves don’t deserve security — seriously, that’s your attitude? Then you don’t deserve to be in any position to make security decisions for other people.
Security should be as automatic as possible. It should be assumed that any step that requires manual intervention will be skipped by most people.
Indeed.
1. If you don't care about security, it still doesn't hurt to have HTTPS. Think of it as "extra" that you get for free.
2. If you care about security, you might still don't have the know-how to make sure everything is secure and don't have time to get into it as you're trying to get things done.
3. Even if you care about security AND have the know-how, you might still forget. Nobody's perfect. So it's good that the HTTPS is there.
If HTTPS could be used to replace PGP signature checks then I'd agree with you but it's not the case. So I go back to my initial point, if you worry about your image being tampered with HTTPS is not enough. If you don't care then you don't care either way.
In a way not using HTTP is kind of an implicit disclaimer on Debian's part. "Don't trust what you get from this website". If they feel like they can't guarantee the security of whatever server is hosting the CD images adding HTTPS might actually be a bad thing because people who might otherwise have checked the signature may think "well, it's over HTTPS, it's good enough".
It's not their responsibility to automate security for people using their repos via different tools.
If the solution was just "install certbot on the server and use a free https cert" then perhaps you could make an argument saying maybe they should just do it. But when the problem space includes aggressively using a global (largely volunteer) mirror network and supporting local caching proxies, I can completely understand why they'd say "Nope. Not our problem, not our responsibility to provide a solution. We've got other more productive ways to spend our and our mirror volunteers time and effort".
Actually, debs are not signed in Debian or Ubuntu. The accepted practice in Debian is that only repository metadata is signed.
The argument is that the design of Debian packages (as in, the package format) makes it difficult to reproducibly validate a deb, add a signature, strip the signature, and validate it again.
Personally, I'm not sure I buy it, as we don't have problems signing both RPMs and RPM repository metadata. Technically, yes, the RPM file format is structured differently to make this easier ('rpm' is a type of 'cpio' archive), but Debian packages are fundamentally 'ar' archives with tarballs inside, and those aren't hard to do similar things. For reproducible builds, Koji (Fedora's build system) and OBS (Open Build Service, openSUSE's build system) are able to strip and add signatures in a binary-predictable way for RPMs.
Fedora goes the extra step of shipping checksums via metalink for the metadata files before they are fetched to ensure they weren't tampered before processing. But even with all that, RPMs _are_ signed so that they can be independently verified with purely the usage of 'rpm(8)'.
This is simply not true. Governments can simply compel your CA to do as they want. Not to mention that "uncovered later" is pretty damn worthless.
That said, I do agree that there should be HTTPS mirrors.
It is not worthless as a deterrent to the CA. Proof of a fraudulently issued certificate are grounds to permanently distrust the CA. So, yes, they can do it, but hopefully only once.
There were some topics about it yesterday.
Eg. https://news.ycombinator.com/item?id=18948195
Some arguments of the blog post are also valid here
Here's are some low level gems:
1. I have installed package X. I want to validate that the files that are listed in a manifest for package X have not changed on a host.
APT answer: Handwave! This is not a valid question. If you are asking this question you already lost.
2. I want to have more than one version of a package in a distribution.
APT answer: Handwave! You don't need it. You can just have multiple distributions! It is because of how we sign things - we sign collections!
3. I want to have a complicated policy where some packages are signed, some are not signed, and some are signed with specific keys.
APT answer: Handwave! You should have all or nothing policy! Nothing policy, actually, because we mostly just sign collections, rather than the individual packages
> APT answer: Handwave! You don't need it. You can just have multiple distributions! It is because of how we sign things - we sign collections!
This is not true at all. If you need to distribute multiple versions of the same package then all you need to do is provide multiple versions of the same package. Just get your act straight, learn how to package software, create your packages so that you can deploy them simultaneously without breaking downstream software, and you're set.
This is not apt's limitation, though.
The way Debian works around this is by doing "<name>-<version>" as the package name. This is a valid approach, though it makes package name discovery a bit more difficult at times...
As for multiple versions of packages in the repo, "createrepo"/"createrepo_c" (for RPM repositories) does not care.
And for Debian, "dpkg-scanpackages" can be made to not care too, using the "--multiversion" switch: https://www.mankier.com/1/dpkg-scanpackages#--multiversion
This is somewhat at your peril, as I've observed APT getting confused when it parses metadata from repositories produced by dpkg-scanpackages that allows multiple versions in the repository.
However, reprepro does not support this at all, so most deployments with semi-large Debian repositories will not have this option available to them anyway.
It seems there is a significant semantics gap between how deb packages work and what's supposed to be a package version.
In deb packages, package versions are lexicograhically ordered descriptions of a version ID that is used to guide autoupgrades.
If a packager wishes that multiple minor version releases should be present in a system then he should build his packages to reflect that, which is exactly how Python packages, and specially some libraries, do. For example, python packages are independent at the major version level but GCC packages are independent at the minor version level.
Not in a system. In a repo. Standard debian tools ( not hacked, not outside the tree, not outside the debian main tree ) do not support this.
If you have a repo with <packagename>-<packageversion> and add a package <packagename>-<packageversion1> where packageversion1 is higher than the package version, the previous version gets deleted from the repo.
Can it be worked around? Sure, you can:
a) have multiple repos. If you want to keep up to a hundred versions, you can just create a hundred repos.
b) you can redefine the meaning of a package name and incorporate the version into the name of the package.
(b) Sounds like a good solution. Except if the org decided to use something like .deb for distribution of artifacts it is probably the org that uses other software. Let's say it uses puppet, which supports Debian package management out of the box, except that now you need to change how puppet uses version numbers because as we started embedding the version name into the package name "nginx-1.99.22" and "nginx-1.99.23" became two different packages not two different versions of the same package.
This goes on and on.
That's the semantic gap I've mentioned.
Your statement is patently wrong, as deb packages do enable distributions such as Debian and other Debian-based distros such as Ubuntu to provide multiple major, minor, and even point releases through their official repos to be installed and run side-by-side .
Take, for example, GCC. Debian provides official deb packages for multiple major and minor release versions of GCC, and they are quite able to coexist in the very same system.
All it takes for someone to build deb packages for multiple versions of a software package is to get to know deb, build their packages so that they can coexist, and set their package and version names accordingly.
Through multiple repos. Not through one repo. That's why testing packages live in a separate repo. That's why security packages live in a separate repo. That's why updates live in a separate repo.
If you are using careful versioning on packaging avoiding the <packagename>-<version> as the convention you are breaking other tools, including the tools that are distributed with Debian. Puppets' package { "nginx": ensure => installed } will not work if you decided to allow nginx to have multiple versions using "nginx-1.99" as a special package name.
Wrong. The official debian package repository hosts projects which provide multiple major and even minor versions of the same software package to be installed independently and to coexist side-by-side.
Seriously, you should get to know debian and its packaging system before making any assertion about them. You can simply browse debian's package list and search for software packages such as GCC to quickly acknowledge that your assertion is simply wrong.
I have provided the challenge in another reply:
I have in a repo:
package: nginx version: 1.99.77
I need to add to the repo:
package: nginx version: 1.99.78
Both of the versions must remain in the repo. Packages must be signed. Repo must be signed. The name of the package cannot change and neither can the version number. Both of the packages must be installable using a flag that specifies a version number passed to one of the standard Debian package install tools (the tool must be in "main" collection )
What is a tool that can be used that is listed https://wiki.debian.org/DebianRepository/Setup as existing in the "main" part of the Debian? For half a point, you can use any tool listed in the wiki.
Everything else is the hand waving.
That's true, but multiple versions of the same package is not the same as multiple versions of the same software. Debian and Ubuntu have access to multiple minor version releases of GCC and they can all coexist side by side.
I want to have
nginx-1.99.77 nginx-1.99.78 nginx-1.99.78ehjkhk-1
in the same repo. It is not possible.
<thispackage>-<this-version>-<this-patchlevel>.architecture
The package name is <thispackage>
I need to add to the repo <thispackage>-<this-version>-<that-patchlevel>.architecture
Packages are signed and repo will be signed. At the end I should be able to install the package using the select the specific version flag without adding another repo.
I use aptly and rely on this everyday.
* gopher://jdebp.info:70/1/Repository/debian/dists/stable/
1) It is not listed in the debian wiki
2) To access the "methods" i need to install a gopher client
That's pretty much a definition of the "handwave! Workaround!"
P.S. I have implemented a workaround. It works. It is just not a solution. I should not need to assign people to resign the packages with our own keys and create utilities that would duplicate Debian tools to provide a nominal access to regular required functionality by a large organization that has packages with hundreds of dependencies and could have 5-20 versions of the apps in different production/test/validation/qa environments. Joe Random Engineer expects that his knowledge of how apt-get works, how dpkg works and how puppet works to be portable.
I'm just not going to laboriously re-type it all into Hacker News when you can just go to the published repository and see it all explicitly laid out right there in the repository itself, with the exact scripts and commands that get run to produce the very repository that you are seeing -- quite the opposite of either handwaving or workarounds.
You could get to it with an FTP or an HTTP client, too. But that would just leave you with the raw files in no particular order, rather than the annotated and organized GOPHER listings. A value-added bonus for the GOPHER version, as I said.
You don't actually have any justification at all for your claim that this is somehow impossible with Debian repositories, given that people like me are doing it with a few simple scripts and even publishing them for the world to see; and clearly neither handwaving nor workarounds for anything are required.
Second of all, you still have not provided a link to the doc that someone can read without installing a gopher client. Come on, you said you already have it!
not going to argue with the rest of your flame (others already did) but the answer to this one is
apt install debsumsIn fact, is the FOSS community running it's own trust network still, it used to be a thing at LUGs.
1. You're in China and you download some VPN software over APT. A seemingly innocuous call to package server is now a clear violation of Chinese law.
2. Even in the US, can leak all kinds of information about your work habits, what you're working on, etc.
3. If it's running on a server, it could leak what vulnerable software you have installed or what versions of various packages you're running to make exploiting known vulnerabilities easier
Every time this comes up, it's always the same handful of incorrect arguments made in favor of HTTPS.
The cargo-cult mentality that HTTPS == security really does more damage than it does good.
That doesn't mean it's impossible to determine what packages you downloaded. But it will be more effort to do so.
And the Debian contributor who wrote TFA says it's possible, and I'm sure he knows a lot more about it than I do.
You can select a list of sizes trading off collisions (2 packages with same size => 1 bit of privacy; 4 => 2 bits, etc). But the most you ever get is (nearly) 16. The amount of padding you need to even get 2 bits of privacy (giving up almost 14) on the long tail of large packages is going to be "a lot" and it grows as you want more bits.
Providing privacy for the packages that, including dependencies, are less than 100MB in size is something that's probably worth doing. The cost of padding an apt-get process to the nearest say, 100MB, is not necessarily infeasible as far as bandwidth goes.
Instead of padding individual files, how about a means to arbitrarily download some number of bytes from an infinite stream? That would appear to be sufficient to prevent file size analysis (but probably not timing attacks).
Exposing something like /dev/random via a symbolic link and allowing the apt-get client to close the stream after the total transfer reaches 100MB would appear to make it harder to infer packages based on the transferred bytes, without being very difficult to roll out.
* https://news.ycombinator.com/item?id=18958679
* https://news.ycombinator.com/item?id=18962084
When I was a 19 year old idiot, I was responsible for a mirror server. As a bad actor, I could easily get access to a valid organizational cert.
Me too! It was even ftp.kr.debian.org! It still is!
Seriously, people, who do you think has root of official Debian mirror servers hosted by universities? University students. Who are 19 years old. This is literally true.
However Debian, where APT is from, relies on the goodwill of various universities and companies to host their packages for free. I can see that they don't want to make demands on a service they get for free, when HTTPS isn't even necessary for the use case.
Also since APT and Debian was created in the pre-universal HTTPS days, it does things like map something.debian.org to various mirrors owned by different parties. That makes certificate handling complicated.
You can not get more security by adding a less secure mechanism to a better one. It's not additive.
Analogy: It's like your hanging on a rope and also add a safety net. If the rope breaks, you only fall on the safety net instead of the ground.
All software has bugs, so if you add two buggy security solutions, you might only be able to exploit a bug in one, but the other still gives you the safety.
To continue your analogy: The rope gets tangled in the safety net, forcing you to jump or proceed up with a loose rope, because you can no longer move the rope..
Digital security strength is measured on orders of magnitude, and two mechanisms providing security with very different orders of magnitude do not add in any practical sense.
I am unconvinced about the last part. More commonly an exploit in either will cause security to fail, so adding more steps just adds more attack surface and leads to less security.
Erm, what?
The number of people, who can listen to (much less — modify) your traffic is very small. It is basically your ISP (who is supposed to offers you services in good faith, not spy on you) and a number of engineers, who maintain Internet backbone. That's far from "everyone". Some SSL evangelists make it sound like everyone's traffic is permanently broadcasted to everyone else in the world, but it is not.
As for "vulnerable packages", the most certain sign, that someone does not install security updates, is lack of traffic between them and update servers. But that's orthogonal to use of encryption.
for example, overlay onion routing, size blurring by appending random length random bits,... with oblivious transfer even the APT-server does not know what you downloaded (but that would require a large amount of information..., nevertheless oblivious transfer might still be a useful tool when used as a primitive, perhaps just to send a list of bootstrap addresses for p2p hosting of the signed files etc...)
The HTTPS everywhere movements are attempting to make privacy the default rather than the exception, but are of course done knowing that the server will always know whats up. The point is to make it so that only the two parties concerned, the server and the client, comprehend the communication rather than the entire world
1. I gain access to a router near you.
2. I rent a sizable server at the hetzner/ovh/… location that's closest to you and volunteer to run a mirror for the OS you're using.
Both are somewhat uncertain (your traffic might flow via a different route, you might be load-balanced to a different mirror) but the uncertainty seems comparable. Option 2 seems so much easier that I have a real problem seeing the point of even attempting option 1 if all I want is the information option 2 would give. Perhaps someone can explain?
However, any argument against using encryption for privacy for APT can equally be applied to any other traffic. Do you trust your public internet route enough to let your traffic run authenticated, but unencrypted? Chats, news, bank statements, software updates?
Even if content cannot be modified, it can still be blocked or made public. There are quite a few nosy governments that would like to know or block certain types of content, software packages included.
As for privacy, eh? It is visible you're connecting to a debian mirror and what size of update list you're getting. Barring that, indexing packages by size is trivial.
You want true privacy, you'd have to use Tor or such.
For example, they might want to know what versions you're running (by looking at what updates you _didn't_ download) so they can target you or lots of people at once.
Edit: some report even over 70% that block GA.
In some countries messaging apps like signal and telegram are illegal.
There is no telling what seemingly benign software will be made illegal in the future for political reasons.
Privacy is always a requirement because of these reasons.
Given that the cost of implementation is high and the protection is minimal the decision to not do so is reasonable.
I'm curious how high it actually is. They say it's high, but that could well just be hand-waving. Sure, prior to things like LetsEncrypt those SSL certs would have been a notable financial burden. There's also some extra cost on infrastructure covering the cryptographic workload, but increasingly the processors in servers are capable of handling that without any notable effort.
The real costs are organizational and technical.
Organizing all the different volunteers who are running the mirrors to get certificates installed and updated and configured properly is work. Maybe let's encrypt automation helps here.
From a technical perspective, assuming mirrors get any appreciable traffic, adding https adds significantly to the CPU required to provide service. TLS handshaking is pretty expensive, and it adds to the cost of bulk transfer as well.
I get the feeling that alot of the volunteer mirrors are running on oldish hardware that happens to have a big enough disk and a nice 10G ethernet. I've run a bulk http download service that enabled https, and after that our dual xeon 2690 (v1) systems ran out of CPU instead of out of bandwidth. CPUs newer than 2012 (Sandy Bridge) do better with TLS tasks, but mirrors might not be running a dual CPU system either.
I see no reason to do that for signed packages from the main repositories, however.
But! As mentioned above, outside entities being able to monitor exactly which versions of which packages are being installed to which hosts is a significant security risk.
Assuming that we consider SSH-ing into a server a negligible effort, then adding HTTPS to a APT repository or mirror is also a negligible effort.
As for whether privacy is worth it: Absolutely, especially in this day and age. There is very rarely a cost too high when it comes to privacy, and in this instance, it comes for free.
1) TLS session negotiation leaks all sorts of useful data about both systems, not to mention TCP and IP stack on which it sits. This data is grabbed in 5 minutes with an existing firewall filter. Combined with IP, it shows the exact machine and web browser (incl. Apt version) downloading the file in many cases.
2) It does nothing to prevent time, host and transfer size fingerprinting.
3) Let's Encrypt helps with deployment but you get rotating automated server certificates. It is reasonably easy to obtain a fake Let's Encrypt certificate so without pinning it is worthless for authentication, pinning a rotating certificate is hard too.
Debian does not have resources to handle impostor mirrors.
it would be great to have the ability to have https but for APT in its current form and for what it is used the cost benefit for adding https is not that compelling to me.
In case of HTTP - Step 1: Read the HTTP request payload. Step 2: There is no step 2.
In case of HTTPS - Step 1: Build an index of all possible packages and their sizes. Step 2: Reassemble HTTPS response traffic into individual HTTP responses. Step 3: Look up the response length to the corresponding package. Step 4: In case of identical file sizes, make some sort of model to find out which packet it looks to be based on other packages downloaded (?).
Yes, it's still possible to track people's packages all the same. But you have to have a have way more determined and prepared attacker - it cannot be as easily be done through casual eavesdropping. It's a false equivocation to say it would not add meaningful privacy, as your attacker model changes from casual eavesdroppers to more determined attackers.
You might not care for that particular distinction, and I agree people should have the choice to use HTTP or FTP for APT when selecting mirrors. Unencrypted APT is plenty secure, but encrypted APT is really a little better. In my opinion, there should not be so much resistance for HTTPS in default configurations (e.g. the Debian project could easily require this for their official mirrors around the world). Let's Encrypt makes this so easy, there's no argument anymore in my opinion.
You might want to think about existing surveillance systems. Analysis of telephone traffic is often done purely on the CDR (the caller, callee, and length of call, in simple terms) without the equivalent of deep packet inspection to read the HTTP request, which would be analysis of the actual audio data themselves. The HTTPS case would likewise need just the total octets transferred over the TCP connection for fingerprinting.
There's a lot of glib handwaving in this discussion about identical sizes, not based upon actual measurements of the Debian archive. I quickly looked at the package cache in one of my Debian machines:
jdebp% ls -l|awk 'x[$5]++'
-rw-r--r-- 1 root root 3314 Feb 16 2018 nosh-run-freedesktop-system-bus_1.37_amd64.deb
-rw-r--r-- 1 root root 35190 Dec 14 2016 redo_1.3_amd64.deb
-rw-r--r-- 1 root root 1114546 Feb 25 2018 udev_232-25+deb9u2_amd64.deb
jdebp %
It turns out that in practice size alone almost does uniquely identify package in this sample. The other file that is 35190 bytes is version 1.2 of the same package, leaving just 2 possible ambiguities out of 847 packages. It seems likely that this holds after encryption as well.So the remaining question is how much HTTP pipelining ameliorates this, which no-one here has yet actually analysed.
I just made the distinction between casual eavesdroppers and determined attackers. Those determined attackers exist and are quite capable, I'm sure. I said as much in my post.
You might also want to look into your use of the word 'glib' here. I find it an uncharitable interpretation of my post to call it 'glib' or 'glib handwaving', to be honest. Makes it seem to me as if I should be defending something I said, but I'm not sure what.
But people spying casually on HTTP traffic in general do exist. People able to spy on HTTP traffic in general casually is one of the main reasons we care about HTTPS in the first place. Even though people can do a targeted content length analysis for nearly all other the stuff we read/watch/download online, too. We still care about HTTPS for all of that. And we should probably care for that with APT too, if only a little bit.
ISPs, for example, eavesdrop us all the time, and they do it quite casually. They will modify your unprotected HTTP requests, inject ads, log everything they are able to, and sell the data if they can.
Looking at stream sizes is not "advanced deep packet inspection", it's in fact the opposite of that.
Presenting it as 120 from 43,000 is a bit of an oversimplification, because the average isn't meaningful. The long tail is going to have the worst privacy and the small packages will (probably) have the most.
A scheme like this might be workable but requires being really careful about the security properties you're claiming (i.e., of those 120, probably half are unique, large packages). And obviously, this scheme requires up to 10% additional bandwidth, in the case of the chosen 10% threshold. If buckets change over time, packages moving between buckets may leak a lot of information.
1) You're building a botnet (or, these days, are crypto mining). In that case you're not targeting a specific machine, you just want many of them.
2) You want to exfiltrate information from a host, or sabotage it. In that case you're targeting a specific machine.
I'd argue that in both cases, the proposed attack vector of inferring installed software versions through apt downloads is inferior, or at least more involved. In case 1) you're better off scanning for known vulnerabilities or make use of shodan and the likes. In case 2) you're probably going to probe the server anyways. It might take a little more time than if you just had a complete list of installed packages and their version (given you were somehow able to eavesdrop on the host in the first place), but you'll most likely determine at least what OS is running and what technology stack their internet-facing services are running on after some nmapping.
Or, looking at it from the other perspective, I wouldn't really feel much safer if apt were using https. I'm not against it, but I don't think it's a priority, especially if it needs a lot of coordination between different people, which always turns out to be very time consuming. Just being fast with updating packages seems a better investment of that time.
This is exactly my position, to be fair. We're all bike-shedding here as far as I'm concerned - including this very website. I think the position that HTTPS doesn't help you is a little bit disingenuous, and the only fair position is that "coordinating this stuff takes time and effort we don't feel is worth the negligible advantages" (as you say, and as this website says) is a more acceptable argument than "the negligible advantages don't exist" (as this website seems to also want to say).
B: if the packages are concatenated to each other and a random length noise string, we can substantially frustrate nation state / ISP level attackers to the point of forcing them to get this information from the endpoints themselves: Either end user, or the APT server must already have been compromised.
1) End user not yet compromised: in order to capture these, they must attack the APT server.
2) End user already compromised: on each update of all compromised users, information is sent to C&C, so this would produce lots of opportunity for attentive users to discover the implant.
When focussing on the APT, which would give the cleanest record of attack surfaces, the community can put man power on designing minimalist APT servers, and inspecting published deviations in communications can lead to uncovering 0days.
EDIT: changed disagree to agree, as I (incorrectly) thought you were arguing it would not make attacks more expensive, woops!
We know there are nation states that build profiles of each user based on their HTTP requests, but we don't know of any that have written custom software specifically to target Debian users.
(Would it help if I wrote that code right now and put it on GitHub?)
It would take a software engineer half an hour to write that custom software. It would take a government years to amass the political will to target such a small section of the population, and then potentially hundreds of thousands of dollars for a government contractor to offer a solution and implement it.
There's always going to be a big difference in threat level between a piece of software which already exists and a piece of software that could exist. For example, when you're snatched off the streets by the secret police, and they go to investigate what you've been doing in the country, they might be able to request from HQ a list of HTTP addresses fetched from the IP address associated with your apartment, but they're unlikely to be able to request that HQ write some software to go back retrospectively and count bytes of individual connections you made.
> (Would it help if I wrote that code right now and put it on GitHub?)
No, but it would help if you wrote a patch for APT which made it use HTTP range requests to hide the size of the files it downloads. That should only take half an hour, right?
This sounds less like some sort of massively impossible barrier to overcome and more like a Project Euler problem, and one not all that far into the sequence, either.
One of the things you have to overcome if you want to think like a security person is that, yes, there are attackers that will put some effort into attacking you if you are a target of any consequence, certainly effort far exceeding what you just described. I've watched some people at the company I work for have to overcome that handicap myself. Yes, there are attackers that are not just script kiddies and actually, like, have skills and such.
Attackers won't jump through infinite hoops, but getting a foothold on a network somewhere where they'd like more access, seeing that they can watch a new system in your network getting provisioned, and cross-checking that against a list of known vulnerabilities by looking at package sizes would be boringly mundane for them, not something wildly exotic.
The thing to note here is that the only reason this seems easy is that there is tooling readily available for such a task. If you didn't have such tooling, you'd find it more difficult to implement then your HTTPS case even _with_ tooling.
The same principle applies to your HTTPS case. Your argument disappears as soon as there is tooling. That tooling only needs to be written once. Perhaps it already has been done and exists in the circles where people want to surveil apt users. One possible reason such tooling isn't widely available is that apt doesn't use HTTPS by default, and one outcome may be that if apt switched to HTTPS the tooling would appear.
I have half a mind to write the tooling and publish it just to eliminate this argument. It really isn't very difficult.
The argument that https obfuscates which packages you download is not a good one, and may cause users to unnecessarily worry about the implications (and conversely, that they end up "more safe" if that was not the case). If that type of privacy is desirable you should probably use something like Tor.
It would be more fruitful to discuss the different ways an attacker might deduce what software was installed from a naive implementation: download sizes, download date (i.e. new update available for package P, then a substantial fraction of users who were downloading from the server that day were probably installing P etc...
In theory an onion router might substantially improve the situation if the attacker has a hard time identifying which server the user is talking to, and thus making it hard to identify if the user is even installing anything at all...
sadly I don't trust TOR as long as I can't exclude a specific attack scenario I have always suspected about TOR but never actually known to be present...
Sure, but then you are not talking about "just use HTTPS", you're talking about creating your own protocol and requiring all APT packet sources to speak that protocol, requiring a specialized server software, where currently they can just use whatever HTTP server they want. Switching the whole infrastructure and installed base over to that would be a massive multi-year project, not just a handful code and configuration changes.
The real discussion is not "blindly use HTTPS, or leave it like it is", for me the real interesting question is: can we design a package distribution system that preserves privacy against nation state level actors? can we virtually force those to attack the ATP servers themselves? could we use oblivious transfer ? could we design a fresh minimalist onion router (as opposed to bloated TOR) for package distribution?
That's a great question. Also, can we do it on top of the current Debian infrastructure (HTTP and everything)? Or do we need to change anything?
I'm pretty sure that is not possible, because the current infrastructure is just plain old HTTP file servers (anything that can sling bits will do) run by whoever fancies being a part of it.
But the question here was why https isn't default in apt, not why tor isn't.
That is certainly an intersting theoretical question, but in practice there is also the question of the costs of something that requires you to change and complicate the whole distributed infrastructure vs. the benefits - are there actually real people who need privacy against "nation state level actors" specifically concerning the Linux packages they install?
[0] https://news.ycombinator.com/item?id=16947652
EDIT: just adding, we also don't know what the cost is of the most efficient privacy presevering distribution method actually is. only when people investigate and try will we find out.
X-Unnecessary-Header-0: [random-string]
X-Unnecessary-Header-1: [other-random-string]
...
This could be enough to insert random-length noise without the need to invent any new protocol. Of course this would only be effective for smaller packages as I assume that headers size is limited and as a result this would significantly change the perceived download size only for smaller packages.other things to keep in mind is the dependency graph, some large packages might be rather unique in total filesize downloaded
If you download exactly one package, it may be easy to deduce which one it was (assuming that the protocol overhead is identical each time, and that changing timestamps and nonces doesn't affect the byte length whatsoever, etc.). If you download more than one at a time, which is common with Debian, then the problem is a whole lot harder.
Wait, why? I don't know of any reason HTTP clients can't pipeline requests.
See, debian does not want to manage specific mirror server features if they don't have to. If they were in that position they'd make their own protocol.
That is, in fact, the main threat everyone should protect against.
Privacy against script kiddies and dickhead ISP is valuable. Not every scenario is a determined attacker targeting you personally.
TM1: attacker does not posses zerodays to installed software
TM2: attacker possesses speccific (perhaps OS, perhaps library, perhaps userland) zerodays, usage of which (including unsuccesful attempts) should be minimized to avoid detection
in TM1: it's ok to use HTTP in the clear, as long as signatures are verified
in TM2: everything should be fetched over encrypted HTTPS, since HTTP would leak information about available attack surface
EDIT: not only would this increase security by not revealing what a user installs (perhaps download some noise as well such that it becomes harder to detect what a user is installing?), it could also improve security by turning the APT servers into honeypots, so that monitoring these can reveal zerodays...
TM4: attacker can impersonate a server using Lets Encrypt certificate and bypass their automated verification, creating a fake mirror or a bunch. (HTTP has same vector.) They can also make DNS fail or reroute.
TM5: attacker has a 0-day against the more complex https client (e.g. curl).
TM6: Attacker fingerprints network connections to given servers by size and os or os + tls fingerprinting data
TM6: we should have an overlay onion router, and I agree that the current complexity is worrying, I'd love to see a minimalist version of TOR (with minimal I don't necessarily mean the code size should be small, but minimal assumptions, and that the safety of the system can be verified from the assumptions)
TM4: I don't understand, Lets Encrypt does not calculate private keys for public keys...
Or it would use the much more modern and more secure Noise, like I think QUIC will end up using through nQUIC:
nQUIC is maybe an interesting approach for some specialist applications but more likely a dead end.
The idea you can replace QUIC with nQUIC is like when Coiners used to show up telling us we're going to be using Bitcoin to buy a morning newspaper. Remember newspapers?
nQUIC doesn't have a way for Bob to prove to Alice that he's Bob beyond "Fortunately Alice already knew that" which is the assumption in that ACM paper. So that's a non-starter for the web.
nQUIC also doesn't have a 0-RTT mode. Noise proponents can say "That's a good thing, 0-RTT is a terrible idea". Maybe so, but you don't have one and TLS does. If society decides it hates 0-RTT modes because they're a terrible idea, we just don't use the TLS 0-RTT mode and nothing is lost. But if as seems far more likely we end up liking how fast it is, Noise can't match that. Doesn't want to.
Noise is a very applicable framework for some problems, and I can see why you might think APT fits but it doesn't.
Adding on to this, nQUIC (and Noise specifically) is significantly better for use-cases where CAs and traditional PKI don't make sense, e.g., p2p, VPN, TOR, IPFS, etc...
I agree that APT is not one of these cases. Currently APT has a root trust set that is disjoint from the OS's root CA set, but they could easily do HTTPS and just explicitly change the root CA set for those connections.
EDIT: from the nQUIC paper:
> In particular nQUIC is not intended for the traditional Web setting where interoperability and cryptographic agility is essential.
On another note, I think it would be helpful to expand some points for other readers:
> nQUIC also doesn't have a 0-RTT mode. Noise proponents can say "That's a good thing, 0-RTT is a terrible idea".
0-RTT is dangerous because of replay attacks. It pushes low-level implementation details up the stack and requires users to be aware of and actively avoid sending non-idempotent messages in the first packet.
> Maybe so, but you don't have one and TLS does. If society decides it hates 0-RTT modes because they're a terrible idea, we just don't use the TLS 0-RTT mode and nothing is lost.
One major point of using Noise protocol is to _simplify_ the encryption and auth layers, remove everything that's not absolutely necessary, and make it hard to fuck up in general. Things like ciphersuite negotiation, x509 certificate parsing and validation, and cryptographic agility have been the source of many many security critical bugs.
From an auditability perspective, Noise wins easily. You can write a compliant Noise implementation in <10k loc, vs. OpenSSL ~400k loc.
> But if as seems far more likely we end up liking how fast it is, Noise can't match that. Doesn't want to.
HTTP is insecure, but faster than HTTPS. Most sites now use HTTPS regardless. 0-RTT is insecure and while it might be OK for browsing HN, removing 0-RTT makes it much harder to fuck up.
you could imagine a situation where https would be optional for APT mirrors. Then the package manager would have a config flag to use any mirror or only https-enabled mirrors (probably enabled by default). This would allow to use https without creating any demands to organizations that host those mirrors - if they can they would enable it, but it would not be required. The https-enabled hosts could also provide plain http for backwards compatibility.
I myself uses HTTPS mirror provided by Amazon aws (https://cdn-aws.deb.debian.org/). I do so because My ISP sometimes forward to it's login page when I browse HTTP URLs. Also, it does sometime include Ads (Yeah, it's really bad, but it does remind me that I'm being watched).
Storm in a teacup.
I'm seriously tempted to start flagging links that point to "bad"/"outrage" bugtracker decisions like this, wide public distribution seems to make things quite a bit worse.
Oh and we haven't even addressed that their "secure signing" doesn't also protect first installs that could be insecurely downloaded.
Egypt or Turkey can issue valid fake certificates so you would have to check it if it's not one of those.
As the article says, replay attacks are voided and an adversary could simply work out package downloads from the metadata anyway.
I personally use https out of general paranoia, but understand the arguments for not changing. It's two extra lines in a server setup script.
EDIT: Over 12 years to be precise.
* added apt-transport-https method
Fri, 12 Jan 2007 20:48:07 +0100So in order to get the HTTPS transport, you needed to first download the required package over HTTP.
Since the image was already apt-get install'ing a bunch of other packages at that point and everything seemed to work, the obvious question that popped in my head was: does this mean none of the other packages I've been downloading used https? That's what led me to this website.
My corporate ISP hijacks HTTPS (MITM with self-signed CA), but not HTTP. Any system that uses any HTTPS security properties will verify certificates and fail on my work's network.
The argument about poorly behaved ISPs for one particular protocol but not the other cuts both ways — there are different kinds of poorly behaved ISP.
Moral of the story, I think, is that having a shorter chain of trust is good. In our case, the chain of trust started with a certificate in the original (sometimes OOB) download, the key for which we directly controlled. But for TLS, there are several links in between: the client host's cert store (under the control of OS vendors, hardware vendors a la Superfish, local administrators, etc.), the mess that is the TLS PKI community, your CA, several hundred other CAs, and finally you.
It’s like saying don’t bother having locks on your doors because they are subject to picking... not a great argument.
My counter-counterpoint is that while OpenSSL had (has?) horrible security issues, it's still worth using HTTPS in principle, because a modern internet connected system that has no trustworthy SSL library is going never not going to have security problems. Whether it's hardening OpenSSL, shipping BoringSSL, or anything else, systems just have to get this right, and once they do, applications like apt can take advantage of it.
A "solution" is to protect the endpoint with HTTPS, making MITM attacks impossible. Except that it's also possible that the HTTPS code could suffer from vulnerabilities which can lead to remote code execution. And if I'm being honest, the code which implements HTTPS is much larger and more complicated than the code which is doing the signature checking in APT right now, so by that measure it's actually a downgrade to something "less secure" since it's just adding on more complexity while not improving security much at all.
In reality I believe HTTPS is more heavily scrutinized than the signature verification code in APT is, and therefore could improve security, and there are additional other benefits to HTTPS aside from added security against implementation bugs (like an improvement to secrecy, even if it's small, and better handling by middleboxes which often try to modify HTTP requests but know to not try with HTTPS requests).
An attacker who wants to exploit a buggy pre-auth (or improperly cert-validating) client-side ssl implementation, when the connection is http, can just MITM the http connection and redirect to https.
I don't know enough about it to know either way. I do know that you need to install a package to get HTTPS support for most connections, but I'm not sure if that package is just "switching" to using HTTPS by default or if it actually adds the ability for APT to read HTTPS endpoints.
Is it really not difficult? I bet if you sorted all the ".deb" packages on a mirror by size a lot of them would have a similar or the same size, so you wouldn't be able to tell them apart based on the size of the dialog.
Furthermore, when I update Debian I usually have to download some updates and N number of packages. I don't know if this is now done with a single keep-alive connection. If it is, then figuring out what combination of data was downloaded gets a lot harder.
Finally, this out of hand dismisses a now trivial attack (just sniff URLs being downloaded with tcpdump) by pointing out that a much harder attack is theoretically possible by a really dedicated attacker.
Now if you use Debian your local admin can see you're downloading Tux racer, but they're very unlikely to be dedicated enough to figure out from downloaded https sizes what package you retrieved.
>> "Is it really not difficult? I bet if you sorted all the ".deb" packages on a mirror by size a lot of them would have a similar or the same size, so you wouldn't be able to tell them apart based on the size of the dialog."
Human readable sizes: Sure. Byte size info: Not so much. And even if: Things would become very clear to the attacker after one update cycle for each package.
If you really want to mitigate information about downloaded packages you would have to completely revamp apt to randomize package names and sizes, and also randomize read access on mirrors...
There isn't a need to randomize package names, or randomize read access on the mirror, given fetching deb files from a remote HTTP apt repository is a series of GET requests. Randomizing order of these requests can be done completely on the client side.
Package sizes are still problematic. Here's a suggestion: if each deb file was padded to nearest megabyte, and there was a handful of fixed-size files (say, 1MB, 10MB and 100MB), the apt-get client could request a suitably small number of the padding files with each download. This would improve privacy with a minimum of software changes and bandwidth wastage.
I am fairly confident that this case is not an outlier. Out of the 847 packages currently in the package cache on one of my machines, 621 are less than 0.5MiB in size.
He didn't mean these file sizes specifically. it would still apply just the same with different file sizes
i.e. create cutoffs every 50 or 100kbyte
I am pointing out the consequences of Shasheene's idea as xe explicitly posited it. Xe is free to think about different sizes in turn, but needs to measure and calculate the consequences of whatever size xe then chooses.
No, it would not apply the same with different sizes. Think! This is engineering, and different block sizes make different levels of trade-off. The lower the block size, for example, the fewer packages end up being the same rounded-up size and the easier it is to identify specific packages.
(Hint: One hasn't thought about this properly until one has at least realized that there is a size that Debian packages are already blocked out to, by dint of their being ar archives.)
Transferring apt packages over tor is unlikely to ever become the default, so it's worth trying to improve the non-tor default.
If you update often, the changeset should be small.
It is easy to match in the small set.
Why bet when you can science? :)
$ rsync -r rsync://ftp.be.debian.org/debian/pool/main/ | grep "\.deb$"
$ wc -l debian.txt
1246733 # total number of deb packages
$ cat /tmp/debian.txt | awk '{ print $2 }' | sort | uniq -c | sort -rn | head -1
463 1044 # The most common package size has 463 occurances
$ cat /tmp/debian.txt | awk '{ print $2 }' | sort | uniq -c | sort -rn | awk '{ print $1}' | grep -c "^1$"
259300 # The number of packages with a unique size
$For example, if the download client uses the byte-range HTTP requests to download files in chunks, there is nothing stopping it from randomly requesting some additional bytes from the server. Then the attacker would have a very weak probability estimate of what was actually downloaded.
The post makes clear in the third paragraph: "The easiest way to prevent the attacks covered below is to always serve your APT repository over TLS; no exceptions."
[1]: https://blog.packagecloud.io/eng/2018/02/21/attacks-against-... [2]: https://isis.poly.edu/~jcappos/papers/cappos_mirror_ccs_08.p...
One side: We want performance / caching!
Other side: We want security!
Both sides sometimes argue disingenuously. It's true that caching is harder with layered HTTPS and that performance is worse. It's also true that layering encryption is more secure. (It's what the NSA and CIA do. You smarter than them?)
Personally, I'd default to security because I'm a software dev. If I were a little kid on a shoddy, expensive third world internet connection I'd probably prefer the opposite.
I just wish it were up to me.
It's a relatively insignificant security benefit for most, but could prove an important one for those who targeted attacks are used against.
It might actually become quite difficult to do such analysis, especially if multiple requests were made for packages at once in a connection that's kept open. You won't get direct file sizes either, you'd have to perform analysis for overhead and such - in any case it's significantly less trivial than an HTTP request logger.
That same mechanism means you can easily make trusted mirrors on untrusted hosting.
This is no longer secure then trusting the CA list in the preinstall Windows in pc.
If the whole PKI approach is to work, client has got to get trusting that public key right. In regular practice, that probably means checking it against a HTTPS-delivered version of same from an authoritative domain.
(How far down the rabbit hole do we go? Release managers speaking key hashes into instagram videos while holding up the day's New York Times?)
If someone steals a signing key then they also need to steal the HTTPS cert. Or control the DNS records and generate a new one or switch to HTTP.
Adding an extra layer of encryption is like adding more characters on a password. Sometimes it saves your bacon, sometimes it was useless and came with performance drawbacks.
If you still disagree with me, that's fine. But I want to hear why you continue to hold this opinion when worked for 1Password during Cloudbleed.
https://blog.1password.com/three-layers-of-encryption-keeps-...
Simply put —and this is, I think, where we disagree— the signing of packages is enough. The design of Apt is such that it doesn't matter where you get your package from, it's that it matches an installed signature.
Somebody could steal the key but they would then either need access to the repo or a targeted MitM on you. Network attacks are far from impossible but by the point you're organised to also steal a signing key, hitting somebody with a wrench or a dozen other plans become a much easier vector.
The people that programmed Windows before 2003 probably didn't consider their jobs with the full national security implications.
Then you take something simple, like Linux on a simple IoT device. Say a smart electrical socket. Many of these devices went without updates for years. Doesn't seem all that bad, right? Just turn off a socket or turn it on? How bad could it be?
At some point someone noticed that they were getting targeted and and said: "But why?" The reason is simple. You turn off 100k smart sockets all at once and the change in energy load can blow out certain parts of the grid.
The point isn't that someone will get the key. The point is that we know the network is hostile. We know people lose signing keys. We know people are lazy with updates. From an economics perspective why is non-HTTPS justified? Right? A gig of data downloaded over HTTPS with modern ciphers costs about a penny for most connections in the developed world.
To me, it's worth the cost.
I do think it would cause a non-zero amount of pain to deploy though. Local (eg corporate) networks that expect to transparently cache the packages would need to move to an explicit apt proxy or face massive surge bandwidth requirements, slower updates.
That said, if you can justify the cost, there is absolutely nothing stopping you from hosting your own mirror or proxy accessible via HTTPS.
I'm not against this, I just don't see the network as the problem if somebody steals a signing key. I think there are other —albeit harder to attain— fruits like reproducible builds that offer us better enduring security. And that still doesn't account for the actions of upstream.
I've personally experienced this too: using apt in the presence of a captive portal replaces random bits of `/var/cache/apt` with HTML pages, breaking future updates until you manually find and fix the problem yourself.
Some of these vulnerabilities have the potential for arbitrary code execution, leaving you worse off than the simpler solution based on the verification of cryptographic signatures that has fewer vulnerabilities by virtue of doing less.
The discussion at https://whydoesaptnotusehttps.com is about the protocol. You can add implementation bug risks to the discussion if you want, but then include the risks from both the approaches being discussed.
APT's methodology avoids this and as the current signing and protection mechanisms are file based, the worst case scenario is introducing a new file with a new cryptographic signature along side the old schema, to support still updating a system running old security mechanism.
In comparison, trying to run multiple HTTPS servers with different configurations for specific versions of the system being updated would be a significant engineering effort, especially for mirrors.
This is what many mirrors already do:
>To mitigate this problem, APT archives includes a timestamp after which all the files are considered stale[4].
Let's take a look at the repo spec then:
https://wiki.debian.org/DebianRepository/Format#Date.2C_Vali...
> The Valid-Until field may specify at which time the Release file should be considered expired by the client. Client behaviour on expired Release files is unspecified.
“Should”, “may”, and unspecified behaviour.
An MitM could selectively block certain package being installed / update. Imagine using this to prevent: Bitcoin being installed / enforce a ban on crypto without backdoors / block torrent installations.
This doesn't work as well with the 'recognize package size' method because you need to download the entire package before you know the size. Given the need for Ack in TCP, an MitM can't just buffer data until they have the entire package size.
All they have to do is corrupt the final packet and the package checksum fails. An attacker only needs to buffer a single packet worth of data.
I say this because I've corresponded with such advocates about a completely common case for SSL-- setting up a LetsEncrypt certicate, say. The response I often get doesn't make any sense unless I assume they read a page like this and remembered the feels while forgetting all the relevant details that separate apt from their common case.
A man-in-the-middle attack could simply work by serving you a signed, but outdated packages list, preventing your distribution from updating and leaving you vulnerable to security holes. It's the same attack an evil mirror could do as well.
So if you want to be really sure you should probably use two independent mirrors over an HTTPS connection.
The time stamp is described here[1], but it is not clear how the expiration date is decided.
1: https://wiki.debian.org/DebianRepository/Format#Date.2C_Vali...
Packages in debian and derivatives are not signed. Instead, the manifest that lists all the available packages and their checksums is signed. That's also where the expiration data is stored.
Scroll to the end for a very simple how-to.
It is a dangerous mistake to decide what kind of privacy people need. Privacy should be absolute and without conditions.
What if you live in Iran? Some Ubuntu packages are already inaccessible due to government's pornography keywords censorship. E.g. I can't download "libjs-hooker" from this http link http://archive.ubuntu.com/ubuntu/pool/universe/n/node-hooker... from Iran. What if the government decides to censor the "tor" package?
I find it strange to have a site that is just about one thing that is not that important to most people on a custom domain. If there were pages and pages of information then yes this might make sense but there isn't.
Coming soon...
howtotieyourownshoelaces.com
The premise of this article per domain reminds me of 1998 when everyone thought that instead of search engines people would be typing in URLs, e.g. 'yescupofteaplease.com' so URLs like 'pets.com' were seen as goldmines-to-be.
Let's Encrypt, when are you going to revoke placeimg.com's certificate? The site has been pushing Exploit Kit's malicious payloads since Jan 18 2019 via SSL. Many Flash/IE users are getting infected because most firewalls are unable to peer into SSL tunnels signed by you.
(To be fair, Let's Encrypt is not the only cert authority getting abused (Comodo, yes you))
How often is this, practically? If I'm understanding this right, each new timestamp would come only with a package upgrade, meaning the time period is quite a long time indeed, long enough for a replay attack to work. I would argue that there should be a mechanism requiring a signed the-latest-package-is-X message updated at least every day or so.
Edit: it looks like this is actually what's going on. The page wasn't clear, but it is a metadata "Releases" file that is timestamped, not the packages themselves.
Other than that, most should be safe ig.
One annoying thing is that the page name is misleading. apt does use/support https, it's just that Debian chooses for its default mirrors to it be optional.
It might be probably true that APT is a simple service that does not require the full TLS capability. But APT is only prepared for active exploitations. Passive exploitations will effectively compromise the availability, by compromising the integrity in the relatively predictable way. I don't think APT is also prepared for passive exploitations---casual users will be much more prone to them.
[1] https://forums.xfinity.com/t5/Customer-Service/Are-you-aware...
"Max Justicz discovered that APT incorrectly handled certain parameters during redirects. If a remote attacker were able to perform a man-in-the-middle attack, this flaw could potentially be used to install altered packages."
Far easier to do a MITM attack when apt isn't using https by default
Security is best in layers
That seems like a false dilemma imho.
What's preventing APT from splitting up downloads into identically sized chunks of say 4kb?
Nobody said each mirror couldn't have its own certificate and its own domain. The host names of these mirrors are usually obtained from a trusted source.
HTTPS in addition to GPG signatures does no harm. It also means that one would have to compromise two entities, my distro's GPG key and my mirror's certificate.
Also, a smart observer might be able to detect APT traffic by analyzing the timing.
If you want privacy, APT supports Tor and there are mirrors providing onion services:
https://blog.torproject.org/debian-and-tor-services-availabl...
If they were genuinely concerned about privacy, security and surveillance they would have a lot to say about the out of control surveillance economy and its participants including their employers, the security model of browsers, mobile operations systems, javascript, systemic user stalking and hoovering up data.
Yet those stories are mainly driven by people outside the tech ecosystem and within there is mostly silence interspersed with vague economic justifications so there is something truly bizarre about https extremism on the grounds of privacy.
If your packages are already signed why do you need a middleman? That is a better trust model for a distributed Internet than a centralized CA that is not accountable to individuals and can be compromised by power. Here people are neck deep in business models stalking people online 24/7 and tracking their location and some are 'concerned' about protecting the list of packages you download?
That's believable. "There are a thousand hacking at the branches of evil to one who is striking at the root." - Henry David Thoreau. Nice word "hacking".
Edit: Uh oh maybe Thoreau was pro CA.
And in case of a zero-day exploit it would be really handy not having to wait for some global timeout during which I won't see any updates.
I'm sorry for not liking your favorite color and distro. Please deal with it.
And btw they don't digitally sign their package too (they sign separated meta data file having checksum which is not equivalent of embeding signature inside the package and validate it).
Compare that to yum/rpm which use secure https and signed rpm and signed metadata (both the medium and the payload are secured)
They say https would not add privacy in this context because the package size (almost) uniquely determine the package name. Why is this invalid?
And given the scope of the attacker, fingerprinting by size and server is trivial so easily https adds nothing related to anonymity nor security.
You can't block based on package length, because you need to let the entire update through before you know the length. At that point, it's too late to block. Buffering the entire message doesn't work because TCP expects ACKs.
B) in the interest of memory usage, you could not buffer, and send selective acks to the server -- once you decide to allow it, stop blocking the first data packet, and let the client ack that without the sack and let the server retransmit.
c) b, but for network efficiency, actually let the client receive all packets but the first, and sack them itself --- then when you do allow the first packet, the rest of the packets won't need to be retransmitted.
There are still timeout issues with the buffering, but it is a lot weaker defense.
Your individual statements are correct, but they do not add up to valid argument in this case.
Kazakhstan forces their citizens to install government-issued certificate to use SSL. This allows Kazakhstan to track their citizens. Which proves, that a regime can track it's citizens even in presence of SSL encryption. In other words, using SSL/PKI does not inherently prevent tracking by powerful entities. You need to create your own government for that.
It is naive to think, that regimes like egypt/syria/US can't track people, while at the same time being able to exert overwhelming physical force over the exact same people. If you can force someone to hand over encryption keys, you can track them. Different countries do the same thing, everyone just picks their preferred ways: physically controlling Certificate Authorities in case of US, handing over encryption keys in case of Great Britain.
> Compare that to yum/rpm which use secure https and signed rpm and signed metadata
No, using more "secure" technologies does not amount to better security.
So why debian/ubuntu vulnteer to remove this layer? Why doing the equivalent of installing random certs for every gov/isp on every user?
Yes, government can force someone to install it, but it won't use force on every single person.
If pip or npm used gpg signing of packages rather than https, this wouldn't be a big deal. But as it stands, it's a nightmare on various Windows systems, Linux boxen, and Docker images to get the code you need.
However, I'm not sure if the benefit in speed would be measurable for apt.
The design of HTTP/2 helps websites more. It supports push since webpages include resources that a server knows a client is going to need anyways. This avoids request latency for websites. It also allows connection reuse (via multiplexing. However, again this is mostly helpful for websites. Decrease latency by avoiding tcp hand shake. It allows for slightly smaller headers since it's a binary protocol.
However, pretty much none these really help with large file downloads, and other use cases like APT. Even the headers will be dwarfed by the files themselves, and the header information is probably already pretty minimal.
Also HTTP/2 does not require TLS it's just pretty much all the browsers ignore that the standard does not require TLS. However, they are trying to push for more encrypted traffic. A decision I don't really think browsers should be making.
That's a very subjective and circumstantial statement being passed off as fact.
Just because one can list a few scenarios where HTTPS wouldn't prevent a malicious actor from achieving their goals, doesn't mean HTTPS does not increase security/privacy for apt in other situations.
Further, in certain contexts, guessing and proving are not fungible.
A good design always aims to have multiple othorgonal security mechanisms in place.
https://wiki.debian.org/SecureApt#How_to_manually_check_for_...
They can argue and justify why they do it as much as they want, that actually shows a bit of security immaturity on their side.
The only sensible argument against HTTPS seems to be infrastructure cost. This made more sense in the past, but nowadays hardware crypto acceleration (e.g. AES-NI) is commonplace, and certificates can be obtained for free.
Ultimately it's up to mirrors to decide if they're willing to provide HTTPS.
With GPG, I only need to trust Debian to give me Debian.
Is there really a question about security here?
> providing a huge worldwide mirror network available over SSL is [...] a complicated engineering task
> A switch to HTTPS would also mean you could not take advantage of local proxy servers
As it stands right now, apt-cacher-ng cannot work with https sources.
We can trust the Release file because it was signed by Ubuntu. We can trust the Packages file because it has the correct size and checksum found in the Release file. We can trust the package we just downloaded because it is referenced in the Packages file, which is referenced in the Release file, which is signed by Ubuntu.
Some basic package manager principles
I work with APK, DEB, and RPM based package managers and each of them behave very similar. Each repository has a top level file, signed by the repository's maintainer, that includes a list of files found in the repository and their checksums. When your package manager does an update, it looks for this top level file.
For DEB based systems, this is the Release file
For APK based systems, this is the APKINDEX.tar.gz file
For RPM based systems, this is the repodata.xml file
These files are all signed by the repository's gpg key. So the Release file found athttp://us.archive.ubuntu.com/ubuntu/dists/bionic/Release and is signed by Ubuntu and the gpg key is included in your distribution. Let's hope Ubuntu doesn't let their gpg key into the wild. Assuming that Ubuntu's gpg key is safe, this means that the system can verify that the Release file did in fact come from Ubuntu. If you are interested, you can click on the previous link, or navigate to Ubuntu's repository and open up one of their Release files.Release file
In the Release file you'll see a list of files and their checksum. Example:55f3fa01bf4513da9f810307b51d612a 6214952 main/binary-amd64/Packages
9f666ceefac581815e5f3add8b30d3b9 1343916 main/binary-amd64/Packages.gz
706fccb10e613153dc61a1b997685afc 96 main/binary-amd64/Release
9eae32e7c5450794889f9c3272587f5e 1019132 main/binary-amd64/Packages.xz
5dd0ca3d1cbce6d2a74fcc3e1634ac12 96 main/binary-arm64/Release
The left column is the checksum, then the size of the file, and lastly the location of the file. So we can download the files referenced in the Release file and check them for the correct size and checksum. The Packages or Packages.gz file is the one we care about in this example. It contains information about the packages available to the package manager (apt in this case but again, almost all of the package managers behave very similar).
Packages file
Since we know that we can trust the Release file (because we have proven it was signed by Ubuntu's gpg key), we can then proceed to download the contents of the Release file. Let's look at the Packages file specifically as it contains a list of packages, their size, and checksum.
Filename: pool/main/a/accountsservice/accountsservice_0.6.45-1ubuntu1_amd64.deb
Size: 62000
MD5sum: c2cffd1eb66b6392f350b474e583adba
SHA1: 71d89bd380a465397b42ea3031afa53eaf91661a
SHA256: d0b11d1d27fe425bc91ea51fab74ad45e428753796f0392e446e8b2450293255
The Packages file includes a list of packages with information about where the file can be found, the size of the file, and various checksums of the file. If you download a file through commands like apt install and any of these fields are incorrect, apt will throw an error and not add it to the apt database.
It's time to debunk some myths!
Can an attacker send me a fake Release file?
Sure, but apt will throw it out because it's not signed by Ubuntu (or whoever your repository maintainer is like centos, rhel, alpine, etc)
Can an attacker send me an old index from an earlier date that was signed by Ubuntu that has old packages in it with known exploits?
Sure, but apt will throw it out because it will have a date (in the Release file) that is older than what is stored in the apt database. For example, the current bionic main Release file has this date in it: Date: Thu, 26 Apr 2018 23:37:48 UTC So if you supply it with a Release file older than that timestamp, it will throw it out because it is older than what it currently knows about.
I hope this helps clear the air!
Shameless plug. If you are serious about security and not just compliance, check out our Polymorphic Linux repositories. https://polyverse.io/ We provide "scrambled" or "polymorphic" repositories for Alpine, Centos, Fedora, RHEL, and Ubuntu. We use the original source packages provided in the official repositories and build the packages but with memory locations in different places and ROP chains broken.
Installation
Installation is a one line command that installs our repository in your sources.list or repo file. There is no agent or running process installed. It is literally just adding our repository to your installation. The next time you do an `apt install httpd` or `yum install docker` you'll get a polymorphic version of the package from our repository. You can see it in action in your browser with our demo: https://polyverse.io/learn/
What does it do?
Many of the replies in this post referenced an attacker tricking a server into an older version of a package that has a known exploit. We stop this. Even if you are running an old version of a package, with a known exploit, memory based attacks will not work on the scrambled package because the ROP chain has been broken or as we call it "scrambled". So with our packages, you can run older versions of a package and not be effected by the known exploits. This also means that you are protected from zero day attacks just by having our version of the package.
FREE! For individuals and open source organizations you can use our repositories for free. I hope you try it out!
At a minimum, HTTPS prevents leakage of information about your configuration but there are several direct attack vectors listed in other threads. Please stop calling it “hysteria”.
From https://www.beauzee.fr/2017/07/04/videolan-and-https/:
> If you use homosexuality in order to insult/make fun of someone, it is homophobic. You might disagree, but that’s irrelevant. It is. Calling members of the projet on their personal phones and insulting them is probably a bit of an overreaction, don’t you think?
Using a non-encrypted connection means that it's trivial to work out what packages you download. Using a secure connection at least makes it 1 step harder to infer that information.
However, whether packages should be kept private is all-together another question. I argue that OS updates and packages related to that do not need to be kept private, but applications packages do.
It is trivial still in the case of APT. That's exactly my point: people start believing HTTPS will protect them against many attack vectors it doesn't. That's OK for uninformed people to believe in the magic of a padlock in the address bar, but technical folks should really know better. If I have access to your network traffic and intend to see which packages you download by apt, I will do it irrespective of whether you use HTTPS or not.
Also... the firewall at work breaks APT HTTP pipelining. Very very annoying. It would not be able to do so if APT was using HTTPS.
Oh, you are in for a surprise when you find out what middleboxes really do these days.
(And convince them all to take the CPU overhead hit of TLS)
EDIT: Since I can't reply to all the downvoters, I'll add here.. LetsEncrypt does not solve this. http://us.archive.ubuntu.com/ - that goes to likely 10's of different mirrors. Which one will the LetsEncrypt verification call hit?
Source: Using multiple dozens of LE-issued certs without even thinking about it.
The mirrors would just need to install a letsencrypt-compatible client and setup SSL via that.
On any modern CPU since 2011 or so, TLS overhead is below a percent and a decent cheap dedicated box should be able to saturate a 1 Gbps uplink.
A hostname like ftp.us.debian.org resolves to many different mirrors, and may not resolve consistently around the world -- if that's the case, let's encrypt will not be able to verify the hosts through http challenges.
Also, 1gbps is pretty small. I can't find any documentation on traffic, but I'd imagine mirrors in popular places are at least on 10gig.
LE offers other challenges to verify hosts, like DNS verification. DNS verification can be done easily with an external API for mirror owners to hit (most ACME clients offer DNS challenge with the standard update protocol for DNS which can be secure appropriately).
Debian could get the certificates, but getting the certificates was never the issue -- some CA would be happy to issue certificates for little or no cost to help Debian and gain mindshare. Coordination between 3rd party, volunteer mirror owners and the Debian organization is the issue.
What kind of CPU and traffic patterns are you using to hit 10 gbps of TLS protected traffic?
Mainly serving a file directory with apache with files ranging between 100M and 1G in size. Should be easily comparable to Debian or Ubuntu repositories.
Debian relies on an assortment of volunteer servers of unknown size, and they're not dedicated.
Now multiply that by the 1000 or more projects that each mirror syncs content from, all with something different because nothing standard exists.
LetsEncrypt is great, I love it, I have 10's of certs for personal stuff from them. I think they've completely changed the CA landscape, hopefully forever.
However I'll say it again, LetsEncrypt does not solve this problem. That's OK. LetsEncrypt doesn't have to solve every problem with TLS!
With DNS-01 you only own up to the domain you verified. If you verify ftp.de.debian.org then you can't issue certs for de.debian.org or debian.org but you can issue for www.ftp.de.debian.org.
I don't see the issue with that.
Either way - assuming restricting issuance to exactly 1 name is a solved problem.. This:
What? No.
The mirrors would just need to install a letsencrypt-compatible client and setup SSL via that.
is still a far cry from reality thanks to all the other issues.If anything, it's: What? No. Certs are just the tip of the iceberg, even if LetsEncrypt solved that problem neatly (and they don't), you have ignored the massive complexity of the issue, both the technical and organisational issues.
HTTPS as default would have severely reduced the attack surface for this bug.
Based upon https://www.debian.org/mirror/list, it seems they all have pretty much unique hostnames (ftp.<COUNTRY>.debian.org). You can easily get a certificate for that.
In brief: It's not a huge issue to get a certificate.
But to the article's point, the DNS requests and IPs and file sizes would all be largely transparent, and that's probably enough to figure out what's being downloaded.
On the other hand, HTTP/2 could improve throughput and ensure proxies aren't tampering or replaying.
[0]: https://packages.debian.org/sid/apt-transport-https
Edit:
It was VSCode:
https://code.visualstudio.com/docs/setup/linux
So Microsoft provides apt packages over https, but Debian doesn't.
See https://deb.debian.org/ -- it works just fine.
HTTPS might prevent some kind of surveillance, entities (ISPs, governments, any other MITMs) from determining what kind of packages you're installing, and this might be beneficial, but plain old unencrypted HTTP is often faster if these aren't a big concern (they likely aren't in most developed nations and for most occupations). Lack of encryption overhead as well as transparent proxies being able to serve files are huge boons to this.
With plain HTTP it remains possible no matter how much pipelining you do.
If they allocate it for the whole connection, then all they get is a total, which still gives the attacker some information (there's a limited combination of packages that sum up to that number).
That is a really hard problem to solve. See the knapsack problem.
Imagine I downloaded 5 packages. Debian has about 68'000 packages. That means there is a total of 1 septillion combinations.
If you could check 1 trillion package combinations per second, it would only take 20 thousand years to solve (40 thousand worst case).
In practice, attacker either does not care about your packages, in which case hiding that information gains you nothing, or wants to be alerted, when you (or anybody else) install one of few specific packages. Those combinations can be computed in advance and identified in traffic.
Even if you were interested in a few packages, if any additional packages are mixed in or if dependencies are already installed, this problems become a lot harder again.
While HTTPS doesn't make such an "attack" impossible, it makes it very hard and compared to HTTP the attacker cannot inject or replace data (replay attacks are possible with APT on plain HTTP)
That said, pipelining over https is surely possible to, and reduces the risk.
That said, if you're installing security updates automatically, as you should, anyone will know anyway, as there are only about 3-5 possible combinations of updates you'll be downloading on a particular day in one session.
With automatic security updates, the risk of an attacker finding out what packages you have is less valuable considering you are installing the latest patches.
It would be more interesting if that doesn't happen, in which an attacker can learn what you have installed and wait until exploits appear. Automatic updates would negate this attack model.
edit: As I've demonstrated in a sibling comment; even 5 packages is already out of scope as solving which packages they are is a task of millenia. If you use 4 it could possibly be done by throwing a supercomputer at it for a few months.
Not impossible, just marginally harder (if timing attacks can be though of as "hard").
Keeping your specific packages of choice in secret does not buy you anything anyway. The attacker with access to your traffic will always knows, when you perform system updates, which is more important than names of specific packages.
Anyway, I assume there are very few (if any?) scenarios with no way to overcome such bug in broken proxy (yes, after that exhausting investigation, but still) and there are obvious potential benefits in scenario when such middlebox does its work well.
With TLS (as I understand it) those potential benefits as well as potential bugs are just tossed away together (I'd not use therm 'bypassed' here) and every single download is forced to be made along full wire length, what could be pretty nasty in some locations. Eric Meyer recently wrote interesting article [0] on this topic. (I understand it is all quite obvious stuff, but well expressed IMO.)
[0] https://meyerweb.com/eric/thoughts/2018/08/07/securing-sites... [0][HN] https://news.ycombinator.com/item?id=17707187
With TLS, Eve cannot see (as easily) which packages Alice is downloading from Bob's mirror. If she could, Eve could use that information to decide which exploitable applications to target on Alice's machine.
Speaking of "scary things"... Don't they default to decentralized updates in Windows 10? Sounds like simply joining the swarm will disclose anyone, what update packages you download and when.
Does that mean, that Windows 10 uses "less secure" protocol for downloading it's own updates, than for downloading WSL updates with apt?