You can be: It isn't difficult to argue that signal overstates the security properties of their solution.
It's not an incoherent position to say that it's more important to be conservative (or at least accurate) about the claimed properties than it is to provide epsilon more security.
If it were the case that signal claimed that this metadata had no privacy towards them (and the hosts running their systems), but when you dig into the protocol or detailed tech docs you find that they're using SGX to minimize risk-- then I think there would be nothing to fault about the addition of SGX. In that case it would be purely additive security.
Security is always about tradeoffs and compromises. A system that is very, very good, but not perfect and widely used is far preferable to a system that is perfect but used by no one.
[1]: I work around this by running Signal in a work profile with no contacts. So it's... kinda sorta optional, but in practice there are many caveats and complications that make it effectively not. Certainly not when we're claiming it for average users, which is Signal's argument for using it.
I think there’s a difference between what you’re saying, which seems to be that you personally don’t upload your contact metadata to Matrix, and what I asked / the previous commenter was describing, which is the idea that there are messaging platforms which have decided to not support storing this kind of metadata.
Notably, last I looked, Signal does allow users to opt out of SGX metadata storage, though the initial implementation didn’t cleanly allow for that.
What Signal does is 1) require your phone number, 2) automatically upload every phone number in your contacts, and 3) send a push notification to people on Signal that have uploaded your number in the past that you have joined.
It's a world of difference. And a huge breach of implied trust for multiple people I've known, who joined and immediately uninstalled when they found out a bunch of random people were notified.
But again, the point I made originally, and which you replied to, isn’t about what users may or may not do. It’s whether any messaging platforms have resolved to not store metadata, rather than attempting (as Signal has) to only store it securely.
We can debate whether people agree with Signal’s approach to the tech or the UX, but Matrix doesn’t bother; if you give them your metadata, they’re just storing it without any attempt at E2E encryption.
An entirely separate system exists to facilitate contact discovery. You have to explicitly choose to publish an identifier to it. You choose the server you want to publish to. You choose if, which, and when to query such servers. You choose what to query them with.
A server implementation can (in theory) support whatever arbitrary thing it wants to for an identifier. The fact that email addresses are commonly supported is already _significantly_ better than being forced to use a phone number.
Per the top recommendation from matrix.org, I downloaded the Elements app for iOS. When I opened it, the very first thing it did was ask me if it could access my contacts for sharing them with my (as yet unchosen?) identity server.
Then, I registered an account and clicked the “People” tab. I was immediately prompted to enable contact discovery, via a modal called “Contact Discovery”. It specifically calls out finding contacts by email and phone number. There was never an option presented to select my own identity server, or change what info is published.
If “contact discovery” isn’t part of the core offering, it certainly makes a great attempt to look like it is. The only choice I had available to me is to decline sharing my metadata, which is the same option Signal makes available for SGX.
I get that due to the magic of open source and federation, Matrix implementations could technically use random numbers or favorite colors or some other identifier, and I’m sure there are other less recommended clients which let me pick my own identity server or only share certain slices of data. But if we can’t agree that the flow recommended on the website is the one that the vast majority of users are going to follow, there’s not much room for a productive discussion.
You are looking at only a single implementation here. The point is that this is a set of federated protocols that has been specifically designed to provide freedom of implementation in this regard.
> But if we can’t agree that the flow recommended on the website is the one that the vast majority of users are going to follow, there’s not much room for a productive discussion.
I never claimed otherwise? We seem to be talking past each other.
You say you followed the top recommendation provided to you and that the end result was essentially equivalent to Signal. Since you appear to approve of how Signal handles things, that hardly seems like a complaint to me?
The key difference, of course, is that for Signal that's the only option. If the central authority changes it tomorrow, then tough. Matrix, on the other hand, provides you the freedom to select (or build, or patch, or whatever) a client that meets your needs. It's providing a superset of what Signal is offering.
My initial question, earlier up the chain, was in regards to this dilemma: Signal didn’t upload contact metadata to their servers for years, until they implemented SGX to allow them to do so securely. By contrast, every other messaging platform (Matrix included) just wrote up a spec for contact metadata storage that doesn’t address protections from the server operator.
We’ve got back and forth about the upsides of federation and open source clients and such, but as I’ve tried to point out, none of that is relevant to the actual question I asked, or the upstream point in the thread: that Signal’s spec didn’t include contact uploads at all until there was a way they could store them securely, and other message platforms just skipped straight to storing them.
In practice I have no way of preventing Signal from turning off SGX any time they feel like it because they aren't open and federated. They could push a software update and I'd be none the wiser until a security researcher noticed and pointed it out!
I guess it would be nice if Matrix integrated some sort of remote attestation support into the identity server protocol. Maybe they should have done so from the start, maybe not. At this point I don't see their offering as being less secure than Signal's though, just slightly different.
It's also worth noting that SGX (and AMD's competitive offering) has been broken an embarrassing number of times at this point. From my perspective, secure storage by the server is more or less orthogonal to any given spec or implementation. You either provide data to a server as cleartext or ciphertext. The server does with it what it will - you as the user have effectively no control over that.
Oh, and any/every device/client is a first-class citizen which means you can run it on more then one phone at the same time, or use it without a phone at all.
One can have metadata for sensitive internal corporate channels stay on a network they own in whatever country they want while still being able to chat with outside parties on matrix.org or other servers. Several friends host their own servers and more recently matrix p2p is rapidly maturing to dump the need for servers at all for many use cases.
Federated systems allow people like me to choose to host a server for my own family in my own home closet rack so data shared between my family and I never leave our network.
Still others can host matrix as a Tor hidden service where it will be very expensive to even learn who to target.
If all participants are using Tor and not using identifying information like a phone number, then bulk deanonymization becomes very expensive, particularly if many small highly targeted social circles roll their own.
Signal does not give users a choice but to trust one SPOF setup built under a one size fits all threat model that flows all IP metadata to one place which can leak information regardless of any encryption.
I don't want a world where one central party holds all communications metadata in one place under one legal jurisdiction with one proprietary memory isolation technology under one threat model.
Moxie here likes to point out that email demonstrates all the things that can go wrong with federated systems, but if had chosen the popular alternative of letting a single party take over this whole class of communication, we could all still be using AOL.
Internet messaging will outlive us all, and if we advocate everyone lock up their communications with one (even benevolent) dictator, it won't end well when the next dictator is not so benevolent.
See: pretty much every social network in China and Russia that is now under state control in spite of early promises of privacy by founders.
This is, granted, not as easy as it should be, but it is an issue the matrix team is working on improving.
You also don't need to use the identity service at all. It is totally optional for user discovery.
Third party implementations of the identity server already exist too.
Someone could even write their own replacement that uses SGX if they really wanted ;)
They either have your phone number and other contact details in their phone or they don't. They either make good decisions or they don't. You choose how much to trust them and what with. Federation and third party implementations of identity servers for one particular app changes absolutely none of that.
If you don't register your phone number or email address (which last I checked is not part of the default account creation flow) then it really makes no difference. I agree they should be federated (and they have been working on that among many other things) but it's not like every user's identity is centralised.
With a centralized approach like Signal, I have to trust one single provider, i.e. the Signal Foundation running the servers, with my (meta)data. Sure, in an ideal world I'd rather not trust anyone. Fair enough. With a federated network, however, I generally have to trust every single host that any of my contacts has decided to sign up with, as my metadata will necessarily leak to those hosts when I'm communicating with those contacts in question. And while I personally put in a lot of time and effort into making sure I only use software and internet platforms that protect my privacy as best as possible, most of my contacts don't. The NSA would have an easy time setting up a rogue host.
Now imagine that, in addition, there weren't just one single version of the Signal app, but dozens of forks by dozens of developers. The security risks would get even greater and I would now also have to make sure that my friends use the right (secure) version.
Am I the exception here? Does everyone here only have friends who are IT security experts and know how to tell apart trustworthy hosts and app developers from non-trustworthy ones?
On a side note, another reason why I no longer strongly believe in the idea of decentralization is the following: There was a CompSci paper a few years ago that argued that in any human-made (initially) decentralized network/graph (whether digital or analog; whether a graph of company relationships or mankind's social graph) there will eventually appear "supernodes" which have a disproportionally high number of edges/connections to other nodes. (In the same vain, even in federated networks like email, some hosts will become bigger than others and will eventually take over most of the users and traffic.) In short: There is no such thing as full decentralization. Nodes will necessarily "accumulate" and gather around local centers. (If anyone finds/knows that paper: Please let me know!)
I can see why, at first glance, this seems strictly worse because there are more entities to trust, but this is not a static system.
As an analogy, think about Facebook. Let's say that your friends all hate Facebook, but continue to use it, because for them the inconvenience and switching costs of moving to a different social network (perhaps one you are offering to host for them) is higher than the cost of continuing to use Facebook.
With a federated network, not only can you encourage (or demand) your friends to use a host that you approve of, but also, the very fact that people can move from one host to another (and there isn't a single basket containing all the eggs) means that these hosts are less likely to risk their reputation or be attacked in the first place.
> even in federated networks like email, some hosts will become bigger than others and will eventually take over most of the users and traffic
It seems you are arguing both that federated systems lead to there being too many nodes to trust, and also that in practice only a few big nodes will be trusted. I may be over-simplifying your points there, and I do think there are challenges with federated networks (particularly when federation is deliberately broken due to spamming/abuse/politics) but it's worth having a clear SWOT analysis here.
Anyway, my point, as before, is that a network coalescing around a few big nodes is not a problem as long as there is portability of accounts between them. I would rather trust 3 big providers who fear losing users to a more secure platform, than 1 monopolistic provider that's too big to fail.
I think it's worth explicitly spelling out that federation reduces the cost to switch hosts to near zero in many cases. There's no technical reason you can't use a different host on your end per contact in Matrix (for one to one messaging at least). It would be absurd, but you could do it provided that your client supported it.
If it's really needed, federation combined with open source software means that tooling can be adapted to facilitate your particular security and workflow requirements. (Not that you should need it generally, but the freedom is always there if you do.)
In theory, yes. In practice, however, it gets very complicated. (See e.g. Jabber.) You end up with a situation where only experts are able to set up a secure system for themselves, whereas the average user can't even accurately assess how well her/his privacy is currently protected.
It's analogous to open source vs proprietary software. End users shouldn't generally need to modify the software they use. Doing so isn't typically easy or straightforward for the vast majority of people. But if you ever need to do so, the option is there.
> ... I generally have to trust every single host that any of my contacts has decided to sign up with ...
In general, you should only have to trust a particular host with communications going to the specific contact that uses it.
> Am I the exception here? Does everyone here only have friends who are IT security experts ...
No, I think you're just looking at the problem the wrong way.
With a centralized model, you generally have to unconditionally trust the central authority.
A federated system offers much more flexibility. Metadata is likely to be spread piecemeal across multiple hosts and network paths, making it much more difficult for an adversary to analyze in the general case. Instead of a blanket trust decision, you make a per-contact decision based on the nature of the interaction you intend to have with them. If you feel the need, you can even self host and insist that a particular contact register with _your_ server to communicate with you.
If you absolutely can't trust someone to make good decisions with security critical information, it's unlikely Signal can do much to change the threat they pose to you in a real world scenario anyway. At the end of the day you're ultimately choosing how much trust to place in the person you're communicating with regardless of which system you use to do so.
> ... there will eventually appear "supernodes" which have a disproportionally high number of edges/connections to other nodes ...
That's certainly an interesting dynamic but it's hardly an argument against federation. Centralization is literally the worst case in that model (ie a single node and nothing else). In a federated network you have a choice of whether to make use of such supernodes. If it really matters, I can (for example) refuse to communicate with a particular contact regarding some sensitive topic via (say) gmail. A centralized system doesn't provide that option at all.
Signal's goal, however, is to protect the masses and, thus, society as a whole, by protecting as many people's privacy as possible. The things that are at stake here are not the data and the social graph of a handful of individuals, but the data and social graph of society as a whole. (I'm sure I don't have to mention the implications for democracy and social order.)
My perspective, therefore, was that of an average user. The reason I mentioned my personal situation and the fact that I myself spend quite a bit of time on making sure my systems are safe, was to emphasize that if the situation is already overly complicated for me, it will be much worse for the average user.
> In general, you should only have to trust a particular host with communications going to the specific contact that uses it.
I don't think the word "only" is appropriate here. Let's say John Doe has roughly ~1000 contacts. This means that, in the worst case, he would have to research ~1000 hosts and their privacy policies. Now the "accumulation effect" I mentioned previously will reduce that number quite a bit but there are still going to be dozens of different hosts. I doubt we could expect John to look into each and everyone of them before he gets in touch with his contacts. (Note that this gets even worse when people can freely switch between hosts, as suggested e.g. by other replies to my comment, as John will then have to repeatedly do the checking.) Therefore, I think it is reasonable to expect that a significant number (millions, if not billions) of users will end up residing on insecure and untrustworthy hosts and their social graph and metadata – and possibly even their data (see below) – won't be protected at all.
> With a centralized model, you generally have to unconditionally trust the central authority.
In the federated model, John has to do that, too. Sure, he can self-host but how many people are actually going to do that? Put differently, the vast majority of all acts of communication in the network is going to get routed not through self-hosted nodes but through the servers of providers whose trustworthiness is at least questionable. I hope you will agree that, for society as a whole, this exacerbates the trust problem you mentioned.
> A federated system offers much more flexibility. Metadata is likely to be spread piecemeal across multiple hosts and network paths, making it much more difficult for an adversary to analyze in the general case.
A federated system also makes it much easier for adversaries to enter the game, as they don't have to compromise a well-known provider like Signal that is under the close scrutiny of the public. Instead, they can just create new hosts (just like they do in the case of the Tor network). What's worse, in this case they won't just be able to access their user's social graphs but very likely also the content of their messages, as the key discovery problem is usually solved by hosts distributing their users' public keys.
> Signal's goal, however, is to protect the masses and, thus, society as a whole, by protecting as many people's privacy as possible.
Sure, by positioning themselves as the only node, which you are then forced to trust. That hardly seems like an acceptable solution to me.
The chance of a given mainstream node being actively malicious seems unlikely to me. Maybe that's misguided, but at least (as you point out) there will be relatively few big ones to research. Non-mainstream nodes don't pose a significant threat by virtue of having little to no use (and so seeing little to none of your traffic and contact graph).
> ... the vast majority of all acts of communication in the network is going to get routed not through self-hosted nodes ... I hope you will agree that, for society as a whole, this exacerbates the trust problem you mentioned.
Actually, it seems to diminish the problem to me. Instead of everything going through one central authority it's now being split across multiple actors. No single entity has access to the complete picture anymore. Moreover, you as the user have the freedom to avoid such supernodes if you feel the need. Yes, that will likely introduce usability hurdles, but at least you have the freedom to do so (as compared to a strictly centralized model).
> Instead, they can just create new hosts (just like they do in the case of the Tor network).
Doesn't this attack model fail to account for how users go about selecting and using servers? If I join the Mozilla instance to chat with them, that wasn't an arbitrary choice - I selected the instance that the organization I want to communicate with is using. So the other party, which I'm going to communicate with one way or another, is the real threat here.
As far as MITM attacks go, that scenario seems to get a bit out into the weeds cryptographically. I'm not sure how viable a large scale attack would be here - presumably at some point odd traffic patterns would become noticeable? And you still have the issue of a malicious node that actively attacks all traffic needing to somehow manage to grow its user base to a significant degree.