The first messenger without user IDs
simplex.chat
simplex.chat
Simplex does something similar. A connects to B. B connects to C. And A and C connect. They all chat. but there is no way to know A, B, or C, because from the outside, it all looks like: X connects to Y, X connects to Y, X connects to Y. So who spoke to whom?
This is great. Even if "the authorities" demand access to chat logs, first, they won't know what to ask for. Chats between whom? Second, they still won't know who spoke to whom even if they have all the data. It's anonymized chats. They would have to sift through all of it.
It still won't prevent someone invading privacy if they have physical access to your device, since the identities are stored locally for your usage convenience.
does it mean if a lot of Simplex users band together and sift through all their local identities they can connect the dots?
In phone A, you will have B's contact stored locally, let's say as "Dan".
B, A, "Pedro"
C, D, "Dan" (yes D is also named "Dan")
D, C, "Olga"
What do you look for?
Everyone can see and read the cipher text on all the papers, but each of the 4 people can only decode the things meant for them.
So, if that's how it works, you could certainly learn who was talking to who if you had access to all the devices. But access to one device only shows you what came and went to the device, but no data about which of the other three users were involved in those reads/writes. You would have to gain access to each device, in turn, to prove whether it was in contact with the first device.
I specifically said a lot of users. A lot is the opposite of one. Imagine a lot (>40% of total users) of impostor devices acting in accord to deanonymize some of the X and Y. Is it vulnerable to that. Like apparently Tor is.
You would at most be able to deanonymize a certain percentage within the impostor network itself. Kind of pointless.
Physical access of devices is the only way to have some chance of deanonymizing some of the users (always less than number of devices you have access to).
That’s my understanding, but the maker of this thing is here, and maybe can respond better?
What makes it work afaict is the combination of:
- there are still queue (inbox) IDs
- key (and (just initial?) queue ID) exchange out of band
https://github.com/simplex-chat/simplexmq/blob/stable/protoc...
So messages are still delivered to an identifier, it's just that every user has tonnes of identifiers (per contact/group), there's no server tracking and handling their exchange, and possibly they rotate via encrypted messages once established anyway.
Exchanging out of band gets you the secrecy, and having one per-chat protects you from a contact turning out bad/leaking/compromised - it's fine that they have metadata about their own chat with you, because they have that & the plaintext anyway.
An identifier is something that relates to more than one thing. A connection is its two endpoints. It is those ends.
Who is at each end is unknown and cannot be known without resorting to grabbing all users devices.
Having not used this chat I don't know how easy thing might be but I do remember, before mobile phones were a thing, being able to remember at least 8 phone numbers that I used to call regularly. Certainly if it called for it you could do this with simplex?
These super-anonymous communication technologies are touted time and again to solve the problem of a surveillance state, while they do nothing of the sort. You cannot solve a social problem with technology.
But if a tool can be devised that no trace of its use can be found, there’s nothing such state can charge you of.
This tool clearly wants to be a step in that direction.
I'm a little put off by parts of their advertising though. Their homepage states Signal can be MITMed given the "operator's servers are compromised", which I don't doubt is true to the extent of the actor stealing, maybe, phone numbers and metadata? But my understanding of Signal's protocol is that a compromised server couldn't intercept _messages_, which is what I feel they're implying.
According to GitHub, their (mono?)repo[1] is split 33% between Haskell, Kotlin, and Swift, which is nice.
What direction does it appear to be heading?
Companies generally trying to do the right thing (e.g. Signal and Mozilla) are held to impossible standards. They are by no means perfect but they're leagues better than the alternatives.
You can't brag about your open source server code and then stop updating it for a year.
You can't spend all day on twitter writing bullshit to attract a certain political demographic and then add a crypto wallet to a chat client.
etc etc
The tech community is seriously guilty of letting perfect be the enemy of good.
For the record, I'm not a huge fan of the crypto wallet thing at all but I just don't use it.
You just used "closed server source code" as an argument against telegram? I also agree that it doesn't matter at all. But then why trashtalk telegram?
Why not, if it solves payments between Signal users (believe the Signal-MobileCoin integration, disregarding Moxie's involvement in both for a second, is a direct response to the now-defunct Libra/Deim)? Personally, I see MobileCoin / wallets as a genuine path to monetization for the Signal Foundation. Though, it remains to be seen if it is any successful.
It's not a "fatal flaw" - it's just a design choice that renders it useless for many use-cases.
I generally believe that for-profit, venture funded company, will some IP help in non-profit, has much better chances of delivering privacy preserving service.
Initially, in 2020, I thought that SimpleX Chat should be non-profit, but then after a long chat with Joseph Jacks in April 2020, who is evangelising VC investment in decentralized open-source tech for a long time, he both convinced me that 1) a dual model is better both for the users and for the scale of change that can be achieved 2) to make a dive into it - the idea at the time seemed too crazy to do something about it for real.
So here we are, with a for-profit company building a privacy-preserving communication network, that will have more than one provider by design.
>SMP is initialized with an in-person or out-of-band introduction message, where Alice provides Bob with details of a server (including IP, port, and hash of the long-lived offline certificate), a queue ID, and Alice's public key for her receiving queue.
So the identity here is the ridiculously long number formed from "Alice's public key for her receiving queue" combined with all that other stuff, isn't it? In other words, a user ID?
Would it perhaps be more accurate to describe this as a system with separate user IDs for each contact?
If Alice's public key is an ephemeral key used only for a single connection -- as with e.g. HTTPS -- then it can't meaningfully be called a "user ID". Then it really is just a random number negotiated as part of making the connection and thrown away when the connection is terminated.
It could be a per user option with new layers of encryption applied to the log at some slow interval. You would get something like a soft delete that would take n hours to decrypt even if you had the keys. Could gradually make stale conversations harder to decrypt with the oldest messages de hardest.
you could have part of your data stored with each contact in decryptable form and the rest stored with them too but in increasingly unreasonable. If enough contacts preserve your data you can easily decode your logs, with less it takes increasingly more cpu cycles.
nearly every 'professional' email app has this feature.
i've been using it on k9 and aquamail for as long as I can remember.
Better yet, I set my email and messaging apps for manual updates. Collateral benefit: it’s really good for battery autonomy indeed
I liked the concept, IMO the accountlessness makes it much easier to invite people via link and start talking immediately (compared to any other popular messenger where some sort of account is needed). The apps’ UI was clean and simple.
The main issue was that one of the user’s Android app would not be able to get timely notifications no matter what we tried. My best guess is that it was an over-eager OS killing background jobs according to a whitelist (Conversations worked fine, for example). iOS app worked really well though (I know it has to use the push notification crutch). Another issue that bit me was that in order to preserve connections between users, server has to be configured to create a form of a backup, and has to be shutdown gently, not with SIGTERM. So a power outage (or carelessness) could sever existing connections.
The home page says that everything is private etc etc. Whatsapp also made similar claims before they were acquired, then suddenly these things didn't matter.
If this gets critical mass, what's stopping them from selling out. Bonus if they "accidently" started tracking users between update v4.2 and v5.2. Also, the TOS magically changed overnight.
The sad part is if they didn't sell out, they'll be buried under lawsuits and eventually banned. Just like KATorrents, yt-dl, Internet Archive, Dread Pirate Roberts etc.
1. https://www.businessinsider.com/suspected-kickass-torrent-ar...
The transmit to every node with fake data all the time approach quickly becomes unsustainable for large groups of >100 people so even if you had a public user ID it wouldn't be seen by that many people anyway. But if they solve this feature disparity without IDs or bloat it will be very interesting.
Largest group I have on WhatsApp has a few dozen people, and after some invisible threshold being passed, activity in the group drops, as people are hesitant to chat with so many unknown people.
The base message queue protocol chooses sensible cryptography, but leaves some important security aspects under-specified. This can lead to mistakes when implementers do not fully understand the security implications. In the context of an open ecosystem, being under-specified here also opens up more risk of implementers causing vendor lock-in (deliberately or otherwise) by breaking interoperability.
For instance, "servers have long-lived, self-signed, offline certificates whose hash is pre-shared with clients over secure channels". Exactly how that pre-sharing is performed is left unprescribed, although SimpleX proposes that clients could introduce other clients to new servers' addresses and public keys. Key distribution is always a challenging problem in cryptography, but traditional x509 PKI is not generally considered to be an issue, so I do not understand what additional benefit to privacy would be found when distributing the keys using the SimpleX protocol itself.
The overview goes on to explain how all client-server communication takes place using blocks of data fixed at 16KiB to make it more difficult for passive observers like ISPs to collude and work out which two parties are communicating. This protection is just taking advantage of statistics - the larger a block is, the more possible messages it could contain, and so the sender is more ambiguous. The downside, though, is that it's very wasteful: assuming that you need to refresh at least once per second for a text chat to feel responsive, that would entail sending a minimum of just under 1MiB of data each way every minute, which is approaching the bandwidth required for a two-way VoIP call! Such an extreme level of resistance against active attempts at monitoring is not usually required by most people, so it would seem more sensible to me to merely ensure that SimpleX is compatible with a lower-level private protocol like TOR for when that protection is necessary.
Moving on to the overview of the chat part of the protocol, I see a fairly uninteresting JSON-based format. There is nothing immediately wrong about it, but it's hardly state-of-the-art either: there's no CBOR to reduce overhead, no JSON-LD to improve extensibility, no MIME types to account for different types of attachment.
From a long-term community perspective, the seemingly arbitrary choice of sub-protocols within the chat protocol (currently group chats, file sharing, contacts and WebRTC calls) makes me hope that SimpleX have a plan in place to avoid what has happened to Matrix. Matrix has numerous built-in features (some of which are in a rather half-baked state), often with a tight coupling between the protocol and the expected user interface for the features, making Matrix notoriously difficult to implement.
Ultimately, I fully support what SimpleX is trying to do: remove the dependence on long-term, globally-unique user IDs in internet communication, which at the moment hampers even federated applications like those that use ActivityPub or Matrix. However, I get the sense from reading the introductory documents that those behind SimpleX aren't aware of (or worse, don't care about) existing standards upon which they could build, and so the novel features are obscured by a lot of rather pedestrian protocol definition that has to be implemented. I fear that SimpleX won't be able to achieve mainstream success unless it fits better into an existing ecosystem of protocols, and whilst I wish them good luck, I must admit that I'd be putting my money on W3C's Decentralized Identifiers (DIDs) to liberate us from centrally-managed online identities.
Do you have examples of protocols that have JSON-LD extensibility mechanism that is embraced by the developer ecosystem to create interoperable extensions. I.e. where the mechanism works more than in theory. I know ActivityPub where JSON-LD is actively shunned, because devs hate it. And Solid project where the specs around using it appropriately get ever more complex. Heard that DID/VC has another extension mechanism fleshed out, but don't know about uptake.
Not sure what you mean by underspecified - it is specified to the level of wire encodings. Possibly you looked at the wrong doc?
> There is nothing immediately wrong about it, but it's hardly state-of-the-art either: there's no CBOR to reduce overhead, no JSON-LD to improve extensibility, no MIME types to account for different types of attachment.
We considered all that, and it seems that they all offer a bad value, compared with lower ubiquity. Also given that messages are padded to fixed 16kb size, there is no value in reducing JSON overhead, and files are sent as binary anyway. Being boring where it doesn't matter is good.
> avoid what has happened to Matrix
Messaging clients are hard to implement indeed, and forking the UI is usually easier than rebuilding it. We purposefully don't want to encourage the development of alternative clients too early, before the spec stabilised, to avoid the fragmentation that happened both with XMPP and with Matrix.
what fragmentation are you thinking of with Matrix? to my knowledge, we have zero fragmentation so far. some clients implement more features than others, but we don’t have any classic “my client sends different reactions to yours” or “my client archives messages differently” or “my encryption is incompatible” style problems. otherwise this smells a bit FUDy…
The Matrix spec has many versions and many features. Clients implement and keep up with varying parts of it due to varying reasons usually involving varying amounts of manpower and funding. Same as with XMPP. I don't see the difference.
In other words, I’m defining fragmentation to be incompatible features - not just clients/servers which haven’t yet implemented a given feature (which is inevitable, just like browsers lag behind specced HTML and CSS features)
One way of putting this is that we’ve traded off the risk of fragmentation (but with free-for-all governance) for the risk of more centralised governance by the Matrix.org Foundation, with associated high drama when folks don’t agree with the curation decisions we make in what gets merged into the official spec.
Both are valid approaches with different tradeoffs; I was just trying to flag the confusion upthread accusing Matrix of being fragmented when it really isn’t (to a fault!)
citation needed? If anything Matrix doesn’t have enough coupling between the API and the expected UI for features - making UIs trivial to implement (which is why there are so many unfragmented Matrix clients - in fact, I’m not aware of any fragmentation?), but the lack of UI coupling then makes performance harder.
So ironically on the Matrix side we’ve been busy adding tighter APIs like Sliding Sync (MSC3575) which make more opinionated choices about the UI in exchange for better-than-telegram perf.
On the other hand, a bigger problem than IDs would probably be IP addresses, etc. Random IDs per-conversation would also go a long way.
https://abcnews.go.com/blogs/headlines/2014/05/ex-nsa-chief-...
If you are considering writing a communications tool for humanitarian reasons, you have to consider that people are thrown into holes and forgotten for looking like they might be in cahoots with someone deemed bad by the powers that be (but not necessarily by the world at large)
Somehow I'm also amused by this clip where woman points qr code at webcam.
Btw. there's Session - https://news.ycombinator.com/item?id=28715627
But in reality there is not an equal distribution between these 3 groups. And there is a high probability that the user base is not as limited as in your pseudo factual simplification. (journalists come to mind for example etc. pp)
In my experince with my volunteer job, we sort of do this. Querying a user's profile via the API only needs a username. I assume they're going to actually use the user ID in the next major release, but I highly doubt it.
If so what is the contingency plan? Out of curiosity I'd assume it's not secret anyways?
Could we make a text chat where all (encrypted) message data is public knowledge in a globally shared bitcoin-like ledger?
There are ~3 billion users on the worlds biggest messenger app, sending ~100 billion messages per day. Assume a message averages 15 words. State of the art compression achieves ~0.65 bytes per word [1]. Thats 1 TB of data per day, for all human conversation.
Imagine any phone can send out an encrypted message, and it gets added to the global ledger.
The data stream of all ledger entries could be broadcast to all phones worldwide (LTE networks support broadcast), and your phone just pick out the messages aimed at it.
End result: Even if every device in the network is evil, nobody can ever determine who you are communicating with.
Current e2e systems might hide the contents of your messages, but they still reveal who you are communicating with to intermediate servers. That info can be used to find and arrest your friends if you're a terrorist, protestor, activist, or just a bit gay in a country that doesn't allow it.
No tor-like system can prevent this if an adversary can run a bunch of the nodes.
Although, to be honest, I still don't understand how nodes propagate messages or find each other.