50k entries isn't a whole lot but as people hosting Mastodon servers have found out, things start slowing down when 1000 people transfer those 50k entries at the same time.
50k entries isn't a whole lot but as people hosting Mastodon servers have found out, things start slowing down when 1000 people transfer those 50k entries at the same time.
You're thinking of password hashing functions like argon2, there's no reason for normal signing and verification operations to be intentionally expensive as that's not where the security guarantees come from.
Compared to something like simple a CRC32 checksum to verify that the data was transmitted correctly, these operations will always be complex and more time intensive than you'd prefer them to be.
The OP was going on about storing all the data on-device and uploading it, but regardless of where it’s stored, if a bunch of people have to move, the thundering herd problem, so to speak, will still exist.
Also, signature verification is not slow. I don’t know what AT uses or how good this source from 2020 is [0] nor what machine it ran on, but ed25519 verification takes about 50us for a 32 byte signature verification. That suggests this guys 50k posts will be validated in 2.5s. If we assume just a single server then that’s about 34k users per day. Or put another way, 30 servers to onboard one million people in a day. None of this seems outrageous to me.
[0] https://safenetforum.org/t/ed25519-vs-bls-performance/32613
Verifying those messages will take about a minute of CPU time per user (assuming no impact from cache misses due to threads swapping in and out and processing new data). I think that's quite significant.
But is this the only possible implementation? I suppose for the OP to have a reasonable point (that it’s “a crock of shit”), this would need to be an intractable problem.
1. They are URI's, and while ActivityPub say they should be https URL's, they don't need to be, and could e.g. point at IPFS or similar.
2. JSON-LD signatures are used by Mastodon, and included in the export, and nothing stops another instance from validating those and serving them up with the original URIs in the id from new URLs, as a means of making it clear the server didn't originate them (there'd be a trust issue if the other servers is unable to get hold of the keys because the original server is gone, but no more so than if the new server had simply republished the content, so the "worst case" is to distrust the original id's).
There are some corner cases there, around trusting the identity of the old and new account represents the same user, so I do think a recovery key type scheme would be nice to allow a user to prove the old and new id is the same (if changing id; I also think we could really use decoupling the expectation that a webfinger id is inherently tied to a Mastodon account - you can sort of do that today; nothing stops you from serving up a separate webfinger result and use it as an alias, but there are usability issues to solve there).
The main problem would probably be keeping track of what server to fetch these messages from after a move (or even a second move) to a different server and keeping the metadata attached in sync.