Interesting, I've never heard that rumor before.
> a single write in a 5000 member public group would result in 5000 writes to the user table.
This hasn't been true since we implemented soft deactivation in 2017. You can read about the feature in our documentation here; I think it's a pretty cool design:
https://zulip.readthedocs.io/en/latest/subsystems/sending-me...
The optimization takes advantage of the fact that public groups with 10Ks of members tend to have a lot of totally inactive users, and so if you have a good way to know that only 600 of them have logged in during the last 2 weeks, your server can run like it only has 600 users.
It is true that Zulip writes a tiny UserMessage table row with (user_id, message_id, flags) for every active recipient of a message, we need to do so in order to track the unread state for all of those recipients (as well as related details, like mobile push notifications state). We need to write to that row a second time when the user reads the message.
With modern postgres and SSDs, writing 1000 rows is really cheap, and so part this is extremely fast when sending messages with 1000 online recipients, which isn't a thing real users do constantly anyway (since that's effectively an announcement, not a chat message).
Folks who are curious can read https://chat.zulip.org/#narrow/stream/3-backend/topic/send_m..., which is related optimization work we did this week that made it into this release.
(You could have a bad experience if you use remote storage that rate-limits "IOPS", though, because AWS at least used to count each row as an IOP and would use your full quota for 20s if you marked 20K messages as read with a 1000 UOP/s plan).
No software architect here so I might be completely off the target but, is not read/unread count something that squares perfectly with the "eventually consistent" model? You don't need to write it down to the persistent storage right now as long as the client UI has it, then some kind of cache has it (so mobile can be in sync when refreshed) and then you persist it to the DB.
Especially since that latency is mostly invisible, thanks to local echo: https://zulip.readthedocs.io/en/latest/subsystems/sending-me...
It's also not very high-value to optimize; that database write is 10-20% of the total time to process sending a message to everyone in a large open community like chat.zulip.org (with 18K total users).