The AT protocol is the most obtuse crock of shit
urbanists.social
urbanists.social
Before I do, let me just say: Bluesky and the AT Proto are in beta. The stuff that seems incomplete or poorly documented is incomplete and poorly documented. Everything has moved enormously faster than we expected it to. We have around 65k users on the beta server right now. We _thought_ that this would be a quiet, stealthy beta for us while we finished the technology and the client. We've instead gotten a ton of attention, and while that's wonderful it means that we're getting kind of bowled over. So I apologize for the things that aren't there yet. I haven't really rested in over a month.
ATProto doesn't use crypto in the coin sense. It uses cryptography. The underlying premise is actually pretty similar to git. Every user runs a data repository where commits to the repository are signed. The data repositories are synced between nodes to exchange data, and interactions are committed as records to the repositories.
The purpose of the data repository is to create a clear assertion of the user's records that can be gossiped and cached across the network. We sign the records so that authenticity can be determined without polling the home server, and we use a repository structure rather than signing individual records so that we can establish whether a record has been deleted (signature revocation).
Repositories are pulled through replication streams. We chose not to push events to home servers because you can easily overwhelm a home server with a lot of burst loads when some content goes viral, which in turn makes self hosting too expensive. If a home server wants to crawl & pull records or repositories it can, and there's a very sensible model for doing so based on its users' social graph. However the general goal is to create a global network that aggregates activity (such as likes) across the entire network, and so we use large scale aggregation services to provide that aggregated firehose. Unless somebody solves federated queries with the sufficient performance then any network that's trying to give a global view is going to need similar large indexes. If you don't want a global view that's fine, then you want a different product experience and you can do that with ATProto. You can also use a different global indexer than the one we provide, same as search engines.
The schema is a well-defined machine language which translates to static types and runtime validation through code generation. It helps us maintain correctness when coordinating across multiple servers that span orgs, and any protocol that doesn't have one is informally speccing its logic across multiple codebases and non-machine-readable specs. The schema helps the system with extensibility and correctness, and if there was something off the shelf that met all our needs we would've used it.
The DID system uses the recovery key to move from one server to another without coordinating with the server (ie because it suddenly disappeared). It supports key rotations and it enables very low friction moves between servers without any loss of past activity or data. That design is why we felt comfortable just defaulting to our hosting service; because we made it easy to switch off after the fact if/when you learn there's a better option. Given that the number one gripe about activitypub's onboarding is server selection, I think we made the right call.
We'll keep writing about what we're doing and I hope we change some minds over time. The team has put a lot of thought into the work, and we really don't want to fight with other projects that have a similar mission.
It's not all roses. I'm not sold on lexicons and xrpc, but that's probably because I am up to my eyeballs in JSON Schema and OpenAPI on a daily basis and my experience and my existing toolkit probably biases me. I think starting with generally-accepted tooling probably would've been a better idea--but, in a vacuum, they're reasonably thought-out, they do address real problems, and I can't shake the feeling that the fine article is spitting mad for the sake of being spitting mad.
While federation isn't there yet, granted, the idea that you can't write code against this is hogwash. There's a crazily thriving ecosystem already from the word jump, ~1K folks in the development Discord and a bunch of tooling being added on top of the platform by independent developers right now.
Calm down. Nobody's taking ActivityPub away from people who like elephants.
Can you elaborate on this?
Correct me if I'm wrong on this though.
Mastodon will already check the local server first if you paste the URL of a post from another server in the Mastodon search bar, so it's already sort-of doing this, but only with the local server.
So you can already do that with ActivityPub. If it becomes a need, people will consider it. There's already been more than one discussion about variations over this.
(EDIT: More so than an improvement on scaling this would be helpful in ensuring there's a mechanism for posts to still be reachable by looking up their original URI after a user moves off a - possibly closing - server, though)
The Fediverse also uses signatures, though they're not always passed on - fixing that (ensuring the JSON-LD signature is always carried with the post) would be nice.
> expected
Because the URI is only _expected_ to be immutable, not required, servers consuming these objects need to consider the case where the expectation is broken.
For example, imagine the serving host has a bug and returns the wrong content for a URI. At scale this is guaranteed to happen. Because it can happen, downstream servers need to consider this case and build infrastructure to periodically revalidate content. This then propagates into the entire system. For example, any caching layer also needs to be aware that the content isn't actually immutable.
With content hashes such a thing is just impossible. The data self-validates. If the hash matches, the data is valid, and it doesn't matter where you got it from. Data can be trivially propagated through the network.
Posts are explicitly not immutable, so they do need to be revalidated, and that's fine.
For a social network immutable content is a bad thing. People want to be able to edit, and delete, for all kinds of legitimate reasons, and while you can't protect yourself against people keeping copies you can at least make the defaults better.
OK that's my point. In the AT protocol design the data backing posts is immutable. This makes sync, and especially caching a lot easier to make correct and robust because you never need to worry about revalidation at any level.
> People want to be able to edit, and delete
Immutable in this context just means the data blocks are immutable. You can still model logically mutable things, and implement edit/delete/whatever. Just like how Git does this.
You need to know that the user expects you to now have a different view. That you're not mutating individual blocks but replacing them has little practical value.
It'd be nice to implement a mechanism that made it easier to validate whole collections of ActivityPub objects in one go, but that just requires adding hashes to collections so you don't need to validate individual objects. Nothing in ActivityPub precludes an implementation from adding an optional mechanism for doing that the same way e.g. Remote storage does (JSON-LD directories equivalent to the JSON-LD collections in ActivityPub, with Etags at collection level required to change if subordinate objects do).
No you don't. Sorry if I'm misunderstanding, but it sounds like maybe you don't have a clear idea of how systems like git work. One of their core advantages is what we're talking about here -- that they make replication so much simpler.
When you pull from a git remote you ask the remote what the root hash is, then you fetch all the chunks reachable from that hash which you don't yet have. If the remote says you need the chunk with hash X, and you have a chunk with hash X, then you have the data. You don't have to worry if it has changed. Once you have all the chunks reachable from the latest head, you have the latest state of the entire repository. That's it.
(I mean simple in the sense of clear/direct/correct, not in the sense of "easy". It's certainly the case that a design based on consuming a stream of change events is a lot less code).
Yes, I know how Merkle trees work, what it allows you to do. In other words you use the hash to validate which blocks are still valid/applicable. Just as I said, you need to revalidate. In this context (a single user updating a collection that has an authoritative location at any given point in time) it effectively just serves a shortcut to to prune the tree of what you need to consider re-retrieving in this context.
It is also exactly why I pointed at RemoteStorage, which models the same thing with a tree of etags, rooted in the current state of a given directory to provide the same shortcut. RemoteStorage does not require them to be hashes from a Merkle tree, as long as they are guaranteed to update if any contained object updates (you could e.g. keep a database of version numbers if you want to, as long as you propagate changes up the tree), but it's easy to model as a Merkle tree. Since RemoteStorage also uses JSON-LD as a means to provide directories of objects, it provides a "ready lift" model for a minimally invasive way of transparently adding it to an ActivityPub implementation in a backwards compatible way.
(In fact, I'm toying with the idea of writing an ActivityPub implementation that also supports RemoteStorage, in which case you'd get that entirely for "free").
> (I mean simple in the sense of clear/direct/correct, not in the sense of "easy". It's certainly the case that a design based on consuming a stream of change events is a lot less code).
That is, if anything, poorly fleshed out in ActivityPub. In effect you want to revalidate incoming changes with the origin server unless you have agreed some (non-standard) authentication method, so really that part could be simplified to a notification that there has been a change. If you layer a merkle like hash on top of the collections you could batch those notifications further. If we ever get to the point where scaling ActivityPub becomes hard, then a combination of those two would be an easy update to add (just add a new activity type that carries a list of actor urls and hashes of the highest root to check for updates).
Does that makes it more difficult to implement the right to be forgotten and block spam and trolls?
It makes life really easy for spammers.
I don't feel we have the correct solution and there is no commercial reason to ge this thing shipped. Now is the time to explore all the possibilities.
Once we have explored the problem space we should graft the best bits together for a final solution, if needed.
I'm not sure I see the value of standardizing on a single protocol. Multiple protocols can access the same data store. Adopting one protocol doesn't preclude other protocols. I believe Developers should adopt all the protocols.
I firmly believe that protocols are developed through vigorous rewrites and aren't nearly as important as the data-stores they provide access to. I would like our data-stores to be stagnant and as required we develop protocols. Figuring out a method to deal with whatever the hosted data-store's chosen protocol is seems correct to me. I just don't see mutual exclusivity. Consider the power of supporting both protocols.
> We sign the records so that authenticity can be determined without polling the home server, and we use a repository structure rather than signing individual records so that we can establish whether a record has been deleted (signature revocation).
Why do you need an entirely separate protocol to do this? Email had this exact same problem, yet was able to build protocols on top of it in order to fix the authenticity problem. This is the issue: instead of using ActivityPub, which is simpler to implement, more generic, and significantly easier for developers to understand, you invented an overly-complex alternative that doesn't work with the rest of the federated Internet.
> The schema is a well-defined machine language which translates to static types and runtime validation through code generation. It helps us maintain correctness when coordinating across multiple servers that span orgs, and any protocol that doesn't have one is informally speccing its logic across multiple codebases and non-machine-readable specs.
OpenAPI specs already exist and do the same job. They support much more tooling and are much easier for developers to understand. There is objectively no reason why you could not have used them, you are literally just making GET and POST requests with XRPC. If you really wanted to you could've used GraphQL.
There are plenty of protocols which do not include machine-readable specs (including TCP, IP, and HTTP) that are incredibly reliable and work just fine. If you make the protocol simple to understand and easy to implement, you really don't need this (watch Simple Made Easy by Rich Hickey).
> The DID system uses the recovery key to move from one server to another without coordinating with the server (ie because it suddenly disappeared).
Why is this necessary? The likelihood of a server just randomly disappearing is incredibly low. There are community standards and things like the Mastodon Server Covenant that make this essentially a non-issue. You're storing all of a user's post history on their own device in the case of an immediate outage. That's equivalent to Gmail storing all of your emails on your device in case you want to immediately pack up and move to another email provider. That is an extremely high cost (I have 55k tweets, that would be a nightmare to host locally) for an outcome that is very unlikely.
> It supports key rotations and it enables very low friction moves between servers without any loss of past activity or data.
This forces community servers to store even more data, data that may not even be relevant or useful. Folks might have gigabytes of attachments and hundreds of thousands of tweets. That is not a fast or easy thing to import if you're hosting a community server. This stacks the decks against community servers.
Most people want some of their content archived, not all, and there is no reason why archival can be separate from where content is posted. Those can be two separate problems.
> That design is why we felt comfortable just defaulting to our hosting service; because we made it easy to switch off after the fact if/when you learn there's a better option. Given that the number one gripe about activitypub's onboarding is server selection, I think we made the right call.
Mastodon is able to do this on top of ActivityPub. Pleroma works with it. Akkoma works with it. There's already a standard for this. Why are you inventing an unnecessary one?
Mastodon also changed their app to use Mastodon.Social as the default server, so this is a non-issue.
theyre tweets, how much could they cost? @ 280 bytes each, that's like 15MB. double it for cryptographic signatures and reply-to metadata. is that really too much to ask for the capacity to transfer to another host at anytime?
(also, leaving aside the fact that 55k tweets puts you in the 0.1% of most prodigious users)
Anybody who has had a smartphone for a decade likely has at minimum 10k photos in their cloud locker with local thumbnail.
A quick search got me to twitter stats from 2013 when people were posting 200 billion tweets per year. Thats 5-6 orders of magnitude more. You don't get a 10000x improvement just by federating and hosting multiple nodes.
I have every email I've ever received or sent (and not deleted) and it's only 4GB.
Should something require I download all that every time I login? No. But having a local copy is amazing, and a truly federated system should have and even be able to depend on those.
The Mastodon Server Covenant is a joke; the only enforcement is to remove the server from the list of signup servers; which if it just fell over dead because the admin died/doesn't care/got arrested/got a job will not matter.
My work email is 10gb.
Seriously. This is setting off all sorts of red flags. I’m old enough not to trust non w3c standards.
Ah, email, where a message of 114 characters with no formatting ends up over 9KB due to authentication and signatures, spam analysis stuff, delivery path information and other metadata. Sigh. Although I doubt this will end up as large as email, the lesson is that metadata can end up surprisingly large.
In this instance, I think 1–2KB is probably more realistic than the half kilobyte of “double it”.
This has actually happened. It's a real problem. For example, "Mastodon instance mstdn.plus with over 4K users suddenly broke" https://lapcatsoftware.com/articles/mastodon.html
As far as I'm concerned, the Mastodon Server Covenant is a joke.
Another example: Mastodon.lol, which had 12,000 users literally shutdown a few hours ago. They did manage to give notice but the point remains that people had to move instances, cannot take their posts with them, and it’s a giant PITA, server covenant or not.
To call this stuff a “non-issue” seems incredibly obtuse, especially when the data portability piece is clearly an after thought by the Mastodon devs, and something that ActivityPub would need some major changes to get accomplished. Changes that the project leads have been fairly against implementing.
However, you are coming across as highly adversarial here. Mostly because you immediately follow your questions with assertions, indicating that your questions may be rhetorical rather than genuine.
I’m not accusing you of anything per say but I very much want a dialog to happen in this space and I think your framing is damaging the chances of that happening.
It is why passersby like me can't get into either Twitter or Mastodon when it is a culture of getting outraged and shouting at each other, to collect a choir of people nodding and agreeing in the replies: "well done for saying it like it is."
These people forgot how humans talk and have arguments outside of their Internet echo chambers.
Anger, insults and hate sells more (creates more engagement) than reasoned arguments. No one would have posted this on HN if is was otherwise. So don't worry, they are not hurting their chances, the next topic that can be summarised with an angry title like "xxx is the most obtuse crock of shit" will get great traction on HN.
The linked article/toot and certain replies seriously makes me want to just shutdown my personal mastodon server and move on from the technology altogether.
"Also I don't care if I'm spreading FUD or if I'm wrong on some of this stuff. I spent an insane amount of time reading the docs and looking at implementation code, moreso than most other people. If I'm getting anything wrong, it's the fault of the Bluesky authors for not having an understandable protocol and for not bothering to document it correctly."
It's unfortunate because there are some valid points in his criticism.
It still does in my circle. The overloading of "crypto", though, has become such a source of confusion and misunderstanding that I have stopped using it and just use the full word, be it cryptography or cryptocurrency, instead.
(Of course it's perfectly possible, for all I know, that SW is not debating in good faith. But what you quote doesn't look to me like an admission of bad faith.)
"I put as much effort in as can reasonably be expected; I tried to evaluate it fairly; but the documentation and supporting code is so bad that I may have made mistakes. If so, blame them for making it impossible to evaluate fairly, not me for falling over their tripwires."
If something is badly documented and badly implemented, then I think it's OK to say "I think this is badly designed" even if you found it incomprehensible enough that you aren't completely certain that some of what looks like bad design is actually bad explanation.
If some of the faults you think you see are in fact "only" bad documentation, then in some sense you're "spreading FUD". But after putting in a certain amount of effort, I think it's reasonable to say: I've tried to understand it, I've done my best, and they've made that unreasonably difficult; any mistakes in my account of what they did are their fault, not mine.
(I should reiterate that I haven't myself looked at the AT protocol or Bluesky's code or anything, and I don't know how much effort SW actually put in or how skilled SW actually is. It is consistent with what I know for SW to be just maliciously or incompetently spreading FUD, and I am not saying that that would be OK. Only that what SW is admitting to -- making a reasonable best effort, and possibly getting things wrong because the protocol is badly documented -- is not a bad thing even when described with the words "I don't care if I'm spreading FUD".)
If your identity is separate from your Gmail account (as it can be with a custom domain, for email and for bluesky), this seems like a very plausible and desirable thing to be able to do. Just recently there was an article about how Gmail is increasing the number of ads in the inbox; for some people that might change the equation of whether Gmail's UX is better than it is bad. If packing up and leaving is low-friction enough, people might do it (and that would also put downward pressure on the provider to not make the experience suck over time)
And that's not even getting into things like censorship, getting auto-banned because you tripped some alarm, hosts deciding they no longer want to host (which has happened to some Mastodon instances), etc.
It happens all the time. mastodon.social, the oldest and biggest Mastodon instance, has filled up with cached ghost profiles of users on dead instances. Last I checked, I could still find my old server in there, which hasn't existed for several years.
But improving on the ActivityPub user migration store is also a minor/trivial change away from doing much better than today: you just need to change ActivityPub Ids to either fully a contentadressable hash or referencing a base that is under user control, plus a revocation key style mechanism for letting the user sign claims about their identity in order to allow unilateral moves.
On the contrary, email has no solution to the authenticity problem that’s being talked about. Even what there is is a right mess and not even slightly how you would choose to build such a thing deliberately.
If you want to verify authenticity via SPF/DKIM/DMARC, you have to query DNS on the sender’s domain name. This works to verify at the time you receive the email, but doesn’t work persistently: in the future those records may have changed (and regular DKIM key rotation is even strongly encouraged and widely practised).
What you are replying to says that AT wants to be able to determine authenticity without polling the home server, and establish whether a record has been deleted. Email has nothing like either of those features.
Which is a risky thing to do, because most people don't associate GPG with positive feelings about well designed solutions, but they're right in that it works well, solves the problem and is built squarely on top of email.
The reason that it's not generally well received is that there's no good social network for distributing the keys, and no popular clients integrate it transparently.
That’s not great. I wonder if the receiver could append a signed message upon receipt with something like “the sender’s identity was valid upon receipt”.
That's exactly what does happen, if you view the raw message in GMail/iCloud, you should see DMARC pass/fail header added by the receiving server (iCloud in your example).
(Well not exactly, it's not signed, but I'm not sure that's necessary? Headers are applied in order, like a wrapper on all the content underneath/already present, so you know in this case it was added by iCloud not GMail, because it's coming after (above) 'message received at x from y' etc.)
I’m also curious how this plays into the original comment about dkim/spf/dmarc not being sufficient due to key rotation still factors into the conversation after having discussed this?
Key rotation would have the same effect as 'DNS rotation' (if you stopped leasing the domain, or changed records) - you might get a different result if you attempted to re-verify later.
I just don't really see it as a problem, you check when you receive the message; why would you check again later? (And generally you 'can't', not as a layman user of GMail or whatever - it's not checked in the client, but the actual receiving server. Once it's received, it delivers the message, doesn't even have it to recheck any more. Perhaps a clearer example: if you use AWS SES to receive, ultimately to an S3 bucket or whatever for your client or application, SES does this check, and then you just have an eml file in S3, there's no 'hey SES take this message back and run your DKIM & virus scan on it again'.)
Everything else aside, this is completely untrue.
I self-hosted my first fediverse account on Mastodon and got fed up with the complexity of it for a single person instance and shut it off one day (2018 or so?).
On another account at some point 50% of my followed people vanished because 2 servers where everyone in that bubble were on just went offline. Took a while to recreate the list manually.
This may be anecdotal but I've seen it happen very often. Also people on small instances blocking mastodon.social for its mod policies comes close to this experience.
Your reply fails to address that push-based systems are prone to overwhelming home servers due to burst loads when content becomes viral. By implementing pull-based federation, the AT Protocol allows for a more balanced and efficient distribution of resources, making self-hosting more affordable and sustainable in the long run.
Servers go down or get flaky all the time for various reasons. Easy relocation (with no loss of content & relationships) and signed content (that remains readable/verifiable even through server bounciness) soften the frustrations.
55k tweets is little challenge to replicate, just like 50k signatures is little challenge to verify, here in the 2020s.
If Mastodon does everything better with a head start, it should have no problem continuing to serve its users, and new ones.
Alas, even just the Mastodon et al community emphasis on extreme limits on visibility & distribution – by personal preferences, by idiosyncratic server-to-server discourse standards, by sysop grudges, whatever – suppress a lot of the 'sizzle' that initially brought people to Twitter.
Bluesky having an even slightly greater tilt towards wider distribution, easier search, and relationships that can outlive server drama may attract some users who'd never be satisfied by Mastodon's twisty little warrens & handcrafted patterns-of-trust.
There's room for multiple approaches, different strokes for different folks.
They have instead created another issue.
Before, it was a usability issue which normal users looking for an alternative social network got confused on the sign up process and gave up.
If that wasn’t an issue why did the Mastodon devs decide to select a default server in the app after seeing this?
Now, they have traded that off and created a centralization issue going against the point of encouraging federation.
This only shows that centralization wins in the end.
> Why is this necessary? The likelihood of a server just randomly disappearing is incredibly low.
The likelihood of a server just randomly disappearing at any point in time is low. The likelihood of said server disappearing altogether, based on the 20+ years of the internet, can & will approach 100% as the decades go on. Most of the websites I know in the early 2000s are defunct now. Heck, I have a few webcomic sites from the 2010s in my bookmarks that are nxdomain'd.
Also, as noted by lapcat, these sudden server disappearances will happen. Marking this problem as a non-issue is not, in any realm of possibility, a good UX decision.
https://news.ycombinator.com/item?id=35883409
This is coupled with the fact that Mastodon (& ActivityPub in general) don't have to do anything when it comes to user migration: The current system in place on Mastodon is completely optional, wherein servers can simply choose to not allow users to migrate.
https://news.ycombinator.com/item?id=35883570
https://news.ycombinator.com/item?id=35884682
> There are community standards and things like the Mastodon Server Covenant that make this essentially a non-issue.
*The Covenant is not enforced in code by Mastodon's system, nor by AcitivtyPub's protocol.* It's heavily reliant on good faith & manual human review, with no system-inherent capabilities to check if the server actually allows user data to be exported.
> You're storing all of a user's post history on their own device in the case of an immediate outage. That's equivalent to Gmail storing all of your emails on your device in case you want to immediately pack up and move to another email provider. That is an extremely high cost (I have 55k tweets, that would be a nightmare to host locally) for an outcome that is very unlikely.
An outcome *that can still happen*. As noted by the incidents linked above, they're happening within the Mastodon platform itself, with many users from those incidents being unable to fully recover their own user data. Assuming that this isn't needed at all is the equivalent of playing with lightning.
No. Just no.
If (IF!) some distributed social network breaks through and hundreds of millions or billions of people are participating, they are going to do things that The Powers That Be don't like. For better or worse, when that happens they will target servers, and servers WILL just disappear. Domains will disappear. Hosting providers will disappear. You can take that straight to the bank and cash it.
Uncoordinated moves are table stakes for a real distributed social network at scale. The fact AT Protocol provides this affordance on day one is a great credit.
But if we started today, we wouldn't build email that way. There are so many baked-in well-intended fuckups in email that reflect a simpler time where the first spam message was met with "wtf is this, go away!" I remember pranking a teacher with a "From: president@whitehouse.gov" spoofed header in the 90s.
Email is the way it is because it can't be changed, not because it shouldn't be.
I literally read about a case over a month ago where some obscure Mastodon-server admin blocked someone's account on their server so it was impossible to move to another instance. The motivation was "I don't want capitalist here, can change my mind for money" (slightly paraphrasing). Basically, it's stupid to use any Mastodon instance other than the few largest one or your own.
That's why BlueSky's approach makes sense.
>with the rest of the federated Internet. You're saying like it's a thing that won and not a niche project for <10M users globally.
In the world I currently live in, all my emails are stored locally on my devices. Also, text files take up little to no storage, so why does it matter?
I’m guessing with more success will come more haters unfortunately. :-/
I’m still waiting for an invite to try it myself but I’m excited to see how it compares to mastodon.
I’m also curious to know if AT Protocol can be used for multiple service types. I’m more interested in a reddit replacement (lemmy) than a Twitter replacement (mastodon) and I’m hoping we can see another full scale fediverse come to fruition.
Then maybe don't invite journalists onto your Quiet Stealthy Beta!!!
I hope you get rest, work is not that important
By complete chance, I happened to interview someone that had, "wrote the debugger for that obscure microcontroller back in the 90's," tucked away in their resume. It was hard not to spend the entire interview session just picking their brain for anecdotes about it.
Fortunately not something ive had to experience!
(speaking for myself, not you, to be clear)
That said, my frustration still gets converted into not-so-polite comments in the source code the culprits will never see.
I want to say that "if there was something off the shelf that met all our needs we would've used it" has been the justification for many over-engineered, not-invented-here projects. On the other hand, in many cases it is completely legitimate.
On my windows PC, I type ‘note’ and slam enter to open notepad a thousand times a day without a problem. On my KDE desktop ‘text’ seems to 50/50 bring up… whatever the text editor is called and 50/50 something else. Apparently “Kate” is what I’m after.
There is real value in naming things after what they do and it’s my sole gripe with KDE that they have stupid names.
I can also see the term becoming controversial such male and female vs. plug and socket. Some see the former as blatantly sexual and the latter requiring dirty mind to be sexual.
Maybe it's best to hold off on that reference until we get more voices to chime in.
Glad it was something else entirely. The localization for Brazil can be rocambole.
It's tough to do in a freestanding way given that it's a command-response protocol. It's very convenient to depend on the specific uart API that's available.
I will nake sure to reuse this name if I ever find myself developing an Obtuse crock.
Yeah, but it's an obtuse crock that you can literally still here in your memories. How many protocols can say that? It has a special place in my heart, I think.
I know Jeremie Miller is on the board, is Telehash or something similar being used within Bluesky?
Also, I'm sure you get this a lot, but I'd love a BlueSky invite please.
=Bill.Barnhill w dot a dot barnhill at gmail.com
As far as I'm aware its developed on Matrix and had some connection with it, so this isn't a case of not knowing about it so there is likely some engineering reasons why this was not done.
Would be interesting to understand.
* Both are work by self-authenticating git-style replication of Merkle trees/DAGs
* Both define strict data schemas for extensible sets of events (Matrix uses JSON schema - https://github.com/matrix-org/matrix-spec/tree/main/data/eve... and OpenAPI; AT uses Lexicons)
* Both use HTTPS for client-server and server-server traffic by default.
* Both are focused on decentralised composable reputation - e.g. https://matrix.org/blog/2020/10/19/combating-abuse-in-matrix... on the Matrix side, or https://paulfrazee.medium.com/the-anti-parler-principles-for... on the bluesky side, etc.
* Both are designed as big-world communication networks. You don't have the server balkanisation that affects ActivityPub.
* Both eschew cryptocurrency systems and incentives.
* Both have names which everyone complains about being hard to google, despite "AT protocol" and "Matrix protocol" or "Matrix.org" being trivial to search for :P
There are some significant differences too:
* Matrix aspires to be the secure communication layer for the open web.
* AT aspires (i think) to be an open decentralised social networking protocol for the internet.
* AT has portable identity by default. We've been working on this on Matrix (e.g. MSC1228 - https://github.com/matrix-org/matrix-spec-proposals/pull/122... and MSC2787 - https://github.com/matrix-org/matrix-spec-proposals/blob/nei...) and have a new MSC (and implementation on Dendrite) in progress right now which combines the best bits of MSC1228 & MSC2787 into something concrete, at last. In fact the proto-MSC is due to emerge today. EDIT: and here it is: https://github.com/matrix-org/matrix-spec-proposals/blob/keg...
* AT is proposing a asymmetrical federation architecture where user data is stored on Personal Data Servers (PDS), but indexing/fan-out/etc is done by Big Graph Servers (BGS). Matrix is symmetrical and by default federates full-mesh between all servers participating in a conversation, which on one hand is arguably better from a self-sovereignty and resilience perspective - but empirically has created headaches where an underpowered server joins some massive public chatroom and then melts. Matrix has improved this by steady optimisation of both protocol and implementation (i.e. adding lazy loading everywhere - e.g. https://matrix-org.github.io/synapse/latest/development/syna...), but formalising an asymmetrical architecture is an interesting different approach :)
* AT is (today) focused on for public conversations (e.g. prioritising big-world search and indexing etc), whereas Matrix focuses both on private and public communication - whether that's public chatrooms with 100K users over 10K servers, or private encrypted group conversations. For instance, one of Matrix's big novelties is decentralised access control without finality (https://matrix.org/blog/2020/06/16/matrix-decomposition-an-i...) in order to enforce access control for private conversations.
* Matrix also provides end-to-end encryption for private conversations by default, today via Double Ratchet (Olm/Megolm) and in the nearish future MLS (https://arewemlsyet.com). We're also starting to work on post quantum crypto.
* Matrix is obviously ~7 years older, and has many more use cases fleshed out - whether that's native VoIP/Video a la Element Call (https://element.io/blog/introducing-native-matrix-voip-with-...) or virtual worlds like Third Room (https://thirdroom.io) or shared whiteboarding (https://github.com/toger5/TheBoard) etc.
* AT's lexicon approach looks to be a more modular to extend the protocol than Matrix's extensible event schemas - in that AT lexicons include both RPC definitions as well as the schemas for the underlying datatypes, whereas in Matrix the OpenAPI evolves separately to the message schemas.
* AT uses IPLD; Matrix uses Canonical JSON (for now)
* Matrix is perhaps more sophisticated on auth, in that we're switching to OpenID Connect for all authentication (and so get things like passkeys and MFA for free): https://areweoidcyet.com
* Matrix has an open governance model with >50% of spec proposals coming from the wider community these days: https://spec.matrix.org/proposals
* AT has done a much better job of getting mainstream uptake so far, perhaps thanks to building a flagship app from day one (before even finishing or opening up the protocol) - whereas Element coming relatively late to the picture has meant that Element development has been constantly slowed by dealing with existing protocol considerations (and even then we've had constant complaints about Element being too influential in driving Matrix development).
* AT backs up all your personal data on your client (space allowing), to aid portability, whereas Matrix is typically thin-client.
* Architecturally, Matrix is increasingly experimenting with a hybrid P2P model (https://arewep2pyet.com) as our long-term solution - which effectively would end up with all your data being synced to your client. I'd assume bluesky is consciously avoiding P2P having been overextended on previous adventures with DAT/hypercore: https://github.com/beakerbrowser/beaker/blob/master/archive-.... Whereas we're playing the long game to slowly converge on P2P, even if that means building our own overlay networks etc: https://github.com/matrix-org/pinecone
I'm sure there are a bunch of other differences, but these are the ones which pop to the top of my head, plus I'm far from an expert in AT protocol.
It's worth noting that in the early days of bluesky, the Matrix team built out Cerulean (https://matrix.org/blog/2020/12/18/introducing-cerulean) as a demonstration to the bluesky team of how you could build big-world microblogging on top of Matrix, and that Matrix is not just for chat. We demoed it to Jack and Parag, but they opted to fund something entirely new in the form of AT proto. I'm guessing that the factors that went into this were: a) wanting to be able to optimise the architecture purely for social networking (although it's ironic that ATproto has ended up pretty generic too, similar to Matrix), b) wanting to be able to control the strategy and not have to follow Matrix's open governance model, c) wanting to create something new :)
From the Matrix side; we keep in touch with the bluesky team and wish them the best, and it's super depressing to see folks from ActivityPub and Nostr throwing their toys in this manner. It reminds me of the unpleasant behaviour we see from certain XMPP folks who resent the existence of Matrix (e.g. https://news.ycombinator.com/item?id=35874291). The reality is that the 'enemy' here, if anyone, are the centralised communication/social platforms - not other decentralisation projects. And even the centralised platforms have the option of seeing the light and becoming decentralised one day if we play our parts well.
What would be really cool, from my perspective, would be if Matrix ended up being able to help out with the private communication use cases for AT proto - as we obviously have a tonne of prior art now for efficient & audited E2EE private comms and decentralised access control. Moreover, I /think/ the lexicon approach in AT proto could let Matrix itself be expressed as an AT proto lexicon - providing interop with existing Matrix rooms (at least semantically), and supporting existing Matrix clients/SDKs, while using AT proto's ID model and storing data in PDSes etc. Coincidentally, this matches work we've been doing on the Matrix side as part of the MIMI IETF working group to figure out how to layer Matrix on top of other existing protocols: e.g. https://datatracker.ietf.org/doc/draft-ralston-mimi-matrix-t... and https://datatracker.ietf.org/doc/draft-ralston-mimi-matrix-m... - and if I had infinite time right now I'd certainly be trying to map Matrix's CS & SS APIs onto an AT proto lexicon to see what it looks like.
TL;DR: I think AT proto is cool, and I wish that open projects saw each other as fellow travellers rather than competitors.
Matrix uses JSONSchema to define event schemas, but how can they be considered strict if the Matrix spec doesn't specify that any of them have to be validated apart from the PDU fields and a sprinkling of authorization events?
> Matrix has an open governance model with >50% of spec proposals coming from the wider community these days.
Do you have a percentage for the proportion of spec proposals from the wider community making it into spec releases?
It’s refreshing to see this perspective, and I appreciate trying to quash the us-vs-them thing, so thanks.
I'd love to get the decentralized protocols to work together. I work on braid.org, where we want to find standards for decentralized state sync, and would love to help facilitate a group dialogue. Hopefully I can connect with you more in the future.
it's literally just one guy on a Mastodon somewhere
If I don't want to use BGS, how can I access the PDSs of the people I follow, and get e.g. 10 latest entries for each:
Starting from a handle, let's say @pfrazee.com, how do I fetch using CURL your content without the BGS?
Of course, there are also other issues regarding monetization, and if BGS becomes the de facto way to get content, then some PDSs might become locked to your official BGS.
I hope you are willing to contribute with known non-profits such as Mozilla to do a wider consortium of players for BGS space, which is mostly inaccessible for self-hosters.
Can you elaborate on this point? There's multiple mature solutions in this space like OpenAPI+Json Schema, GraphQL, gRPC. All try to solve the same problems to varying degrees and provide similar benefits like generated static typed and runtime validation. Was there something unique for you that made these tools not appropriate and prevented you from building upon an existing ecosystem?
For example, the JSON Schema structures: `anyOf`, `oneOf` and `allOf` are fairly clear when applied toward validation. But how do you map these to generating code for, say, C++ data structures?
You can of course minimize the problem by restricting to a subset of JSON Schema. But that leaves others. For example, a limited JSON Schema `object` construct can be mapped to, say, a C++ `struct` but it can also map to a homogeneous associative array, eg `std::map` or a fixed heterogeneous AA, eg `std::tuple`. It can also be mapped to an open-ended hetero AA which has no standard form in C++ (maybe `std::map<string,variant<...>>`). Figuring out the mapping intended by the author of the JSON Schema is non trivial and can be ambiguous.
At some level, I think this is an inherent divide as each language one targets for codegen supports different types of data structures differently. One can not escape that codegen is not always a fully functional transformation.
I don't really understand why this is a problem. Unless you're using things like Haskell, Julia, or Shapeless Scala, you generally accept that not everything is modeled at the type level. I don't know the nuances of the C++ types you mentioned, but I have not encountered the ambiguity you described in Typescript or the JVM. E.g. JSON Schema is pretty clear on that any object can contain additional unspecified keys (std::map I assume) unless additionalProperties: false is specified.
> `anyOf`, `oneOf` and `allOf` for [...] C++ data structures?
Like I said I don't know C++ well enough, but these have clear translations in type theory which are supported by multiple languages.I don't know if C++ types are powerful enough to express this.
allOf is an intersection type and anyOf is an union type. oneOf is challenging, but usually modelled OK enough as an union type.
What I wanted to express is that using JSON Schema (or any such) for validation encounters a many-to-one mapping from multiple possible types across any/all given programming languages to a single JSON Schema form. That is, instances of multiple programming language types may be serialized to JSON such that their data may be validated according to a single, common JSON Schema form. This is fine, no problem.
OTOH, using JSON Schema (or any such) for codegen reverses that mapping to be one-to-many. It is this that leads to ambiguity and problems.
Restricting to a subset of JSON Schema is only goes so far. For example, we can not discard JSON Schema `object` as it is too fundamental. But, given a simple `object` schema that happens to specify all properties have a common type `T` it is ambiguous to generate C++'s `class` or `struct` or a `std::map<string,T>`. Likewise, a JSON Schema `array` can be mapped to a large set of possible collection types.
To fight the ambiguity, one possibility is to augment the schema with language-specific information. At least, if we have a JSON Schema `object` we may add a (non `required`) property to provide a hint. Eg, we may add `cpp_type` propety. Then, typically, the overhead of using a codegen schema is only beneficial if we will generate code in multiple languages. So, this type hinting approach means growing our hints to include a `java_type`, `python_type`, etc. This is minor overhead compared to writing language types in "long hand" but still somewhat unsatisfying. With enough type-theory expertise (which I lack) perhaps it is possible to abstractly and precisely name the desired type which then codegen for each language can implement without ambiguity. But, given the wealth of types, even sticking with just a programming language's standard library, this abstraction may be fraught with complication. I think of the remaining ambiguity between specifying use of C++'s `std::map` vs `std::unordered_map` given an abstract type hint of, say, `associative_array`. Or `std::array`, `std::list`, `std::vector`, `std::tuple` if given a JSON Schema `array`).
I don't think this is a failing of JSON Schema per se but is an inherent problem for any codegen schema to confront. Something new must enter the picture to remove the ambiguity. In (my) practice, this ambiguity is killed simply by making limiting choices in the implementation of the codegen. This is fine until it isn't and the user says, "what do you mean I can't generate a `std::map`!". Ask me how I know. :)
What DID are you using? Is it unique to AT or something more widely available?
Any tips for getting an invite to Bluesky for us nerds on here?
They host Mastodon instances, contribute to OSS, edit Wikipedia etc etc
AT and Bluesky are unfinished. Not ready for primetime. It's not fair to compare it to mature, well-developed stuff with W3C specs and millions of active users on thousands of servers with numerous popular forks.
But, also, everyone who hates Mastodon and spent the last months-years complaining about it is treating your project like the promised land that will lead them into the Twitterless future, somehow having gotten the impression that it's finished and ready to scale.
I think most critiques of Bluesky/AT are actually responding to this even if the authors don't realize it. They're frustrated at the discourse, the potshots, and the noise from these people.
Really? So rather than try to compare, contrast, and course-correct a project in its early stages by understanding the priors and alternatives, we should only do retrospectives after it has matured?
I would have thought this was the whole point of planning in early development: figuring out what you actually need to make? And that is usually a relative proposition, a project is rarely in a vacuum!
We should never just uncritically go ahead with the first draft in anything, especially not in the protocols we use, as they have this annoying habit of sticking around once adopted and being very hard to change after the fact.
As someone who wants to like bluesky, I feel like a lot of the scepticism comes from the seeming prioritisation of a single centralised server, over federation, for what is "sold" as a decentralised system.
An open protocol exists that broadly does what you want to do. That protocol is stable and widely used. That in itself, regardless of the quality of the protocol, already represents an OKish argument to strongly consider using it. If you're going to go NIH your replacement needs to not just be better but substantially better, and you should also show understanding of the original open spec.
Sofar, Bluesky's already been caught redoing little things in ways that show a lack of reading/understanding of open specs (.well-known domain verification), and the proposed technological improvements over ActivityPub fall into three categories:
1. Something ActivityPub already supports as a "SHOULD" or a "MAY": there's arguments to be made these would be better as "MUST", but either way there's no reason AT couldn't've just implemented ActivityPub with these added features.
2. Highly debatable improvements - as highlighted by this article. I do think some of this article is hyperbole but it certainly highlights that some of the proposed benefits are not clear cut.
3. Such a minor improvement as to be nowhere near worth the incompatibility.
All that coupled with the continuous qualifiers of it being incomplete/beta/WIP when there's a mature alternative really just doesn't present well.
However from an Ops perspective ActivityPub is incredibly chatty. If this had to scale to a larger instance the costs would spiral fast. Operationally and cost efficiency wise ATProto is a better looking protocol already. From a single individual user this won't necessarily be obvious right off the bat. But it will tend to manifest in either overworked operations people or slow janky instance performance.
While it's certainly a reasonable question whether the world needs another federated social protocol or not ATProto definitely solves real problems with the ActivityPub protocol.
> ATProto is a better looking protocol already.
Are there benchmarks for this? What's the level of difference here? Request frequency seems closely linked to activity and XRPC bodies are JSON just as ActivityPub so message size should be within order of magnitude at least. Are there architectural differences that reduce request frequency significantly?
> If this had to scale to a larger instance the costs would spiral fast.
Are we talking bandwidth costs or processing. I know the popular ActivityPub implementation is widely considered to be pretty inefficient processing-wise for reasons unrelated to the protocol itself: is that a factor here?
Now on my single instance it's not too bad because 1. I follow maybe 40 people and 2. I have like max 10 followers. For an instance with people with high follower counts across multiple other instances it could get to be a problem fast.
Edit: my previous description used fetch when it should have used push.
I think the weirdness is with Bluesky all that cost is still there but it's now handled by a small group of massive megacorps which is a real tangible benefit to self-hosters but you could have that on top of AP by running your service off what would essentially be a massive global cache of AP content which is what the indexer is.
I haven't implemented AP from scratch so I may be missing technical details precluding this, but it sounds like it may at least partially be covered by https://www.w3.org/TR/activitypub/#shared-inbox-delivery
... in which case it may be an implementation issue?
Mind you, there is liberal use of "MAY" there which I find is always a problem with specs: that would likely lead to mandatory chattiness of outgoing requests if you're federated with a lot of instances without shared inbox support, but should at least solve for incoming.
Which is why most browser-server communitcation is still limited by HTTP/1.0.
Except it isn't. That's not how this works out when there are benefits to all participants to implement these optimizations.
1. server with the sender constructs a map by receiving server that contains a list of all users on that server who should receive the message.
2. sending server iterates the map. If the receiving server has multiple recipients do a check to see if the receiving server supports this kind of 'bundled' delivery.
2a. If so send the message once with a list of receivers.
2b. receiving service processes the one message and delivers it to all the users.
3. If not sending server sends it in the traditional way, multiple pushes.Starting from scratch, just because you can theoretically design a better system, is one of the worst thing to do to users. Theoretically better also rarely wins in the marketplace anyway.
If you want a slightly lighter position: Software needs to be built to be migrated from and to, as a base level requirement. Those communities that do this tend end up with happy users who don't spend lots of toil trying to keep up. This is true regardless of whether the new thing is compatible with the old - if it doesn't take any energy or time, and just works, people care less what you change.
Yes, it is hard. yes, there are choices you make sometimes that burn you down the road. That is 100% guaranteed to happen. So make software that enable change to occur.
The overall the amount of developer toil and waste created en masse by people who think they are making something "better", usually with no before/after data (or at best, small useless metrics sampled from a very small group), almost always vastly dwarfs all improvement that occurs as a result.
If you want to spend time helping developers/users, then understand where they spend their time, not where you spend your time.
From a purely technical analysis ATProto looks better as a protocol to me. But I don't use Bluesky I use ActivityPub because the people I want to be connected to are there and not on Bluesky. I do think you could probably make improvements to ActivityPub that reduce operational costs. It's not something I feel the need to tackle right this moment because my usage doesn't incure those costs really.
The whole point is that it needs to be reasonably easy for people to run and scale their own servers. If people are constantly being burned out, quit, or run out of money, then it has an impact on regular users.
I don't think most projects do ;)
> If you want to spend time helping developers/users, then understand where they spend their time, not where you spend your time.
You're spending your time with ActivityPub, so this advice should apply to you, too. The bulk of potential users are spending their time on anything but ActivityPub. And as for developers, one of course needs to attract them, but I hear the people behind BlueSky have a couple of bucks, the ambition to create a huge potential new market, and a track record of creating a couple of things. I don't think they'll have trouble finding developers.
If BlueSky comes up with a better architecture, ActivityPub clients should rebase. BlueSky should pretend they don't exist, except if they have some nice schemas or solved some problem efficiently, try to maintain compatibility with that unless there's even the slightest reason to deviate.
Didn't ActivityPub have enough of a head start? Why didn't ActivityPub just use Diaspora? Why prioritize ActivityPub over OStatus?
Says everyone everywhere who thinks they made something better!
"If BlueSky comes up with a better architecture, ActivityPub clients should rebase. BlueSky should pretend they don't exist, except if they have some nice schemas or solved some problem efficiently, try to maintain compatibility with that unless there's even the slightest reason to deviate."
Look, i'm not suggesting whoever does it first gets to dictate it, but literally everyone thinks their thing will be better enough to attract lots of users or be worth it, and most never actually do/are. They do, however, cause lots and lots and lots of toil!
Your position is exactly what leads people over this cliff - better architecture does not matter. it doesn't. Technical goodness is not an end unto itself. Its a means, often to reduce cost or increase efficiency, and unfortunately rarely, deliver new features or better experience. Reducing cost or increasing efficiency are great. But architecture is not the product. The product is the product.
So, why is the Fedi not built on RSS/Websub/etc. then?
This is true everywhere?
* Improvements that'd layer cleanly on top of ActivityPub if they'd made any attempt at all.
E.g. being able to "cheaply" ensure that you have a current view of all a given users objects is not covered in ActivityPub - you're expected to basically want to get the current state of one specific object, because most of the time that is what you'll want.
So maybe it falls in the "highly debatable" category, but we also have a trivial existing solution from another open spec: RemoteStorage mandates etag's where a parent "directory" object's etag will change if a contained object changes, and embeds that in a JSON-LD collection as the directory listings. If you feel strongly that this is needed to be able to rapidly sync changes to a large collection, an ActivityPub implementation can just support that mechanism without any spec changes being needed at all (but documenting it in case other implementations wants to do so would be nice). Heck, you can "just" add minimal RemoteStorage compatibility to your ActivityPub implementation, since that too users Webfinger, and exposing your posts as objects in a RemoteStorage share would be easy enough.
Want to do a purely "pull" based ActivityPub version the way AT is "pull"? Support the above (w/fallback of assuming every object may have changed if you care about interop), make your "inbox" endpoints do nothing and tell people you've added an attribute to the actor objects to indicate you prefer to pull and so to not push to you.
Upside is, if any of the AT functionality turns out to be worthwhile, it'll be trivial to "steal" the good bits without dropping interop.
(Also, I wanted to see exactly how AT did pulls, and looked at the AT Proto spec, and now I fully concur with the title of this article)
Account portability is in no way incompatible with ActivityPub. It's not built into the spec., but it's also not forbidden / prevented by the spec. in any way. From ActivityPub's perspective it's an implementation detail.
Would it be nice if it was built into the spec: yes. Does that justify throwing out the baby with the bathwater?
I was hoping pfraze would offer some better reasoning, but while the post above is a great read for the technically curious, its very "in the weeds" so doesn't really address the important high-level questions. Textbook "technologists just want to technologise" vibes: engineers love reinventing things because working out problems for yourself is the fun part, even if many people have come together to collaborate on solving those problems before.
I've ported my account on ActivityPub a couple time, and it's a horrendous experience -- not only do I lose all my posts and have to manually move a ton of bits, but the server you port from continues to believe you have an account and doesn't like to show you direct links on there anymore.
The latter could probably be easily solved, the former needs to be built into the spec or it will continue to be broken.
> the former needs to be built into the spec or it will continue to be broken.
Absolutely, but ActivityPub doesn't preclude that. There's no reason for that proposed feature to be incompatible with the spec., or to have to exist in an implementation that is incompatible with ActivityPub.
Fwiw Mastodon, the most popular ActivityPub implementation, is (as is often the case with open standards) not actually fully compliant with the spec. They implement features they need as they need & propose them. This is obviously a potential source of integration pains, but as long the intent to be compatible is still there, it's still a better situation.
The ActivityPub spec says that object id's should be https URI's, not that they must. The underlying ActivityStreams spec just requires them to be unique.
All that's needed to provide full portability without a "transfer" is for an implementation to use URI's to e.g. DID's, or any other distributed URI scheme. Optionally, if you want full backwards compatibility, point the id to a proxy and add a separate URI until there's broader buyin.
I think there's be benefit in updating the ActivityPub spec to be less demanding of URIs, and instead of saying that they "should" be https require them to be a secure transport, and maybe provide a fallback mechanism if the specific URI mechanism is not known (e.g. allow implementations to provide a fallback proxy URL), but the main challenge there is not the spec but getting buying from at least Mastodon. The approach of providing a https URI but give a transport-neutral id separate to the origin https URI would on the other hand degrade gracefully in the absence of buyin.
Which is why promoting a new competing and incompatible protocol is a good way to push for change.
You want Microsoft to become more open? Then start promoting a pure FLOSS alternative until they can't ignore it anymore.
SSB is also decentralised, while ATProto is federated - like ActivityPub.
That's actually one of the things that bugs me about ActivityPub... unless I'm running my own one-single-user instance I won't have control over my own identity.
It's also very weird how even on a supposedly "federated" system the only way to ensure you can access content from all instances (even if they differ in philosophy or are on opposite sides of some "inter-instance-war") is to have separate accounts for each side... it kind of defeats the point of federation. There's even places like lemmy which instead of using blocklists use allowlists, so they will only federate with pre-approved instances.
I'd rather favor a more decentralized structure that allows the users to directly access content from any content provider that hosts it through the protocol (ie. without necessarily requiring another specific instance to index that content from their side.. if they index it great, but if they don't it should still be possible to access it using the same user account from a different index), with a common protocol that allows a separation between identity management and content provider.
From what I understood, ATProto is closer to that concept.
The W3C spec leaves all the hard parts to vendors, which is why the only DID implementation up to now has been Microsoft's, which relies on an AD server in Azure. Much decentralise.
Bluesky's isn't that, but a hash of some sort, which is centrally decentralised ... on their servers? I think this is one of the bits of AT that isn't finished yet.
But "W3C DID" is not a usable spec in itself, it's a sketch at best.
It is but that's largely conceptual so doesn't really mean anything in the context of the protocol. These are message exchange protocols so the defining element is whether the messaging is federated or decentralised.
Fwiw AP also supports DID I just haven't seen any implementations use it, since it strongly recommends other ID (a mistake imo).
> It's also very weird how...
What you're describing here is a cultural phenomenon, not a technological one, so isn't really relevant to the discussion: ActivityPub siloing isn't a feature of the protocol, it's an emergent feature of the ecosystem/communities.
It's also worth mentioning it isn't usually implemented as you describe, unless you're specifically concerned with maximising your own reach from a publishing perspective: most instances allow individual users to follow individual users on another "blocked" instance - its usually just promotion/sharing & discovery that are restricted.
If they are still allowing access to all forms of third party content through their own instance (even if they restrict the discoverability) then they are still risking being held responsible for that content. So imho, that would be a mistake.
Personally, if I were to host my own instance under such a protocol, I'd rather NOT allow any potentially illegal content that might come from an instance I don't trust to be distributed/hosted by my node.
The problem, imho, is in the way the content needs to be cached/proxied through the node of the user in order for the user to be able to consume it. This is an issue in the design of how federation typically works.
I'd rather favor a more decentralized approach that uses standards to ensure a user identity can carry over across different nodes of content providers, whether those nodes directly federate among themselves or not.
There should be a separation between identity providers and content providers, in such a way that identity providers have freedom to access different content providers, and content providers can take care of moderation without necessarily having to worry about content from other content providers with maybe different moderation standards.
I'm not saying ATProto is that solution... but it seems to me it's a step in the right direction, since they separate the "Personal Data Server" from the "Big Graph Services" that index the content. I can host my own personal single-user server without having all the baggage of federating all the content I want to consume. The protocol is better suited for that use case.
In services using ActivityPub, instances are designed for hosting communities, they come with baggage that's overkill for a single-user service but that's still mandated due to how the communications work, they expect each instance to do its own indexing/discovery/proxying. So they are bound to be heavier and more troublesome to self-host, and at the same time, from what I've seen, the cross-instance mechanisms for aggregation in services like Mastodon are lacking.
You will find that many people do not dig into details. You can post on Mastodon, you can post on Bluesky, therefore they must be similar.
It does mean that learning about things becomes a superpower, because you can start to tell if a criticism is founded in actual understanding or something more superficial.
I just disagree with this in principle. I wonder what the tech equivalent of "laissez faire" would be.
To this day I don't understand how anyone in the tech world thinks they can make a single demand of anyone else. Even as a customer I believe you can really only demand that the people you are paying deliver what is contractually and legally required. But outside of that ... I just don't understand people's mentality on this subject.
What is it about this specific area of the web that attracts these idealogical zealots? I had the same head-scratching moment in the early 2000s when RSS and Atom were duking it out.
Ultimately, yes you're right, no-one can make actual "demands" of anyone else. My language above is demanding, certainly, but ultimately I'm just arguing opinion. I cannot control any outcomes of what Bluesky or any other enterprise choose to pursue.
> What is it about this specific area of the web that attracts these idealogical zealots?
I think it stems from the unprecedented success of such zealots in the 1980s, which have differentiated the technological landscape of software technologies from previous areas of engineering by making them more approachable, accessible, interoperable and ultimately democratised. That's largely been the result of people arguing passionately on the internet to advocate for that level of openness, collab & interop.
I would challenge this belief, that the success of Linux, or the Web technologies (TCP, HTTP, HTML, etc.) were primarily the result of the zealots from the 80s. I would challenge the belief that protocols for Twitter-like communication fall into the same category as things like POSIX, TCP/IP, HTTP, HTML, etc.
> That's largely been the result of people arguing passionately on the internet to advocate for that level of openness, collab & interop.
My own opinion is the key to success was a large number of people writing useful code and an even larger number of people using that code.
One example that comes to mind is how HTML was spun out from w3c into WhatWG. Controversial at the time to say the least but IMO necessary to get away from the bickering of semantic web folks who were grinding the progress to a halt. HTML 5 won the day over XHTML (actually to my disappointment). The reason wasn't the impassioned arguments of the semantic web zealots, it was the working implementations delivered by the WhatWG members to serve the literal millions of people using their applications. Another example is the success of Linux over Hurd - the latter being a technology initially supported by the most vocal and ideological of all the zealots.
It is a simple fact that if AT garners sufficiently useful implementations and those implementations garner sufficient numbers of users - all of the impassioned arguments against it will have been for naught. So those arguing should probably stop so that they can focus on implementing ActivityPub (or whatever they think is best) and attracting users. I guarantee you that if ActivityPub attracts millions of users then Bluesky will suddenly see the light. They will change course just like every big tech company did when Linux became successful.
It is also why I think all of the hot-air on RSS and Atom was literally wasted. As technologies both have thus-far failed to attract enough users to make it worth it. I would bet that the same will be true of both AT and ActivityPub. Unless someone develops an application that uses one or the other and it manages to attract as many users as Twitter then people are just wasting time for nothing.
That you have no legal basis to enforce those demands does not mean that you cannot make them. They can be ignored of course but they may not be. Not all interactions have to be governed by a literal contract as opposed to a social contract.
For such users, any offering without these is a non-starter, dead-on-arrival.
People with resistance to this epiphany sound like those who used to insist, "HTTP is fine" (even when it put people at risk) or "MD5 is fine" (long after it was cryptographically broken). Most will get it eventually, either through painful tangible experiences or the gradual accumulation of social proof.
A bolt-on/fix-up of an older protocol might work, if done with extreme competence & broad consensus. And, some in the ActivityPub world had the cryptoepiphany very early! Ideas for related upgrades have been kicked around for a long time. But progress has been negligible, & knee-jerk resistance strong, & the deployed-habits/technical-debts make it harder there than in a green-field project.
Hence: a new generation of systems that bake the epiphany in at their core – which is, ultimately, a more robust approach than a bolt-on/fix-up.
Because so many of those recently experiencing this cryptoepiphany reached it via experience with cryptotokens, many of these systems enthusiastically integrate other aspects of the cryptotoken world – which of course turns off many people, for a variety of good and bad reasons.
But the link with cryptotokens is plausibly inessential, at least at the get-go. The essentials of grounding identity & addressing in cryptography predate Bitcoin by decades, and had communities-of-practice totally independent of the cryptoeconomics world.
A relative advantage Bluesky may have is their embrace of cryptographic addressing behind-the-scenes, without pushing its details to those who might confuse it with promotional crypto-froth. Users will, if all goes well, just see the extra security, scalability, and user sovereignty against abuses that it offers. We'll see.
Crytography and security in general are often cargo-culted without any consideration for the negative implications.
> this cryptoepiphan
Bro are you for real.
It's not fine here in 2023.
If you need a secure hash, it's been proven broken for 10 years now.
If you don't need a secure hash, others are far more performant.
Using it, or worse, advocating for its use, is a way to signal your thinking is years behind the leading edge, and also best practices, and even justifiable practices.
HTTP's simplicity could make it tolerable for some places where world-readability is a goal - but people, echoing your sentiments here, have said it was "fine" even in situations where it was putting people at risk.
Major browser makers recognize the risk, and are now subtly discouraging HTTP, and this discouragement will grow more intense over time.
A protocol like this doesn't work without community adoption. And the best way to get the community to adopt is to build on what's already there rather than trying to reinvent the fridge.
IMHO, any protocol that isn't signing content (like mastodon) is merely moving the problem. Signatures allow people to authenticate content and sources and allow people to build networks of trust on top of that. So bluesky is getting that right. Unsigned content should not be acceptable in this century.
Signed content immediately solves two issues:
- reputation: reputation is based on a history of content that is liked and appreciated by trusted sources that is associated with an identity and the associated set of public keys. You can know 100% for sure whether content is reputable or not. Either it is signed by some identity with a known reputation or it is not.
- AI/bot content could sign content with some key of course but it would be hard to fake reputation. Not impossible of course but you'd have to work at it for some time to build up the reputation. And people can moderate content and destroy the reputation of the identity and keys.
The whole problem with essentially all social media networks so far is a complete and utter lack of trustworthiness. You could be reading content by an AI, a seemingly bonafide comment from somebody you trust might actually come from some Chinese, Russian or North Korean troll farm, or you are just wading through mountains of click bait posted by "viral" marketing companies, scammers, or similarly obnoxious/malicious publishers. Twitter's blue tick is laughably inadequate for dealing with this problem. And people can yell whatever without having to worry about their reputation; which causes them to behave in all sorts of nasty ways.
Signed content addresses a lot of that. You can still choose to be nasty, obnoxious, misleading, malicious, etc. but not without staking and risking your reputation.
Mastodon not having a lot of these issues (yet) is more a function of its relative obscurity rather than any built in features. I like Mastodon mainly because it still feels a bit like Twitter before that got popular. But it's not sustainable. If a few hundred million people join, it will get just as bad as other networks. In other words, I don't see how this could last unless they address this. I don't think that this should be technically hard. You need some kind of key management and a few minor extensions to the protocol. The rest can be done client side.
That would be more productive than this rant against the AT protocol.
It's annoying how the word blockchain got associated with the worst of cryptocurrency excesses, isn't it ?
Also :
https://medium.com/@shemnon/is-a-git-repository-a-blockchain...
for some reason people think what we have now is good, and they say "dont reinvent the wheel" but there is no wheel, what we have is just garbage, 50 years and later we still cant beat the "unix pipe"
we have to keep trying to make a wheel
[1] Users with a profile are at about 2.2 million, see : https://stats.nostr.band/
There's the protocol, but then there's also the community – and communities are ecosystems not just technical problems.
I can get com.atproto.server.createSession to return me a token, but then when I try to transition into another endpoint (such as app.bsky.richtext.facet) I only get 404s. Is their any examples of using y'all with rest? Happy to take my question elsewhere if y'all have any places for folks to ask questions?
Technical merits of the different approaches aside... this is a calm, kind and well reasoned response to a rather, uh, emotional critique. Thanks for bringing the discourse back to where it should be.
I suspect it's due to me using an uncommon tld for my primary email.
Email: blade -at- coates dot life.
This... doesn't seem too bad at all?
Let's assume an average of 150 bytes of text per tweet, and an additional 50 bytes of actually important metadata. That's only 10 megabytes for the entire archive of 50,000 items. A single HDR photo from a modern smartphone is larger.
IMO this was the original point of Twitter: the extreme limitation on post size (140 characters) made it feasible to build applications that work with large amounts of content. There was a time around 2010 when building a Twitter API client was the standard demo for a new UI framework — the data model was simple enough that a basic client was the next step up from "Hello world" complexity.
This was actually a nice vision for a social network! Lots of simple near-realtime data, a jungle of interesting clients to make sense of it. I wish someone was building a Twitter competitor with this kind of minimalist approach.
Yes, there will be broken links. But IMO that’s better than having all your data in one centralized location where it can eventually be taken over by private equity looking to make a buck, or a billionaire who wants to be a media mogul.
Edit: Looking back at historical Twitter news, it’s clear that in 2010 their vision was still to have images embedded from other sites. Here’s a relevant TechCrunch article:
https://techcrunch.com/2010/09/14/new-twitter-tips/
Note how they added display support for embedded Flickr photo set links. Your pictures could be on the dedicated photo service and you’d still get the nice set browsing UI inside Twitter. They should have stuck with this and built a protocol that lets sites interoperate on things like “here’s a set of photos to be embedded” (so you could use something else than Flickr). But instead they wanted to chase the wannabe-Facebook dream of sucking everything into their servers and using a closed client to sell inline ads.
https://www.theverge.com/2012/7/9/3135406/twitter-api-open-c...
For instance it was basically the end of Flattr 1.0 :
https://venturebeat.com/social/flattr-the-crowdfunding-payme...
The storage on your phone would probably still be enough for nearly everyone and for the rest doing the account moving process on a PC or through cloud storage would probably be ok.
A lot of users simply don't know or don't recall just how barebones the original Twitter app was.
One thing I do like about the decentralized model is the ability to modularize this.
I would love a Bluesky/ATProto/whatever client that allowed me to specify some compatible image service where images I post get seamlessly uploaded and linked.
Then I can choose what I use, maybe it's something I pay for and have better guarantees around durability and availability.
Or maybe it's a free or ad-supported service where my image gets deleted after 6 months and I don't care that's fine with me.
Giving users more choices is good, IMO.
It's fascinating the Twitter we have today and the Twitter from that era are even considered the same app. I wrote a simple Twitter client in Cocoa and Python and even NodeJS. It was such a simple way to get a working, functional, usable demo going that you could immediately show to friends.
Then they realized it's impossible to make money advertising this way and shut it down.
You might suspect that OP has a familiarity bias here but actually there is objective evidence that ActivityPub based implementations are (relatively) simple: there are dozens of implementations of both servers and clients, will all sorts of functionality that is not emulating the "twitter/mastodon" experience. Heck, even a Wordpress plugin in the works.
How well all these things will scale etc. is still somewhat of a question mark but this aspect feels important beyond specific details and choices. The simpler, more generic an approach, the more likely it is to find fertile ground and grow as a decentralized architecture. The original decentralized Web was simple and generic and this has been touted as key factor for its explosive adoption.
In a sense this was also its ultimate downfall: it did not provide (out of the box) the tools to build the connectivity / social graph experience that was so enormously desirable (and was not provided e.g., by RSS). ActivityPub is somehow making up for that original gap. There might be other ways to do this. But unless there is a bigger agenda (commercialization, financialization), gratuitous complexity driven by not-invented-here or desire for centralized control is more likely to hinder than help.
AT Protocol has a very active ecosystem already too (even though Bluesky is still invite-only): https://github.com/bluesky-social/atproto-ecosystem
And this isn't a complete list; there's a very active Discord for people doing bluesky projects which currently has 1.2k+ members
But on the other hand, the AT protocol is not all what it seems either. It's fundamentally not possible to do what they promise unless there is some kind of synchronization point where your most up-to-date identity can be searched. The documentation goes on extensively about the format of DIDs and the fact that there _is_ a resolution scheme, but it fails to mention that this resolution is being done by Bluesky themselves, and is not planned to expand into something that others can control. As a decentralized protocol, you would expect this to be something DNS-based, or if you really wanted something more peer-to-peer, DHT lookup, or at the worst case, a distributed blockchain. Seeing as this is a fundamental protocol-level change, the fact that they're rolling with the current approach makes me believe they will not be changing this in the near future. Bluesky is just converting one problem into another.
The only options that seem reasonable for PLC is for it either to be operated by some trusted multi-stakeholder institution like ICANN or to be a closed-group blockchain. I'm open to either approach but they require the creation of a consortium, which requires buyin from a set of stakeholders that don't exist yet. So until that happens, we run it, and that's why we called it Placeholder. We'll backronym it into something else once we get there (the C most likely being Consortium).
AtProto is designed to be a federated protocol. The issue I have is that it is not interoperable with the major standard used on the federated internet right now: ActivityPub. You can built protocols on top of each other. Instead of doing that, Bluesky built a confusing alternative that is difficult to implement and difficult to work with.
> In fact, it misses that having pull-based indexes is part of the idea of separating data publication vs. data curation.
Mastodon does this already. There is literally an 'explore' part of the app that is completely separate from the feed and does not rely on ActivityPub.
You cannot 'publish' content without having some sort of a federation protocol to go overtop of it, to network things together and send your content to the people who follow you. There is nothing about ActivityPub that makes a single global view impossible or hard to implement.
The only winning move is to onboard users and given hundreds of millions of people may be looking to leave Twitter, Mastodon utterly shit the bed and missed the moment.
That’s the end of the story: if the system can’t catch users raining down from the sky during a generational upheaval in social media it ain’t never gonna make it.
personally I hope it doesn't actually become "the one thing" with how opinionated and callous it is about some features
The only worse idea than Bluesky using ActivityPub would be to build something new making all the same design decisions.
Deciding “ActivityPub is the standard” (seriously?!) and demanding we give up already is the opposite of what we need.
I don’t know if AT is the best long term solution, but — and I’ve tried it multiple times — ActivityPub/Mastodon sure as hell isn’t. There’s no value in being interoperable with it that I can see, beyond the potential short term boost to vanity metrics on user numbers.
I absolutely want my identity to use public key cryptography, and I absolutely want to store all my emails, tweets (or whatever), and DMs locally first. I don’t like how accounts and servers work in Mastodon and the federation and global discovery approach sucks too.
The best thing we can do now is keep an open mind and try as many approaches as possible, because if something better than ActivityPub doesn’t come along society is stuck with centralised social media forever.
ActivityPub is very generic; it can accommodate all kinds of changes. E.g. if you want to introduce activity / object types that have improvements over what Mastodon supports, you can do so. If you want to introduce vocabulary within existing object types that Mastodon wouldn't understand, you can do so without breaking federation with Mastodon or other Fediverse servers. I'd be a lot more sympathetic to anyone who decided to extend ActivityPub.
The minimum viable subset of ActivityPub is largely: Provide an endpoint that returns an Actor with a list of the required endpoints (inbox, outbox, follower, following etc.) - the spec requires a list of them, but you don't even need all of them for basic interop - and handle POST's to your announced inbox, and GET requests to the outbox. Add unique URL's as "id" fields in the JSON, and support GET to them. Follow the format of the JSON to provide at least the minimum set of fields to address the activity, provide a type, and provide the minimum fields for the given type.
For interop w/Mastodon you'd want to support Webfinger to find the Actor. Nothing stops you from also supporting other mechanisms, like BlueSky's domain validation.
Nothing stops you from supporting additional federation mechanisms. Nothing stops you from providing additional fields. Nothing stops you from storing data locally. Nothing stops you from adding additional ways of signing claims about individual objects or a whole repository of objects. Nothing stops you from providing additional mechanisms for distributed lookup of objects by id. Many of those things would be welcome if people wanted to do it on top of ActivityPub.
Why bother with Mastodon interop? Why bother building on top of a standard that doesn’t do what you want if you don’t think being part of the “Fediverse” is particularly interesting or a goal of the platform you’re building?
Something new is a better bet at this point than being anchored to or seeming like part of Mastodon IMO.
Because they pretend to want to be open, and it's sending a very clear signal that is not their goal if they're not even trying to work with the existing ecosystem.
If they just want to be a silo, that's fine. But in that case be honest about it.
No matter how open I wanted to be, if I was setting out to build a social network, I’d please precisely zero value on being connected to that “existing network”.
It is irrelevant to me, despite what a vocal minority might want me to believe.
This might be a true statement for microblogging, but the major-est "standard used on the federated internet right now" is surely SMTP+IMAP, and I wouldn't be surprised if RSS/Atom were in the #2 spot even though "nobody" is using it anymore (esp. considering "nobody" includes most podcasts).
Good? HN also uses crypto. Unless I'm missing it, it's not using cryptocurrency or 'tokens' or anything, just good old fashioned cryptography.
Wait, is that not what we're talking about? So, it has a confusing name too?
ATM
ATS11=50
And a script to keep doing ATDT##### in a loop till it got through a busy signal, only to find out that short DTMF tones aren't always recognized, and seemingly the tones for "9" and "1" and "1" work when others didn't.Explaining to my parents why there was a cop at the door complaining that the 911 operators were mad at us was fun.
While I am here: is anyone aware of a solid client library for handling AT commands? Ideally, suited for an embedded environment so C, no POSIX, no malloc, but really anything would do. Even just a solid implementation of handling the client to steal ideas & code from?
I hope I have just disconnected someone browsing this page with a dialup modem.
This only works because Hayes patented the idea of the escape code being +++ followed by a delay, so to evade the patent, most other modem OEMs removed the delay requirement.
NO CARRIER
but, when I upgraded to touchtone, I remember the local access number for Compuserve in Watertown MA played The Camp Town Ladies (without the doo dah)
ActivityPub is no walk in the park either. It's a protocol closer to "the HTTP of social media" than a way to interoperate with other servers. Basic functionality is easy, all you need is a few static JSON files, but if you want to write an application around it you're going to have to dive deep into the docs.
What I don't really see is why BlueSky decided to make its own protocol. Like Mastodon is an API that works on top of ActivityPub, BlueSky could've just been a better ActivityPub server as far as I can tell. Most of the big problems ("my toots disappeared after the server shut down") can be resolved without an entirely new protocol.
Going back to the drawing board seems like an excellent way to find all the hurdles every other protocol already encountered. The recent "official" s3 account is just one example of that, and I'm sure there will be more.
Maybe the value add of the AT protocol will become clearer once it's finished. I'm sure there will be ATProto <-> ActivityPub bridges to make both networks integrate for the people who wish to do so. As far as I can tell, BlueSky is just a new, exclusive Twitter with an API at this moment.
Honestly I assumed that was the USP - Twitter sans Musk.
Almost. AT yes, but not protocol. I always considered it to be a user interface.
Given that a user owns their feed, I'm not sure why this is a bad idea?
I want to take this critiques here seriously but there's not a ton to grasp onto. It definitely didn't feel great that AtProto built from bare ground, reinvented json-schema & OpenAPI for no reason. But ultimately this is one of a number of grievances that feels like doesn't really matter. It's stupid & dumb but in the end it doesn't matter. It's hard to tell which of these points really have hard & real impact. And which are just bias.
The general feeling is an unencumbered letting out of biases, which makes it harder to trust.
In the scenario where every authors can only change their own files, you can avoid those sort of conflicts, as they'll never happen.
Remaining is how to solve files referencing other files (in the case of a federated social media, posts that are replies to other posts) which are way easier to solve, as it's not really conflicts but invariants that have to be created.
A single author can still run into these issues with git if they use multiple devices.
Can the same happen in this social media network for a user that uses the same account on their phone, their tablet and their computers?
No, using multiple devices are fine, as long as you're not editing the same content in two different ways on them. Meaning changing "A B" to "A C" on one device, and to "A D" on the other.
> Can the same happen in this social media network for a user that uses the same account on their phone, their tablet and their computers?
When using AT protocol, you'd still be connected to the same PDS, so no conflicts. The conflicts mentioned earlier are about server<>server federation, not server<>client communication.
If I am correct then conflicts are impossible and you are the only one that can write to your* "branch"
* I doubt that forking and merging are possible so it does not really have any practical UX relations to git rebase.
https://initialcommit.com/blog/pijul-version-control-system#...
The functionality would be obfuscated behind easy to use UX constructs.
Anyone worrying about having to rebase, is over thinking it. This is a protocol level detail hidden behind shiny UX elements.
I’ve been using Bluesky for 2 months now; it’s very easy to use.
that by itself is a bad idea
users being able to arbitrary edit tweets in a social network is never a good idea
because the moment you say (i.e. tweet) somethingpublic it's not just your thing anymore,
Don’t spread misinformation please.
To use the Mastodon web application, please enable JavaScript. Alternatively, try one of the native apps for Mastodon for your platform.
Good job, Mastodon. (No, I am not the same guys that complain they don't see javascript games or utilities without javascript. But I think I should be able to see a text on the Internet?)Still, it doesn't show the rest of the toots, so content posted on Mastodon is still not readable without javascript. Which is ridiculous, we have medium, twitter, facebook, reddit etc, and they decided to publish somewhere not readable without javascript.
The issue is being tracked here: https://github.com/mastodon/mastodon/issues/19953
I mean it's an application, applications contain code and tracking that code state server side can involve a non small amount of additional cost.
Sure they could have a web-side like non-logged in viewer mode and then somehow using hydration magic transition to being a app on demand, but that is a bunch of additional work (i.e. cost) for a very small number of people.
But here is a good thing, it's open source so you or any of the view people which care about no-js could try to contribute such a "view only web-site" functionality, or crate an alternative web client which solely uses server side rendering or similar.
But in the end that's not worth your time right (at least for me it wouldn't be if I where in your situation)? So why expect anyone who doesn't even get any benefits from it to invest that time?
At this point I could just flag every mastodon content, since it is not available. Good idea, I will just do that.
If HN requires sources to be available in no-js only form yes, that would be appropriate, if not it would be abuse of the flagging feature.
- Brutaldon: https://brutaldon.org
- Source: https://gitlab.com/brutaldon/brutaldon
Alice on server X follows Bob on server Y. When are Bob's posts delivered to X such that Alice may see them? Are they pushed from Y? Are they pulled by X?
Polling seems like an absolute disaster in an environment where people have an expectation of semi real-time communication.
> It uses pull-based federation instead of push-based like Mastodon.
I was hoping they would employ both: push and pull. Some scenarios I could imagine pull being more efficient, and sometimes push. For instance if I run my own server I don't need to be pushed all content when I sleep. However when I return it could do pulls and then continue pushing. It doesn't look like it's clever like that.
Thrift, cap’n proto, grpc, etc have already been in production for years now.
WE HAVE THE ABILITY TO DO THAT AS SERVER ADMINS!!! MASTODON HAS THIS ALREADY!!“
I’m sorry, but how can you take someone seriously who makes comments this absurd.
This reads like jealousy.
Other systems such as Mastodon avoid running into these problems because of direct and implicit use of DNS for the namespace. (And use of DNS embedding in DIDs ends up undermining the value of flat naming.)
> Imagine if I had to store the 50k+ tweets I've made on Twitter on my device, and upload ALL of them to a new server whenever a community server went down.
doesn't make this person seem particularly competent. I'm getting a vague feeling that this is normal discourse on Mastodon though, and that not using social media much shifts your personal overton window for what is acceptable communication until you essentially don't overlap with the very online crowd anymore.
After that, it just felt like rant.
50k entries isn't a whole lot but as people hosting Mastodon servers have found out, things start slowing down when 1000 people transfer those 50k entries at the same time.
1. They are URI's, and while ActivityPub say they should be https URL's, they don't need to be, and could e.g. point at IPFS or similar.
2. JSON-LD signatures are used by Mastodon, and included in the export, and nothing stops another instance from validating those and serving them up with the original URIs in the id from new URLs, as a means of making it clear the server didn't originate them (there'd be a trust issue if the other servers is unable to get hold of the keys because the original server is gone, but no more so than if the new server had simply republished the content, so the "worst case" is to distrust the original id's).
There are some corner cases there, around trusting the identity of the old and new account represents the same user, so I do think a recovery key type scheme would be nice to allow a user to prove the old and new id is the same (if changing id; I also think we could really use decoupling the expectation that a webfinger id is inherently tied to a Mastodon account - you can sort of do that today; nothing stops you from serving up a separate webfinger result and use it as an alias, but there are usability issues to solve there).
The main problem would probably be keeping track of what server to fetch these messages from after a move (or even a second move) to a different server and keeping the metadata attached in sync.
You're thinking of password hashing functions like argon2, there's no reason for normal signing and verification operations to be intentionally expensive as that's not where the security guarantees come from.
Compared to something like simple a CRC32 checksum to verify that the data was transmitted correctly, these operations will always be complex and more time intensive than you'd prefer them to be.
The OP was going on about storing all the data on-device and uploading it, but regardless of where it’s stored, if a bunch of people have to move, the thundering herd problem, so to speak, will still exist.
Also, signature verification is not slow. I don’t know what AT uses or how good this source from 2020 is [0] nor what machine it ran on, but ed25519 verification takes about 50us for a 32 byte signature verification. That suggests this guys 50k posts will be validated in 2.5s. If we assume just a single server then that’s about 34k users per day. Or put another way, 30 servers to onboard one million people in a day. None of this seems outrageous to me.
[0] https://safenetforum.org/t/ed25519-vs-bls-performance/32613
Verifying those messages will take about a minute of CPU time per user (assuming no impact from cache misses due to threads swapping in and out and processing new data). I think that's quite significant.
But is this the only possible implementation? I suppose for the OP to have a reasonable point (that it’s “a crock of shit”), this would need to be an intractable problem.
https://www.theverge.com/2012/7/9/3135406/twitter-api-open-c...
On Twitter, it was the entire world. And if you wanted to hold the conch shell, you did what the algorithm wanted which was OUTRAGE.
On Mastodon, it's theoretically scoped down to just your server and a handful of deliberate federations. This exists. You can find these servers and have a BBS-like experience. Tooters give you the time of day and assholes are shown the door.
The Very Online folks, however, use Mastodon differently. For them it's more like "I'll build my own Twitter!" And what they want is more OUTRAGE feed to tap into 20 times a day, but with them in control of the algorithm. So they federate freely, with very large servers, making no material progress on the thing that makes Twitter unhealthy.
The fact that Mastodon can seemingly be used both ways is a good accomplishment. It's not the technology's fault the second use-case is socially toxic. But it does make it very hard to talk about Mastodon because you don't immediately know which way anyone is using it.
We tested about 10 different modems. All of them had their own unique bugs I had to work around. Or maybe one of them didn't have any bugs (that I ran into) -- it's been thirty years.
And to be honest it's a good thing: many non-tech friends were a bit confused about Mastodon, where you have to understand parts of the technology behind to do basic things (like follow someone from another instance, IIRC). BlueSky is (currently) a bit more friendly UI/UX-wise. I don't know how that will evolve with more instances, though. I also found that discovering people was easier, but YMMV.
It seems like you're conflating the technology with the user experience. Yes, the user experience is similar. No, the technology underpinning it isn't.
The twitter experience has been in decline for years, and took a sharp downward turn when ownership changed. But people on twitter just have to live with it, especially now that the (already limited) API has been closed down almost completely, because it's not an open platform like bluesky. If bluesky ever starts to make the official client a bad experience, or starts to make moderation (or non-moderation!) decisions people don't like, users don't have to just live with it. They can go find or create the experience they want without losing anything
Interesting. About two weeks ago at the HIMMS conference (Healthcare Information and Management Systems) in Chicago I ran into this. DID has crypto stink on it and people are actively avoiding it as a result. In this case it was the CTO of an established US consortium involved with CMS standards.
This is a shame and it seems irrational to me. Is W3's work on DID doomed? Do they know just how bad the optics of DID are?
I mean, I'm sure they have some points in this ranty thread of toots, but it's hard to take it serious when it ends up blaming it all on capitalism (which, I hear you brother, I'm no fan either) and it's FILLED WITH SHOUTING FOR NO GOOD REASON.
Sometimes, when you feel strongly about something, it's useful to write a first draft, and come back to it after cooling down for a day or two, and rewrite it to be more nuanced.
The AT Protocol has many flaws, like any protocol, that much is evident. But I don't think this toot-thread gives a accurate view of those flaws. It also doesn't seem to consider that all these different protocols make different tradeoffs, and none of the protocols try to be "one protocol to rule them all". They are simply better at some things, and worse at others.
> Also I don't care if I'm spreading FUD or if I'm wrong on some of this stuff. I spent an insane amount of time reading the docs and looking at implementation code, moreso than most other people.
> If I'm getting anything wrong, it's the fault of the Bluesky authors
This is a really disappointing way of reviewing things, I hope that it doesn't become more popular, because no one actually learns anything from it. It ends up being just a rant, but masqueraded as education.
I dunno, this seemed an entirely on-message thing for a “literal communist” to say.
Here's the last generation of critique, when Diaspora developers discussed the good and bad in ActivityPub, and why they wouldn't support it:
- https://overengineer.dev/blog/2018/02/01/activitypub-one-pro...
- https://overengineer.dev/blog/2019/01/13/activitypub-final-t...
You might be surprised. A lot of cellular modems and Bluetooth modules are controlled using a variant of the AT command set.
There's large forums such as wirelessjoint, ispreview and sierra where AT is still extremely relevant.
Twitter had what, 160 or 280 char limit per message? 50k*280 is 14MB. What's to imagine here?
Isn't everyone on a "competing" social network before they join Bluesky?
I managed to snag an invite (sorry, I don't have any others), so I'm on both at the moment. Though I've found Bluesky to be pretty boring and don't check in much.
> And understand better why this guy is so upset about this destined-to-fail network to curse and foam about it.
It didn't sound like he wanted to fail: 'And I went into this with an open mind. I was like "I'll just make a simple alternative to the BlueSky server in Elixir".'
This seems troubling, if it's accurate.
[I've edited the above quote to remove some all-caps and exclamation points.]
The author totally glosses over that and scoffs at it being a real problem both in the thread linked and here in the comments, yet real people have experienced the extraordinary pain of watching their accounts evaporate overnight because of some capricious server admin. It's a total nonstarter for ever using Mastodon.
1) If you add image files and especially video files to the mix, the size of your data can get huge.
2) Smartphone users especially don't have a ton of free storage space.
3) In general, people aren't great at maintaining their own backups.
The question is why a so-called distributed network can't maintain distributed copies of user data, rather than forcing the burden onto the users.
Another VC-funded bait and switch; nothing to see here.
> But I think what's key is to keep an anticapitalist mindset. We can make things easier for users without allowing in what makes social media so fucking awful: capitalism.
Normies do not care about this stuff in the slightest and it isn’t worth confusing them to include it and doesn’t benefit the site now it’s gunning for being “new Twitter”
- Dodges any free-speech issues
- Gives individuals more choice and control over what they see
- Allows labellers to not worry as much about false-positives because their impact is limited by the above, which means they can use more automation, etc.
- Allows hate speech to hide in plain sight.
- Allows plausible deniability that you aren't the nazi bar.
- May allow moderation labeling to be used as a form of harassment, by intentionally using labels inaccurately.
- Does not actually absolve you of the write-side responsibility to filter illegal content.
But no, everyone's all just butthurt over Elon Rocketman and will accept whatever garbage anyone puts in front of them.
Meanwhile I was expecting a Linus rant about a kernel merge for it
And it's awkward to write code for and the people who inhabit it range from right-wing reactionaries to, as was so effectively put, the homeowners' association. Turns out that people don't care about "open source" unless it's easy to work with and people especially don't care about "open source" when they want to talk and shitpost with their extended friend group. (Which no, Mastodon doesn't do a good job of! Even setting aside that quote-toots are Good, Actually, Mastodon doesn't let me see replies to a toot unless I go digging so I don't know if I'm just being one of another set of replies that might already have said what I was going to say, so why post at all?)
I tend to think that the Bluesky crowd seems like they have their shit together as well as having a small beta explode can allow it (the web app is literally at `staging.bsky.app`, come on) and I think the AT Protocol docs identify real and probably intractable shortcomings in ActivityPub. But it doesn't somehow render ActivityPub moot if you want to use it. Go for it. It's still there.
Except that they are objectively wrong about nearly everything they talk about with regard to ActivityPub. Quoting from the FAQ:
> Account portability is the major reason why we chose to build a separate protocol.
There is a widely-accepted account portability protocol built on top of ActivityPub that multiple servers, including Mastodon and Pleroma, all support.
> We consider portability to be crucial because it protects users from sudden bans, server shutdowns, and policy disagreements.
There is nothing inherent about their protocol that solves this. The app on iOS (the Bluesky app) solves this by downloading all tweets locally, which is incredibly space-inefficient and keeps the server from... doing the job of a server (storing that data for you). Additionally, user data is still accessible and downloadable after a suspension on Mastodon and Pleroma.
> Our solution for portability requires both signed data repositories and DIDs, neither of which are easy to retrofit into ActivityPub.
There is quite literally no need for this and they absolutely could have built something that addresses these issues on top of ActivityPub. We're talking about the people who couldn't use OpenAPI, but instead built a shittier version of GraphQL while bold-faced saying 'there was no alternative'.
> a preference for domain usernames over AP’s double-@ email usernames
No need to build a separate protocol for this.
> and the goal of having large scale search and discovery (rather than the hashtag style of discovery that ActivityPub favors).
Nothing about ActivityPub, Mastodon, or the general Fediverse prohibits you from scraping it to make this happen. There are services that do this right now. Mastodon has discovery built into it, this literally completely ignores that.
ActivityPub is the standard for federation on the internet, and for allowing interoperation between social networks. That is the key here. I don't give a shit if people use Mastodon. I myself probably wouldn't use it if I wasn't running a Mastodon server.
What I do care about is whether a service is built on ActivityPub. Even if my friends want to use another social media service (maybe they're on PixelFed or whatever it's called), I can still follow them and interact with them over there while using Mastodon. You cannot do that with Bluesky.
The problems that Bluesky identified with ActivityPub objectively could've been solved by building something on top of ActivityPub, retaining the 'interoperable social network' quality of it. Email had the same problems, and instead of throwing out the email protocol, we built DMARC and co on top of it.
So no, I don't want people to use Mastodon, I don't give a shit about people using Mastodon. I literally criticized Mastodon later on in the thread. What I care about is interoperability, ease of use, and open standards, and AtProto is objectively not that.
The only thing missing there to make this better - and this is not an ActivityPub thing - is including those JSON-LD signatures in more contexts, so that you don't need to rely on the export functionality (of Mastodon) to get an archive, allowing clients to choose to keep a local copy (or nominate someone to back it up for them), and providing an upload functionality for posts (Mastodon doesn't do this, but that's also a Mastodon thing, not an ActivityPub thing).
I wouldn't have had an issue if they added extensions to ActivityPub. There are even things Mastodon refuses to add that I'd applaud people for forcing the issue on by adding extensions to support. But their choice to reinvent everything puts me off.
https://en.wikibooks.org/wiki/Serial_Programming/Modems_and_...?
But it was a different protocol that I have no experience with. It might be a crock of shit, too, but it's hard to tell from that rant.
Interestingly, the original modem AT protocol was reasonable. Then as modems became more featureful, the protocol kept getting extended, and extended, and then I stopped using POTS modems and thought I'd never have to deal with it again.
Until I built my own cellphone discovered that not only does it live on, but it's been extended even more.
ATZ
OKWe might be getting old...
Pretty interesting rant until he harped on about capitalism, and this
> Criticism of Mastodon should be looked at seriously, especially when it's from black people, as this space is overwhelmingly white.
I'm not sure why they're even using this because the docs state that the current DID stuff is all placeholders until they can find something better.
DID is there to stay, it's that the DID spec has a concept called "verification methods" https://www.w3.org/TR/did-core/#verification-methods. They are providing a verification method called "DID Placeholder (did:plc)", because none of the existing methods suited their goals. The idea is that if and when a better verification method appears, they can move to that, but that doesn't mean abandoning DIDs altogether.
Yes.
Centralization is inevitable and normal users only care if they can use it easily or not and don’t have to choose a instance or set up their own mail server, instance, or whatever.
As always with typical techies, the emotions put into this post were already running high given Bluesky itself has gotten someone extremely angry over the tiniest things.
You (probably) and I have both had a more tech oriented upbringing. Current and future urban generations seem to be more tech oriented, anyway, so perhaps we, who probably form the majority of tech consumers, will become these "normal users" we condescendingly refer to in this forum.
With sufficient investment in education or at least awareness, the "normal users" may eventually be privacy and freedom oriented.