How to self-host all of Bluesky except the AppView (for now)
alice.bsky.sh
alice.bsky.sh
How do I secure the webserver and the data? Where is the data on my disk? How to backup and restore? High availability?
There might be detailed documentation somewhere, or I can even read the code. But these are the important things an open source software should tell its users right off the bat.
1: https://github.com/bluesky-social/pds/blob/main/README.md
Meanwhile mastodon is incredibly easy to self host w/ relays.
[1] https://docs.pleroma.social/backend/installation/otp_en/
Here's the kicker, the install script could be called from a Dockerfile pretty easily, no? Sure, there might be things to sort out, but it doesn't seem unreasonable.
I agree having a docker image is super handy and can be quick to try, as well as update, and put into a larger self-hosted environment how you need.
- https://raw.githubusercontent.com/bluesky-social/pds/main/in... has all the expected outside docker-compose setup, you can read it through in like 5 minutes
- Heavy-duty part of the setup is running https://raw.githubusercontent.com/bluesky-social/pds/main/co... which you should be familiar with
I guess the shellscript is for people who want a one-line install, which I wouldn't do myself either, but I guess some people prefer.
It's also over complicated, like, it even tries to handle race condition of multiple apt processes! What kind of environment do they expect the users have? As the project become more popular, the script will need to handle more edge cases. Let's see if it is still a 5 minutes read one year later.
> I guess the shellscript is for people who want a one-line install, which I wouldn't do myself either, but I guess some people prefer.
This is the problem in lots open source projects -- providing a one-liner installer and bragging about how easy the initial setup is, without an easy path for long term maintenance. Give it some time, many happy users of the one-liner will be unhappy when they encounter issues.
Like hosting an application in the cloud, you also will never stop improving how you self-host.
If there's questions lacking about a software package, it's often could be reflected in your self-hosting environment too.
Running this type of an installer is excellent to quickly introduce yourself to any technology - to then start learning about how you want to run it long term.
The questions expressed above are not new. How SRE's solve it today also can be different and more complicated than needed.
Easy answer - if they have an install script, it's getting run inside a VM, or Docker which itself is a baseline backup and HA automatically if needed.
If generally anything is run inside of a self-hosted hypervisor like Proxmox, it can be setup to automatically backup, mirror, HA as-is, while you figure out what you want. This includes running docker inside a Proxmox VM, there is not a big performance hit anymore for doing this for things that are largely idle most of the time.
There is a big difference between SaaS, PaaS, and IaaS. It's easier and easier to get the benefits from all three by being willing to build up the foundation instead of pointing at the gaps in each package for not filling it for you.
It's encouraging to see things becoming more possible :)
I feel like that it's kind of out of the scope from an article describing the steps for application/protocol specific infrastructure. You need to look for resources, guides and such for general self-hosting instead, somewhere else.
For example, if you use TrueScale NAS/unraid/proxmos or whatever for local self-hosting, you'd setup those things via those platforms. If you use Kubernetes/Nomad/Incus/Containers, you'd solve those things via that tooling.
One thing I have found with many open source/selfhostable projects is just how much running them yourself can vary. It can go from a simple compose file with everything included to having to dig for obscure services and piece together how they all form the whole.
For example, I recently looked into self hosting Zotero. It is so under documented and complex that there is almost no way one could self host that (even for just one user) without that being ones job. So one needs to make a distinction between something being open source and being feasible to use/maintain.
In the end I gave up with Zotero. Even though it could have replaced Obsidian Notes, Calibre and Syncthing all at once for me.
I don't self-host my PDS yet because there is no migration path back yet (but there will be). Though maybe I'll just yolo one day and do it anyways.
I've come across this a lot too. But what I've found is that it mostly applies to open source projects that offer a hosted paid version, so it kind of makes sense they'll make the experience slightly worse than it could be (consciously or subconsciously), as it pushes people to their hosted solution. I don't particularly like it though.
Doesn't seem to be the case for Zotero specifically, but your comment reminded me that I've noticed this more often lately.
Just for the benefit for anyone that want to go through this rabbit hole. You cannot selfhost Zotero. In theory but in practice it is no feasible. If you find their free storage limiting then store them on webdav (all clients support that).
zotero team explicity said that they don't see this as a priority [1] and with the release of zotero 7 and transition it is not realistic to think they will ever do.
[1] https://github.com/zotero/dataserver/issues/105#issuecomment...
I've been thinking a lot about the relay, though. 4.5 terabytes is, well... A lot, to say the least. If Bluesky grows 100x larger, running a relay will become pretty insanely expensive. I guess if the Bluesky organization remains fairly neutral about the relay part, it's not a huge deal, but:
- It always eventually becomes hard to stay neutral. Eventually someone will get mad at something going through your network that isn't just obvious network abuse like SPAM.
- It seems like drinking from the firehouse itself will eventually become expensive. Will it be possible for something this high bandwidth to remain freely-accessible?
I love that people even has the choice, so much better than not even being able to.
Is plc.directory a single point of failure for BlueSky users who want to take advantage of the benefits of a did:plc? And if so, is that a permanent thing or down the road will there be multiple interoperating did:plc directories?
The backstory to PLC is that we picked up the DID standard and looked for an existing registry-method that would satisfy requirements¹. None of them really did. We then surveyed mechanisms for decentralized operation: DHTs, open blockchains, permissioned blockchains, and federated databases. Of them, the two blockchain variants seemed perhaps promising, but still premature since (as of 2022) you there's cost variability due to load and in some cases bad transaction latency (eg 10 minutes).
We decided the best decision was to create PLC, which matches all of the requirements except for longterm meta governance. The way we designed it was to make the registry mechanics transferrable to a different protocol in the future, so that if for instance we decided (say) a DHT was suitable (it's not) we'd be able to use the same identifiers but change resolution and mutations to a new process. Then we started talking to other SMEs to get their take.
Ultimately the solution that's gotten the most favorable response has been setting up an ICANN-style independent organization to operate it. This can be joined with a couple of interesting systems, such as mirrors which tail a certificate-transparency-style audit log, and which could even serve as transaction witnesses to indicate when the core registry might be rejecting updates ("write censorship").
What can I say, some things take time and stakeholder-building. Look up the history of DNS and Network Solutions Inc for a bit of a wild ride that people have forgotten about. One other thing I should point out is that the DID spec enables multiple registry methods. Atproto currently supports did:web, and if other methods show up which satisfy the requirements then we are interested.
¹ Secure against manipulation by the registry operators, longterm meta governance, highly available, reasonable transaction latency, reliably low cost that's not dogged by token speculation, low ecological impact.
*Key Event Receipt Infrastructure
This is a fun party trick in some sense, but also a real meaningful feature in another. If I ever decide to move from steveklabnik.com to steve.klabnik.com, a thing I have been considering for a few years, my stuff on @proto/Bluesky will be one of the only services that doesn't have the issue that's kept me from pulling the trigger: updating the entire world that that's where I am now.
https://www.w3.org/TR/did-core/#dfn-verifiable-data-registry
DIDs delegate trust and authority to a data registry, in exactly the same way that DNS delegates trust and authority to ~ICANN.
The system model is exactly the same. The difference is only in the properties of the authoritative entity.
That said, you're not wrong that a registry is a registry.
The ongoing WordPress fiasco is a good sign of what happens when you set up an independent organization too soon. You won't have the people or the commitments from those people to maintain that independence, so the independent thing ends up not being able to do anything to protect the thing that was supposed to be independent from the commercial interests looking to exploit it.
What time window does it cover? A rolling N day window? Everything since year dot?
Can it be pruned? e.g. only data of accounts followed or messages interacted with
This site tries but has limits:
* https://bsky.jazco.dev/stats
They broke 14 million yesterday and it seems to be snowballing now since the election:
* https://bsky.app/profile/jaz.bsky.social/post/3laetwhztdk2x
(And all of this is a fork of my friend's Samuel's blog, https://mozzius.dev, see https://github.com/mozzius/mozzius.dev)
Does that... bureaucracy of documentation not infuriate anyone else or is it just me. I guess I'll try and reset my password to bluesky website, assuming it's this .app one, but then it's asking me to maybe select a provider ... of my password.
Does whoemever made this user experience not have enough emotional intelligence realize how infuriating it is?
It's asking what the host of your data is. If you're not running your own server, then the default value of Bluesky itself is the correct one.
(...also, the title, as the original has the caveat)
(I can't tell if Dan has an alert set up on his handle or whether he just sees everything, but hopefully that works :))
<link rel="canonical" href="https://alice.bsky.sh"/>
HN will replace submission links with the canonical link if it's found.