Syncthing Usage Data
data.syncthing.net
data.syncthing.net
1) Advanced NAT traversal via UDP, which is is less efficient than TCP, but orders of magnitude better than relying on relays. More about this here: https://kastelo.net/2017/03/08/syncthing-kcp.html
2) Internal filesystem notification watch facility for near realtime sync (currently possible via an external service -> syncthing-inotify). For more see pull request https://github.com/syncthing/syncthing/pull/3986
The future looks bright!
Addendum: Have a look at the community forum https://forum.syncthing.net
But the same suggestion could be made about almost any product with a web-site these days. Indeed Syncthing is doing much better than most; they have a box at the top of the page with a one-sentence explanation/plug.
similar to dropbox
many separate sync folders
many computers to many computers syncing.
Personal storage, no data in the cloud
Source: I'm the maintainer of the Syncthing Android app.
Is there any way to get additional information about why a phone and a desktop client on the same LAN can't see each other? (or take 10+ minutes to sync?)
Disclosure: I work on the Dropbox sync engine.
Fwiw, i frequently get notifications from Syncthing that it's found a conflict it can't resolve and asks me to resolve manually.
We've likely put 25-100x more resources into solving this problem over the last 10 years, and we just have a lot more data b/c we have 100s of millions of users on lots of platforms using Dropbox with every application you can imagine. So we're able to tease out the "long tail" of weird file system and application behavior in a way that's very difficult for smaller projects. Truly durable conflict management in the face of arbitrary mutations by user applications on the filesystem ends up being a really, really hard problem to cover exhaustively. The Dropbox client handles literally hundreds of special cases.
So, yeah, I believe it is (in general) safe to assume that Dropbox is probably doing A-more-correct-thing for a complicated (and admittedly confusing) reason when it comes to sync behavior. But we're not perfect--we do still find surprises from time to time, so feel free to contact support if you see something that looks wrong!
Sounds worth finding out in detail.
The issues I had didn't result in conflicted files. Rather, after making a big change (i.e. switching git branches) some files were never updated or synced. Dropbox stopped picking up changes in the folder and eventually removed new changes once restarted.
The order of events were something along the lines of:
1) Did work on computer A that caused massive file changes (i.e. moving between git branches). 2) Moved to computer B to continue work. 3) Noticed files were old or missing on B. 4) Syncing files in some other folders worked, but nothing happened in the folder with missing files. 5) Restarted Dropbox on both machines in hope that this would trigger a fresh sync. 6) Observed files being reverted to old versions or deleted on machine A.
The end result was that Dropbox threw away the changes I had made on A and left me with the original state of B. I was able to recover the changes from a backup, so it was no big deal in the end (although it left me a bit scared I could have lost those files without noticing).
I was in contact with Dropbox support about the issue and explained in detail what I had done and what happened. I was offered help to recover the files, but since I had already done so, I just told them I didn't need any more support on the issue. I thought it might be because /proc/sys/fs/inotify/max_user_watches had a low value on one machine, so I wrote back that they might want to add back the old warning about this. However, the same problem with deleted files happened again after I had verified that this value was high enough on all machines.
I have also seen how a script run by a colleague managed to confuse Dropbox. The script was running a test which repeatedly created and deleted the same file before checking its correctness. Running the script in the Dropbox folder left him with some old version of this file and a failed test. Running the scirpt in a folder outside Dropbox left him with the correct final version of the file. He was only working on one machine.
And yes, I know it's "bad" to run scripts like this or switch git branches on top of sync software, but it happens, and it is interesting to see how different software handles these cases.
It should be noted that Dropbox usually handles these massive file changes well, so moving to Syncthing has for me been more about it being open source and the possibility to keep files on my own machines. I was just glad to see that Syncthing also handles heavy use cases gracefully.
Last I recall it didn't, which kind of soured me on it - in an ideal world every sync operation would be prefaced with "send a backup to one of my backup endpoints first".
(But Syncthing is intended for file syncing, not backup.)
The homepage recommends Syncnthing-GTK (https://github.com/syncthing/syncthing-gtk) (Cross-Platform) and SyncTrayzor (https://github.com/canton7/SyncTrayzor) (Windows only) to help with these pain points. The provide installers, more native-feeling UIs, tray icons, etc.
There are also a large number of other Community Contributions (https://docs.syncthing.net/users/contrib.html).
Photo sync on iOS without iCloud is a huge pain atm.
My backup strategy involves a NAS folder that stays up to date (i.e. also deletes photos that I delete on my phone), and which I occasionally sort into final location. After that I can free the space on my phone.
All other options (including Resilio Sync) are a add-only "sync"[1], which makes it annoying to sift through all the photos including those I already weeded out a week ago.
[0] https://docs.syncthing.net/users/faq.html#why-is-there-no-io...
[1] In theory, Resilio has a real sync, but for some reason, not for the iOS camera roll.
Keep up the great work, team!
https://docs.syncthing.net/users/strelaysrv.html
Eg that VM running a DNS adblocker, from another front page article :)
What I find most shocking is that somehow my instance (in SF via DigitalOcean, surprise) is the only relay[2] on the West coast the majority of the time. How is that? It also is one of the busiest relays by connection count showing that plenty of people in the Bay Area are selecting it likely due to latency.
Checkout more relay stats: http://relays.syncthing.net/
[0] https://github.com/kylemanna/docker-syncthing-relay
Do they generally then run their own clients to only use their own servers?
With Tor I'm convinced over the "speech should be free, I'm making the world a better place" type argument.
With syncthing I'm incredibly impressed that this optional (but incredibly useful) infrastructure is being run by people for free since it doesn't quite make the world a better place in the same way? It just saves a bunch of other people setting up their own relay/discovery server on a cheap linode/aws instance?
Impressed!
Same. The devs the contribute their time to build syncthing, the least I can do to help their (our?) cause is to spend 10-60 minutes to setup a relay and spend 5 minutes to babysit it every other month.
[1] http://relays.syncthing.net/ [2] https://docs.syncthing.net/users/strelaysrv.html#strelaysrv
The only thing missing is an encrypted external sync node so I can put it on a VPS somewhere
Didn't came around to testing gocryptfs, but if you don't need the chunking it looks like the better alternative.
https://nuetzlich.net/gocryptfs/comparison/ https://www.cryfs.org/comparison
Other than that, Syncthing is absolutely amazing!
Repeat it for 5 or 6 machines.
If any machine is switched what do you do? Start over again establishing the link with all other machines.
In general, Resilio Sync's 'key per folder' model seems to work better. I can give a key to other people and they can join the swarm without extra work on either sides. Also, it has a logical extension to encrypted-only peers: they get a derived key that can be used to sync in the swarm, but cannot decrypt the data.
If there was a stable open source program using Resilio Sync's sharing model, I would switch in a heartbeat. (There is Librevault, but it still seems to be mostly in development.)
1) There's no way to revoke access, and 2) Once the secret is out, it's out: you have to scrap it and start again from scratch.
You can always close the folder and start a new one.
Arq uses this method when encrypting blobs.
Also, you get auto-complete in the field where you select the id, so you don't actually have to type it in …
So yes, it's a bit painful (and is the price you pay for the level of security you get), but it's not quite as bad as you make out.
Basically redundant but distributed storage with N node fail tolerance. Preferably interfaceable via a "/mnt/cloud" directory so I can use it from my laptop and easily write backup scripts and store data into it.
If this existed I'd have something big to work on setting up next week.
Now, I think you could definitely set up a system that could handle the distributed data like this and allow you to access it, but I am skeptical that you could construct such a system that would be able to read/write/update/delete files on the fly and retain N-node failure tolerance at all times. You would probably have to allow the system to replicate data slowly.
There is also the problem of network splits. When they rejoin, how do they resolve conflicts and such?
<node_name>_<year>-<month>-<day>_<hr>-<min>-<sec>.tar.xz
into the filesystem and another server would be reading through everything and processing it (not in real time, but as a batch system).It's ok for things to be left over for collection at the next cycle.
Also netsplits aren't much of an issue if you're largely just doing writes and reads on seperate files. It only becomes a problem for short-term and currently processed files (that are being read and written to in real time).
This is more of a long-term cold storage system and not nodes talking to each other.
TL;DR
"It's not a bug it's a feature"
Patient: "Doctor, how do I stop the pain I get from doing this"
Doctor: "You stop doing that"You set up one introducer, where storage nodes and clients meet and exchange each other's address. After that your clients can use the storage as one big single, end-to-end encrypted, resilient space where N nodes out of M can fail without any effect on your data.
Edit: Also, can you say "this group of nodes count's as a single location" for failure protection? So if I have two locations where I store servers and then 10 other one off data collection locations can I say "I want you to treat these datacenters as one node since they are very likely to fail together if they fail".
Here it is directly from the Q&A...
" Q12: If I had 3 locations each with 5 storage nodes, could I configure the grid to ensure a file is written to each location so that I could handle all servers at a particular location going down? "
" A: Not directly. We have a wiki page and some tickets (linked from the wiki page) about this but it's deeper than it looks and we haven't come to a conclusion on how to build it.
The current system will try to distribute the shares as widely as possible, using a different pseudo-random permutation for each file, but it is completely unaware of server properties like "location". If you have more free servers than shares, it will only put one share on any given server, but you might wind up with more shares in one location than the others.
For example, if you have 15 servers in three locations A:1/2/3/4/5, B:6/7/8/9/10, C:11/12/13/14/15, and use the default 3-of-10 encoding, your worst case is winding up with shares on 1/2/3/4/5/6/7/8/9/10, and not use location C at all. The most likely case is that you'll wind up with 3 or 4 shares in each location, but there's nothing in the system to enforce that: it's just shuffling all the servers into a ring, starting at 0, and assigning shares to servers around and around the ring until all the shares have a home.
The possible distributions of shares into locations (A, B, C) are:
(3, 3, 4) 1500 (2, 4, 4) 750 (2, 3, 5) 600 (1, 4, 5) 150 (0, 5, 5) 3 sum = 3003
So you've got a 50% chance of the ideal distribution, and a 1/1000 chance of the worst-case distribution. "
But RAID and Syncthing aren't substitutes for backups, for the same reason: no restore if your data get hosed.
RW times would be slower but storage would be bigger and more redundant.
I am using Resilio Sync since its beta version, but they have added a lot of restrictions since then and made some functions only available for paying users.
I am concerned about ease of use (setup and everyday usage) on Windows and Android.
We used Resilio Sync (aka Bittorrent Sync aka BT Sync) for a while for the Tron project ( https://reddit.com/r/TronScript ). It worked great for our use-case: distributing a large number of files to a large number of nodes who required read-only access. Anytime I changed a file on the master node it would blast out to everyone else while simultaneously preventing any other node in the swarm from propagating changes.
Unfortunately with all the restrictions they introduced, it stopped working reliably. Internally they introduced a hard-coded peer cap of something like 32, while our swarm was already over 500.
We switched to Syncthing and it's been working well. There is still one major issue and that's that anyone can mark their copy of the folder as "read-only" (formerly called "Master"), and then they will attempt to propagate changes out to the rest of the swarm. So, there's that. But strictly on a distributed file syncing level, it works great.
Here's the thread in question: https://forum.resilio.com/topic/30525-please-please-please-p...
https://forum.resilio.com/topic/41520-max-number-of-users-fo...
How do you deal with that?
Fallback is just to provide a static .exe on the mirrors and leave the Resilio Sync node running alongside it.
In the meantime, I have checked Syncthing forums and I have found this topic: https://forum.syncthing.net/t/why-im-moving-back-to-btsync-r... It is exactly complaining about the ease of use on Windows and Android, saying there is no native GUI, etc. From this description it seems, that Syncthing is not there, yet.
See the Community Contributions page (https://docs.syncthing.net/users/contrib.html) for the full list.
Works quite well, with a few gotchas (ie make sure you don't sync cache files, session files etc).
Good software!
This page indicates that your usage should be around 80 megs (with the gui closed): https://forum.syncthing.net/t/huge-ram-usage-2gb-on-synology... and that an unofficial NAS-specific build had a problem with high ram usage.
https://github.com/syncthing/syncthing/issues/468#issuecomme...
Since that bug was closed with no solution I'm not too hopeful but maybe I should try it again.
I also reject any sort of statistics though, so I'm probably not helping.
Don't upset your Network Administrators!
On the negative side, your stats page doesn't have a link back to your main page, so you won't even know how many folk had a look based on this HN post. It's a good idea to ensure that all your pages have a link back to your main page.
People aren't really supposed to link to it directly as a means of introducing new people to Syncthing...
For example, they could like to your Forum page (which doesn't have a link to your main page). Or to your Github page (which has a link to the Forum, but not to the main page). Or your Docs page (which...well, you get the idea).
Sites that have witty and inventive 404 pages sometimes get those publicised. You really never know what someone will choose to submit.
Have a look at librevault [1] instead, it is written in C++ [2].
I wish it could be made more efficient somehow.