Rclone syncs your files to cloud storage
rclone.org
rclone.org
Just seems like a cool dude. Glad to see he's able to do it full time now, and hope it's still as fulfilling.
EDIT: Thanks to archive.org I found it[0]. Even more wholesome than I remembered:
> Rclone is a pure open source for love-not-money project. However I’ve had requests for a donation page and coding rclone does take me away from something else I love - my wonderful wife.
> So if you would like to send a donation, I will use it to buy flowers (and other pretty things) for her which will make her very happy.
[0]: https://web.archive.org/web/20180624105005/https://rclone.or...
It can also e2e encrypt locations, so everything you put into the mounted drive is getting written encrypted to Dropbox folder /foo, for example. Nice, because Dropbox and other providers like s3 don't have native e2e support yet
All in all, Rclone is great! One of those tools that is good to have on hand, and solves so many usecases
Thank you
And i've yet to see a pi-native dropbox client for doing so.
PS: i actually do sync dropbox to my main workstation but selectively do so. The lion's share of my dropbox is stuff i don't need locally so don't bother to sync it. The rclone approach gives me easy access to the whole dropbox, when needed, from anywhere in my intranet, and i can use my local file manager instead of the dropbox web interface.
nohup /usr/bin/rclone serve ftp --addr MYIP:2121 $PWD &>/dev/null &
No configuration needed (beyond rclone's per-cloud-storage-account config (noting that serving local dirs this way does not require any cloud storage config)) and some variation of that can be added to crontab like: @reboot /usr/bin/sleep 30 && ...the above command...
Noting that $PWD can be a cloud drive identifier (part of the rclone config) so it can proxy a remote cloud service this same way. So, for example: rclone serve ftp --addr MYIP:2121 mydropbox:
assuming "mydropbox" is the locally-configured name for your rclone dropbox connection, that will serve your whole dropbox.- startup only when the network is up
- proper logging
- automatic restarts on failure
- optional protection for your ssh keys and other data if there's a breach (refer to `systemd-analyze security`)
Run:
$ systemctl --user edit --full --force rclone-ftp.service
this opens a text editor; paste these lines: [Unit]
After=network-online.target
Wants=network-online.target
[Install]
WantedBy=default.target
[Service]
ExecStart=/usr/bin/rclone --your-flags /directory
and then enable and start the service: $ systemctl --user enable --now rclone-ftpsystemd is far from perfect, and Poettering is radically anti-user. But it's the best we got and it serves us well
What does that mean?
It's not immediately apparent what this means—does it use FUSE, 9p, a driver, or some other mechanism to convert FS calls into API calls?
EDIT: it's FUSE.
FWIW, i've been using Insync on Linux since it went online (because Google never released a Linux-native client). Aside from one massive screw-up on their part about 8 or 10 years ago (where they automatically converted all of my 100+ gdocs-format files to MS office and deleted the originals), i've not had any issues with them. (In that one particular case the deleted gdocs were all in the trash bin, so could be recovered. Nothing was lost, it was just a huge pain in the butt.)
Why. Just why. How does that shit ever happen in a public release?
I've been using it to great effect for over 10 years on a daily basis.
This pain has always stopped me using Unison whenever I give it another go (and it's been like this since, what, 2005? with no sign of them stabilising the protocol over major versions.)
For client side encryption they have a whole encrypted S3 client and everything. (https://docs.aws.amazon.com/amazon-s3-encryption-client/late...)
I use it on my desktop and laptop to mount Google drives. The problem on the laptop is that the OS sees the drive as local, and Rclone doesn't timeout on network errors. So if you are not connected to wifi and an application tries to read/write to the drive, it will hang forever. This results in most of the UI locking up under XFCE for example, if you have a Thunar window open.
I shorten it to prevent lockups like you are describing.
https://forum.rclone.org/t/how-to-get-rclone-mount-to-issue-...
The timeout param is listed as "If a transfer has started but then becomes idle for this long it is considered broken and disconnected". This seems to be only for file transfers in progress.
I traced it once, and Rclone gets a "temporary DNS failure" error once the network is down, but just keeps retrying.
Most cloud space providers don't show you how much space each folder and subfolder actually occupies. Enter rclone ncdu.
This is top-tier tooling right here.
I assume you are pulling from Google photos? If so, then I think the only way to get original quality is to use takeout?
(1) https://github.com/gilesknap/gphotos-sync#warning-google-api...
It's really stupid that the API doesn't support this. For now I'll stick to regular takeout archives. Syncthing directly from the phone might be a better option for regular backups.
I'd been using sshfs for some years until I learned that rclone can mount remotes to the file system, and I've been using that happily since then.
https://rclone.org/commands/rclone_mount/
> at present SSHFS does not have any active, regular contributors, and there are a number of known issues
The rclone backend means I can backup anywhere.
Duplicacy is slower and also only free for personal use.
Fortunately there is https://github.com/creativeprojects/resticprofile to solve that problem.
I hate such an approach when someone assumes that whatever happens their opinion is the best one in the world and everyone else is wrong.
I do not want to encrypt my backups because I know what I am doing and I have very, very good reasons for that.
Restic could allow a --do-not-encrypt switch and backup by default.
The arrogance of the devs is astonishing and this si why I will not use Restic and i regret it very much because this is a very good backup solution.
Try Borg.
I'm using duplicity because none of these can sync to a rsync server that I run on my NAS.
As for remote backups - I use ssh with Borg and it works fine. If this is a NAS you can probably enable ssh (if it is not enabled already).
BTW for my remote backups I do encrypt them but this is a choice the author of Borg left open.
There are other issues with Borg such the use of local timestamps (naive date format, no timezone) instead of a full ISO8601 string, and the lack of capacity to ask whether a backup is completed (which is a nightmare for monitoring) because the registry is locked during a backup and you cannot query it.
Restic for various machines to hourly backup on a local server. Rclone to sync those backups to S3 daily.
Keeps the read/write ops lower, restores faster, and limits to a single machine with S3 access.
This is currently the only way to get Proton Drive on Linux.
When I last checked it doesn't use the AWS SDK (or the Go version is limited). Anyway, it isn't able to use all settings in .aws/config.
But it is kind of understandable that it doesn't support all backend features because it's a multifunctional tool.
Also the documentation is full of warnings of unmaintained features (like caching) and experimental features. Which is a fair warning but they don't specifically tell you the limitations.
It's also great for direct cloud-to-cloud transfers if you have lots of data and a crappy home connection. Put clone on a server with good networking, run it under tmux, and your computer doesn't even have be on for the thing to run.
tmux isn't strictly necessary:
nohup /usr/bin/rclone ... &>/dev/null &
then log out and it'll keep running.It does if you redirect the output to a file other than /dev/null, but in my experience checking on the progress is irrelevant - it's done when it's done.
The instructions look mostly like a reverse engineering effort that can fail anytime for any reason. Of course you'll get some notification, but still seems a thing to avoid in principle.
I’ve been using rclone for 3 years without issues. Across many different backends.
Man I wish I lived in a universe where providers actually supported tooling. You're completely delusional tho.
I sincerely doubt most tools will be supported by most providers (by usage), however, and I question the evaluation of tools by this metric. Support is generally a fairly bad metric by which to evaluate tooling—generally speaking, if you care which client is consuming an API someone has done something horrifically wrong.
ssh user@rsync.net rclone blah blah ...
... so you don't even need to install rclone for some use-cases - like backing up an S3 bucket to your rsync.net account:To help avoid getting anyone's hopes up: rclone does not do automatic two-way sync. In rclone parlance, "sync" means to either pull all files from a remote and make a local copy match that, or do the opposite: push a local dir to a remote and update the remote to match the local one. Note that "match" means "delete anything in the target which is not in the source."
Oooh, nice. That's not _quite_ the same as real-time two-way sync, but i guess it's the next best thing.
I want to have a local directory in my computer that bi-directionally syncs to both Google Drive and DropBox as well as to my Synology NAS. There will be two Google Drive accounts (one for me and one for my wife) and one Dropbox account. I also have two Linux workstations running Ubuntu 22.04 and two Macbooks where I want the local directory to exist.
Basically anytime we put anything in either the cloud drives or the NAS or the local directories in the respective computers, it should get synced to all the other destinations.
There could be other stuff in the cloud drives or the NAS that I would just have the tool ignore. The sync should just be for that one specific folder. I do not want to deal with separate local folders for each cloud storage solutions.
Is this feasible? I am okay to run cron tasks in any of the computers; preferably the Linux ones.
Turns out rclone can present an FTP server proxy that is backed by any storage you want. I put it in a little Docker container and now I don't have to spend $400 on a new scanner.
Running migrations on the server side is faster and more reliable. I can monitor transforms and transfers in tmux and then run quality checks when it’s complete
And having a vm lets me filter and transform the data during the migration. Eg pruning files , pruning git repos, compressing images .
There’s a framework waiting to be made that’s like grub but for personal data warehouse grooming like this
I used the normal ways of accessing AWS, GitLab, etc., but `rclone` made it easy to access the less-technical services.
The simplest thing you can probably do is use rclone to copy your most important files to a B2 bucket. Enable Object-lock on the B2 bucket to ensure no deletion just to be safe. You can then run rclone on server and from your devices with cron jobs to archive your important stuff. This is not a proper backup as I said, if you rename files or delete unwanted stuff it wont leave on the backup bucket but it's usable for stuff like photos and the like, anything you don't want to lose.
(I lied, simplest thing is actually probably just copying to an external hard drive, but I find having rclone cron jobs much more practical)
my gripe with it is to be able to use and sync from my smartphone. at least on ios, there is no robust tool that allows it afaik. there is cryptcloudviewer (https://github.com/lithium0003/ccViewer) which has not been updated since 2020.
appreciate any suggestions or more details of your workflow for these scenarios.
If there is a data corruption bug in ZFS, it will propagate to your remote and corrupt data there.
I hope you have something else in place besides those two tools.
My fallbacks are:
- an external drive that I connect once a year and just rsync-dump everything
- for important files, a separate box where I have borg/borgmatic [1] in deduplication mode installed; this is updated once in a while
Just curious: Do you have any reason to believe that such a data corruption bug is likely in ZFS? It seems like saying that ext4 could have a bug and you should also store stuff on NTFS, just in case (which I think does not make sense..).
[1]: https://www.reddit.com/r/zfs/comments/85aa7s/comment/dvw55u3...
https://github.com/openzfs/zfs/issues/7401
Corresponding HN discussion at the time: https://news.ycombinator.com/item?id=16797644
I think it makes sense and thank you for the sensible reminder.
Most filesystems from x years ago do not cope well with current trends towards increasing file numbers. There are robust filesystems that can deal with Petabytes, but most have a tough time with Googolplexian filenumbers. I speak of all the git directories, or venv folders for my 100+ projects that require all their unique dependencies (a single venv is usually 400 to 800k files), or the 5000+ npm packages that are needed to build a simple website. or the local GIS datastores, split over hundreds and thousands of individual files.
Yes, I may not need to back those up. But I want to keep my projects together, sorted in folders, and not split by file system or backup requirements. This means sooner or later I need something like rsync to back things up somewhere. However, rsync and colleagues will need to build a directory and file tree and compare hashes for individual files. This takes time. A usual rsync scan on my laptop (ssd) with 1.6 Million files takes about 5 Minutes.
With ZFS, this is history and this is the major benefit to me. With block suballocation [1] it has no problems with a high number of small files (see a list of all filesystems that support BS here [2]). And: I don't have to mess with the file level. I can create snapshots and they will transfer immediately, incrementally and replicate everything offsite, without me having to deal with the myriad requirements of Volumes, Filesystems and higher level software (etc.).
If I really need ext4 or xfs (e.g.), I create a ZFS volume and format it with any filesystem I want, with all features of ZFS still available (compression, encryption, incremental replication, deduplication (if you want it)).
Yes, perhaps this has nothing to do with the _cloud_ (e.g. rsync.net offers zfs snapshot storage). But the post was about rclone, which is what my reply was pointed to.
[1]: https://en.wikipedia.org/wiki/Block_suballocation
[2]: https://en.wikipedia.org/wiki/Comparison_of_file_systems
When I tried to back up that and my other data to a hard drive, Google takeout consistently failed, downloading consistently failed. I went back and forth with support for months with no solution.
Finally I found rclone and was done in a couple days.
I get extra angry about situations like this. I have paid Google many thousands of dollars over the years for their services. I just assume that one of the fundamentals of a company like that would be handling large files even if it’s for a fairly small percentage of customers. I get that their feature set is limited. I kind of get that they don’t provide much support.
But for the many compromises I make with a company that size, I feel like semi-competent engineering should be one of the few benefits I get besides low price and high availability.
-- Charlie Munger, Poor Charlie's Almanack ch. 11, The Psychology of Human Misjudgment (principle #1)
But that's only one possible reasoning.
It is something providers do to keep regulators off their back. They are not going to put money into making it work well.
Giving away a feature that is a competitor's cash cow could weaken that competitor's stranglehold on customers you want to connect with or drive the competitor out of business.
At a particular point in time, increasing market share could be more important for a company than making a profit.
It could be culturally uncommon to charge for some essential feature (say bank accounts), but without that feature you wouldn't be able to upsell customers on other features (credit cards, mortgages).
Ad funded companies give away features and entire services for free in order to be more attractive for ad customers.
Of course, if a feature is bad for a company in every conceivable direct and indirect way, even in combination with other features, now and in the future, under any and all circumstances, it would not introduce that feature at all.
Introducing a completely broken feature is unlikely to make much sense. Introducing lots of low quality features that make the product feel flaky but "full featured" could make some sense unfortunately.
Yes, if buyers think that far. Consumers do not. SOme businesses may, but its not a major consideration because the people making the decision will probably have moved on by the time a swtich is needed.
> Ad funded companies give away features and entire services for free in order to be more attractive for ad customers.
Yes, but that means those services do make a profit.
The same applies to the banking example but they can make money off free accounts as well.
> Introducing a completely broken feature is unlikely to make much sense. Introducing lots of low quality features that make the product feel flaky but "full featured" could make some sense unfortunately.
The latter is similar to what I am suggesting here. Being able to export data will probably satisfy most customers who want to do so, even if it does not actually work well for everyone. It will also mollify regulators if it works for most people. If they can say "data export works well for 95%" of our customers regulators are likely to conclude that that is sufficient not to impede competition.
They absolutely do. It's hard not to, because migrating data is the very first thing you have to think about when switching services. It's also a prominent part of FAQs and service documentation.
>Yes, but that means those services do make a profit.
Of course. What I said is that not every single feature can be profitable on its own. Obviously it has to be beneficial in some indirect way.
I have a primary and backup NAS at home, and a separate archive copy updated weekly offsite locally. Then Glacier is my "there was a nuclear bomb and my entire city is completely gone" solution.
And in that case, assuming I had survived, I would be willing to pay most anything to get my data back.
I've already got backblaze (behind rclone) set up for my backups so adding a glacier archive would be quite easy. Thanks!
It also incurs a delay of at least 6-12 hours before the first byte of retrieval occurs, when you need to restore. So there are tradeoffs for the price.
> a delay of at least 6-12 hours before the first byte of retrieval occurs
Huh. I wonder how it actually works behind the curtains. Do hey actually use HDDs to store all of that, or maybe is there some physical work involved in these 12 hours to retrieve a tape-drive from the archive…
Most of the time the poster links to a specific news item at least (which is also not okay without a specific headline) but sometimes the link just points to the home page.
Regardless, it has been mentioned before.
EDIT: Just to be clear, I'm not again an occasional mention of a project. Hacker news has allowed me to find some real gems, but most of those posts are new features added to the software, not just a generic link to the homepage.
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
Like you mentioned, some real gems are found, and for me are often found in the comments. So, these reposts are helpful.
It’s like how many posts were made for “What do you use for note taking?” Some gems are are in the comments.
This is discussed in the site docs, it's fine to get news from HN but HN is not really about news.
I guess I will just be more aggressive in hiding posts that I didn't want to waste my time reading.
The only source of truth that is really reliable is a password memorized in your brain. Because keyfiles that can decrypt data can be corrupted, that is to say, they rely on a secondary data integrity check. Like the information memorized in your brain.