A Storage Crisis
blogs.harvard.edu
blogs.harvard.edu
I was shocked to come back and find the RAID 100% full and only a few GB copied. It took me a few minutes to figure out what happened.
Maddeningly, the USB drive has a feature that it goes to sleep after 15 minutes or so, even if Ubuntu is actively using it to write files! Maybe the drivers on Windows do something that keep it awake, but on Linux it just goes to sleep in the middle of being used.
Now here's the thing, the drive DOES realize it has more to do and wake back up. But this sleep-wake causes a USB disconnect and reconnect. Which causes Ubuntu to unmount and remount the drive.
Now here's the problem, perhaps because the backup program is still making the the original mount point "busy", Ubuntu doesn't re-mount the media to the path. Instead it, gets mounted at "/media/path-1".
Since Linux uses regular folders as mount points, the old mount point at "/media/path" becomes a valid folder on the local disk. So the backup program keeps going, but now it's filling up the local disk.
I haven't found a solution for this problem that will allow me to complete a backup (or even complete a long-duration manual copy).
chattr +i /mnt/foo
This makes the mount point immutable when nothing is mounted there.You could shuck the drive and connect it directly or find a reliable USB enclosure.
Putting the disk in /etc/fstab by uuid should keep it mounted to the same directory at least, but I'll be surprised if the backup software properly handles the potential errors well. The following link has an example to remount the drive with udev when it goes away, and if the backup software isn't running as root you can mount with uid,gid set to the backup user (or chown an ext filesystem) but make the mountpoint directory 0700 owned by root to prevent the backup from writing to it while unmounted.
https://blog.backslasher.net/automatically-mounting-usb-driv...
Smartctl on a timer didn’t work. Touch on a timer didn’t work. I can’t recall if hdparm or some other tuning of sleep and head parking settings worked.
Taking the drive out of the enclosure worked. It got so much quieter too.
Just to be clear, the drive does go to sleep when it's actually idle. Not while it's being used, though. IOW, they work as expected.
What distro/kernel version are you on? I mostly use Debian stable, which means a pretty old kernel. I wonder if it could be a regression in more recent ones?
I also use the drives with FreeBSD and have never had an issue there, either.
YMMV but I've found rsync to handle this kind of fuckery well.
Re; the issue of external drives spinning down causing a remount at a different location, I've never had that exact experience before, but I have had drives spin down more frequently than I'd like, so my answer to that was to just disable the spindown entirely. (There's articles about various ways to do that easily findable on your favorite search engine if you're curious how to do it.) Once disabled, if I find that I still want some form of spin-down for the drive, I install one of the handy available daemons that allows a much more easily configurable form of this feature, and then set a longer delay than the drive's default. (This can also be done by configuring the drive itself to a longer delay, but it's more of a pain than just editing a simple text configuration file for a daemon.)
sustained data transfers crush those little controllers, so spend more on a better enclosure, actively cool the one you have, or just connect the bare hard drive directly to your PC if that is an option.
I've had 2 USB-3 to mSATA enclosures fail permanently after 5 minutes of sustained data transfers.
I use several USB storage devices, enclosures on Linux everyday including using separate USB ssd caches.
Can you check with your storage device managing software on Windows on whether it has the sleep option or spin down set? I remember seeing one in the Seagate SW on mac although it has an OS level option to put HDD to sleep when not in use.
If that doesn't work you can try using hdparm or sdparm to modify the power management and spin down timer. If everything else fails then disable USB autosuspend in the kernel boot option.
The other problem is how the backup software reacts to the possible failure modes.
A backup tool should identify the backup destination in a reliable way and nountpoints is clearly not a reliable way to identify a backup destination
[0] rsync -ahv --progress -e 'ssh -p2222' /data/0/Pictures bu.mydomain.tld:/mnt/data/backups/ --info=progress2 #(assumes cert based auth set up)
As an example; Many folks commenting above appear to have never experienced a problem with hard drives spinning down during a backup, or it causing a remount at a different mount point. This leads to thinking it might not be an operating system issue, but rather something strange about that specific external hard drive.
My solution since then was to work through a soft link to a directory inside the mounted drive.
Whatever the failure mode, when the storage disappears, even if the users do a “mkdir -p” (which seems to be the underlying issue you are describing)
When ever I take pictures on a trip or vacation, i go through them at the end of each day, delete most of them and keep maybe 2 or 3 max, beautify them and the rest goes into the bin. No matter how long the trip, i try to only keep at max the 15 best Pictures. All filler no killer. That way I am comfortable to show others pictures of a trip without boring them and I also like to look at them from time to time as I know those are the best moments.
At the end of each year we create a calendar with photo collages of 4 to 6 pictures per month. The calendar goes to relatives and we create a photo book from the print outs. That is what gets archived.
This... this is a truly universal first-world problem. The average person did not need data retention policies until the past decade or so, and now it feels like I need a retention policy for literally every physical or digital object that enters the premises just to maintain my sanity.
Adam: ~15TiB.
I recently got into film photography and the 36 pictures I get from the lab after a weekend or two make me so much more happier than 6 frames per second that my Fuji does.
I made the mistake of building a NAS for one at one point but it turns out that you need to be really diligent about documenting the setup as rebuilding a raid5 array when it inevitably fails can be very difficult if you don't remember what the actual configuration was.
The crux of the issue is that they need massive storage capacity but whatever the solution is needs to be able to manage merging disjoint Adobe Lightroom catalogs while de-duplicating and without the possibility of data loss, all while needing basically no maintenance because these people tend not to be incredibly computer literate.
This is actually the reason why I bought a NAS for my primary storage at home. Just because if it crashes, I will have support or I can just buy another one with the same configuration and make it work.
If it crashes, the last thing I want to do is have to fiddle with settings around my data.
Photoshelter looks to have a $50/month unlimited cloud storage plan[0], which seems pretty cheap if they're going to really take advantage of the "unlimited" aspect. For a professional photographer $50/month to not have to worry about carrying around this stuff seems like a deal (and less if they don't have that much to store).
I've never heard of Photoshelter, but it sounds like you upload high-resolution final images that are ready to view and print--it doesn't seem all that different from Smugmug or Flickr. Most photographers want to archive the source files and edits (project files). I see you can technically upload RAW files, but unless that integrates very well with your photo management software, I'm not sure if it's of much use. Lightroom databases can get relatively large since it includes caches and thumbnails. So you'd need a separate system for backing up those that ties it back to the RAW files.
Dropbox might work, but I'm not sure how approachable "smart sync" is and their pricing tiers don't seem like a great fit.
My guess is a well-loved open source project that wraps around b2 or AWS might be the best fit. But, honestly, a pile of hard drives will probably win out. You make your investment upfront and it's a simple enough mental-model. Unfortunately, they tend to die and with large data sizes photographers don't always go for redundancy.
My wife has a bunch of photos/videos for a small business and I have to admit we haven't been diligent in backing up. It's a lot of data to save the larger sources. And we also went down the NAS route, mainly for access convenience though.
However that being said, I use a combination of NAS of which I have about 40TB of pure photos (yes a lot of images). I've been dabbling with LTO to backup images, but you need to know what you are doing with tape. It's not a simple drop type opperation as a hard drive is. You can get a lot of shoeshining if you don't do things properly.
So for long term easy access, I recommend to others to look into 4-Bay NAS units (at a minimum) for future flexibility. QNAP or Synology to aid the config and use a blob storage service to store the data (if upload speed permits). I used to use Backblaze but even with basic B2 storage, I found it more expensive than Google Cloud storage (coldline).
But with anything your use case, preparedness to learn and financial situation will dictate what you end up doing.
Of course you also need a really good, sharp lens to take advantage of such a high-res sensor. Any flaws in the lens or your technique or settings will show up more dramatically in a deep crop. So it's not going to save you money on a shorter focal length lens, it's just going to save weight.
ZFS mirror pools is where it's at.
I can pop out any one drive of the pool and take it to any random unrelated machine and restore everything.
Not to mention some tools to help with all the duplication we had on different HDD or USB Sticks.
See this blog post and comment from nearly four years ago for a better description. [1]
And here’s Backblaze’s support page on the same, last updated a month ago (June 2021) indicating that this behavior remains. [2]
Any backup solution that doesn’t, at a minimum, follow the 3-2-1 rule will cause more instances of regret in the future. If the data is of value, it needs better and constant care.
[1]: https://mjtsai.com/blog/2014/05/22/what-backblaze-doesnt-bac...
[2]: https://help.backblaze.com/hc/en-us/articles/217665498-Why-h...
These days, openzfs native encryption is the best of all worlds, I think.
In short, ZFS works fine on capable hardware with plenty of memory. It offers more amenities than a system based on mdraid/lvm/filesystem_of_choice. The latter combination can work on lower-spec hardware with less memory where ZFS would bog down or not work at all.
As a tangentially related aside I wonder why bringing up (potential) downsides of ZFS tend to lead to heated discussions where nothing like this happens when the same is said about e.g. hardware raid/$file_system or a modular stack like mdraid/lvm/$file_system.
re: tangent; I wouldn't really call my response heated, I've just ran into workloads with ZFS where limiting the ARC cache solved problems. I've also ran into frustrations with ZFS that there really aren't easy solutions to (The slab allocator on FreeBSD not playing well with ZFS on this particular servers workload, so having to change the memory allocation system being used to one that doubles CPU usage. Not having ZFSs extended ACLs exposed on Linux, meaning migrating one of our FreeBSD systems to linux will require serious effort)
An old motherboard with 4G of RAM is not hard to come by.
Is this to guard against house fires?
- If you want "personal storage" (no programmatic access) the options are generally around $5 per TB/month - but that's a lot of drag-and-dropping and praying the connection doesn't die during transfer.
- If you want object storage (think S3) the cheapest is $5 per TB/month, the average is $10 per TB/month, and the "high end" is $20 per TB/month, with extra costs for bandwidth ranging from $10 to $120 per TB/month.
Honestly, just store less crap. Marie Condo your digital life. Does it not spark joy? Have you not looked at it in the last 6 months? Does it not serve a useful purpose, such as tax records? Ditch it.
Even if you want to keep some pictures/video, either print a copy in original-quality, or compress/downscale it. I took a 3 minute video on my phone and it's 425MB. And it was still grainy! I used to download two-hour movies that were 725MB and looked like a DVD! errrrrrr... I mean, I heard about a guy that did that.
Very similar to just buying a few more items at the grocery than I normally would. My time is so incredibly important that it's not worth $20 in wasted items to have to potentially drive back there if I've missed something.
If I don't have as much time as possible to spend on working/innovating/personal health and family, my income source will dry up and then I'm in dire straits ;)
It is kinda like how the old Yahoo-organized-like-a-library died out in favor of the Google-just-search-everything approach. Unless you're triaging as you go (which still takes tons of time; do people really go through 1,000 photos after their 3 week holiday to Iceland and keep just the 15-20 best?) it is easy to build up an insurmountable backlog where the only viable option is just "delete everything".
I actually did that with my old collection of ripped-from-CD mp3s and just resigned myself to streaming anything I really cared about. But you can't exactly do that with "family photos".
Next, I might need to build a little blueray changer to burn massive quantities of optical media.
Because the grain is not from video compression but because the camera sensor wasn't receiving enough light so it boosted the ISO to compensate which causes grain. Also the fact that phones have to live encode video which is very non optimal while movies are shot uncompressed and then a powerful CPU can spend as long as it needs for a perfect compression.
I wrote a very simple video editor: video playback, some scrubbing controls on the arrow keys to jump around, and space bar to mark / unmark segments to keep (capturing timestamps). Run ffmpeg to cut the bits out then move source video to a different directory, ready to be deleted.
This cut down on my video volume 100x, while I still retained a few old commute videos, as a keepsake for that time in my life.
Removing those tax-forms isn't going to help you at all in freeing space, they are probably compressed XML or just PDFs. Yet, for example, that single slideshow for granma's 90th birthday gobbles up 95% of all the diskspace all your presentations combined use.
It is far more effective to hunt those large files and go through those only. I have largest.sh and baobab for that, but there are numerous other (GUI) tools for this.
Less fun, though. But after seeing my wife spend an entire afternoon sifting through her files (look! I already deleted 22 GB) and me going: did you see the 'Photoshoot 2012-04-01 waalkant (kopie)' full of Raws? That is 210GB. Can it go? done! Then I realised we need better tools and educate people a little here.
Another issue with storing all these photos for your whole life is of course that once you die, what happens? Who will take charge of all these photos? Who will continue to pay? Who will clean up and take ownership of it? Or will they just be lost once you are dead and all this money for what?
The third issue is of course that the more random unsorted pictures we have the less we want to look at it because 90% are bad and only 10% are good. It's important to clean up just after taking the photos.
On a long enough timescale we all die and are probably forgotten. I've actually spent a decent amount of time resisting the urge to just delete all my old photos (I have things going back to early high school - taken on a Sony Mavica that used floppy disks).
There's a very liberating feeling to the idea of just having no history at all - probably similar to the appeal of the idea of having all your works crumble to dust when you die, which is also something I've seen people refer to at different times.
But, since I haven't done that yet, and since managing an on-site server or disks is it's own stress (i.e. eventually my house will burn down or flood and I'll lose everything...so what was the point?), then paying that $x/month is basically the price I pay to defer having to really think about it.
The problem of "what to keep" isn't even digital-specific. In my youth I probably bought two dozen disposable cameras, had the pictures printed, even paid for the CD-ROMs when they became available. Where are those pictures now? Who knows! That's why I like Marie's message: it's not about throwing away junk, but keeping what you love close to you.
I'm definitely not immune and I find that the satisfaction I get from my photo collection is inversely proportional to the size of my collection. At this point, I'm just lugging around this huge mass of data "just in case". There's no way I will ever have time to sort my images. There are probably 2000 wedding images alone, let alone the tens of thousands thousands of random snapshots that may or may not be something I ever care to see again.
At this point, I would almost prefer a smart solution opt-in solution similar to what Google photos provides for smart albums: "We found these images that would be good long-term. Save them indefinitely?"
I reference mine quite often, usually for a purpose I didn't originally think of. The name of that cool restaurant in a city you visit rarely? How about how many double rainbows did you get last year in June? Pinning down dates from a previous pet. Or the funny pet photo with a squirrel in it? Where were you on March 2nd when your credit card got charged for $1202.20? When did you meet your new significant other 3 years ago or so? What trail head was it that you saw a bear? What is the serial number, VIN, or similar for just about anything valuable you own... or owned. How old where you when you won that race?
The cost of a store a photo for life isn't much, seems silly to spend hours and hours trying to delete them, doubly so because mistakes will be made. Just tag them so you can find what you want. I was surprised how much tagging helped. Grand mothers poured over ever damn photo I tagged with the grand kids. Starting collections of wild animals we've seen by state. etc. When bored I find it rather fun to relive a vacation or hike. Even my kid seems to quite enjoy "visiting" places I've been.
Sure 95-99% of my photos don't get viewed, ever. So? Why waste man hours on useless photos, just make sure you are organized enough to find the 1-5% that you do care about.
Think of all the grandmas you would indirectly be providing more grand kid photos to!
For me that was the selling point of Google Photos. I switched in 2016 maybe - prior to that using Photos.app - and it was mind blowing. Thanks to machine learning, if you didn't think you wanted to tag waterfalls at the time of capture, it would still let you find them (I have 12 waterfall photos).
Unless you are a professional photographer, why waste man hours organising them when some algorithms and machine learning can do a good enough job?
It's very different from when a hobbyist takes photos just for their personal fun
It has been running with zero maintenance (other than occasional partial restores) since late 2018.
I transitioned off of AWS cloud storage when they raised the prices.
I’m not sure if Wasabi is still the cheapest and fastest, but they have been great to deal with. And Arq is an excellent set-it-and-forget-it encrypted cloud backup app.
I also run a server with a 16tb RAID 1 array and a set of local backup drives. Sadly it is almost full, and the volume of data makes it a hassle to upgrade (not to mention the cost).
I’ve found standard 1Gig Ethernet to be just barely fast enough for editing photos over the local network. However, for my own sanity, I usually do the initial editing on a local drive before sending the files to the server (and from there to the cloud backup).
The AWS or Azure price is high, but it's a scalable price, a real price.
I'm currently not backing that data up, though. It's not critical, and the chances of AWS losing it are low enough that I'm not too worried. The biggest risk would be myself accidentally deleting it.
One thing I have decided for sure, though, is that archiving lots of small files is awful. Much better to wrap them up in a tarfile, uncompressed.
I personally wouldn't recommend Amazon Drive as a backup destination to anyone.
Backblaze has been offering their 'unlimited data' for over a decade [1], and it hasn't ended badly for them yet.
It's fully sustainable because they only lose money on a few customers (like that one customer storing 430TB for $6/month [2]), while most customers use much less storage than that, so the service remains profitable overall.
The soft limit of its 'unlimited' comes from it being a personal-backup mirroring service and not a fully-external cloud storage, so the service only fits certain use cases.
[1] https://www.backblaze.com/blog/all-in-on-unlimited-backup/
[2] https://www.reddit.com/r/IAmA/comments/b6lbew/were_the_backb...
Also their upload speed is capped so you can't just upload at 1GB/s or max out your connection to ingest data into their system, like something you can do with S3.
Glacier Deep Archive is still the cheapest thing really at $1/TB, but the retrieval times and also egress data charges are a big catch.
Glacier/Deep is good, but with the 180-day minimum object lifetime, you want to be sure that the data is ready to go into the archive before pushing it there. (You can use tiered storage, but then you're storing all data in standard S3 for 30 days before it gets into Glacier, and that one month of storage in standard S3 will cost you the same as 10-months of Glacier storage.)
The web site is hosted through AWS S3 and the Cloudfront CDN. The "web side" is carefully optimized for storage and transfer costs. (Glad webp is good-to-go, I am hopeful JPEG XL rolls out fast.) With Glacer/Deep I can afford to archive the Camera RAW, superresolution PNGs and the other assets that go into the print sides just in case I lose my workstation and my storage server.
One of my (to be implemented) backup strategies was to use Glacier Deep Archive as last-resort recovery, and just stick yearly tarballs (e.g. 2018, 2019, 2020) in there. That should save me a bit on retrieval requests as well.
That might be prohibitive if you have a huge NAS to back up, but for moderately large amounts of data (say 5 TBs or less), it seems pretty reasonable.
For me Backblaze Personal is anything but a backup service.
https://www.backblaze.com/blog/subscription-changes-for-comp...
Now, a person might have a big backlog of data. This guy has 7TB of photos. If you start uploading, you’ll get caught up eventually! If you want to get caught up faster, you just need to find a faster pipe and park your computer/HDs on it for a little while. For me, it was a friend with FiOS. For many professionals, it might be their office.
The point is, once you’re through the backlog, 10mbps is probably enough to maintain. Again: for most people.
So far I've been able to get by with retail HDDs; I have a 5TB drive at home and another that we keep off-site and refresh a couple times a year. This seems to be a sustainable strategy for me, as affordable (~$100) portable HDDs are growing at a rate that is faster than my storage needs. I don't know if it would be cheaper to pay for iCloud, but for whatever reason I feel safer having two HDDs with my media than trusting Apple's (or anyone else's) cloud.
I'd suggest paying for iCloud just for the ability to automagically backup your iPhone(s). You can exclude Photos from being backed up if you prefer.
I use iCloud Photo Library as well on a larger plan, but the backup feature alone is definitely well worth $0.99.
Prior to that I had this monsterous, 4U server that I had to maintain, update, debug, not to mention build and move from house to house, for 8 years. At the time it seemed like a good move, but over time, as my time got shorter to work on personal products, I started to look for other ways to solve the storage problem.
I much prefer the Synology NAS to the 4U custom built system. It's about as close to a toaster-style appliance as I can think of. If it died tomorrow, I'd buy another one (then restore from my Glacier backups).
SSD has 0 benefits over spinning rust. and is still more expensive.
I can buy decent spinning rust 10TB disk anywhere, if i want that in ssd's, id need 5 drives, or start hinting for enterprise grade equipment.
So for me for raw storage (if you need more than a couple of TB's) spinningg rust is still the winner.
(I use B2 for cloud storage backup since it’s a lot cheaper than S3 at US$20 for 4 TB, but that still is ~$105/year more than local)
And that's with bad pricing! Last year, 4TB was $120 and 6TB was ~$150.
Building out a NAS for $1000 ($600 in 4x hard drives, $400 for other components) is very reasonable. Last year that was 4x6TB == 12TB storage + 12TB redundancy, but this year prices are worse so you "only" get 8TB + 8TB redundancy.
$400 can afford a Synology or various NAS devices. It can also afford a new desktop that you can install FreeNAS or whatever onto.
-----------
Eventually, when the 8TB is not enough, just buy a new HBA card and shove 4x more hard drives in there for a 2nd storage on the same NAS. Maybe 8TB x 4 == 16TB usable + 16TB redundancy.
Except no need to copy everything over, just keep the old 8TB cluster working, and just start writing to the new 16TB storage.
If you have 5TB of photos, chances are you're not looking through 5TB of photos all that often, and S3's IA storage is very appealing at ~$65/mo.
If you only want a cold storage backup, glacier will keep them for you at $20/mo. If you need storage that's only accessed once or twice a year (and you don't mind waiting a bit to get your file) you can pay less than $5/mo with Glacier Deep Archive.
The cost of S3 that you quoted is for _nine nines_ of durability and milliseconds latency. If you don't need that, you can pay for far cheaper storage that better-matches your needs, with the convenience of never needing to buy/transfer/replace/sell drives.
> That doesn’t include power costs, but over a year you get $990 to spend on that and other things from the cost difference vs. S3
If you spend ~days each year worrying about storage and paying for ever-larger disks (and spending time selling your old ones, for whatever you can get), chances are the few hundred dollars you might come out ahead doing it yourself isn't really worth it.
You can still beat that, of course, but it's assuming you have time and skills to do so and don't mind spending that time playing sysadmin.
Moving up to larger drives every 6 months is insane. Maybe every 3-4 years. Depending on your actual storage requirements of course. Not sure why OP's requirement is to have "few devices".
Local storage is always cheaper than cloud storage unless you're doing it really wrong. Fully (3-2-1) backed up local storage less so, but if you do it smart it doesn't have to cost an arm and a leg.
Offline storage is dangerous. Readers go obsolete. Media degrades. Data sets go missing. Protocols and storage formats become obsolete. Migrating data from one online storage to another while upgrading hardware solves a lot of these issues, the only exception being storage format for straight file copies between systems.
The more pix there are, the less likely it is you'll ever even open any of these.
> The more pix there are, the less likely it is you'll ever even open any of these.
Well, there are also future generations to think about, but their interest will fall off too (until you reach the genealogical profile level, which maxes out at a few portrait-type pictures of any regular individual).
It's pretty essential to aggressively curate and organize data like this.
The way I think about it is that you need relatively small "SAVE THIS" drive/disc. Sure, storage is getting cheaper, but you don't want to burden someone with some massive storage array, and you can't count on whatever software you used to tag to still work in someone else's hands.
And yet Hardrive price / GB hasn't fallen much at all. From 2013 to 2021 now, when it was $0.035 to $0.04 /GB. Mostly because much HDD patter density has stalled.
I turned an old desktop PC into an 8-bay storage machine with a PCIe SATA card and now have all the flexibility I want, without the price penalty for Synology hardware.
The best place to find hard drives is https://diskprices.com/ which scrapes Amazon and a few other sites to show TB per £/$. It's almost always cheaper to buy a retail packaged external disk and extract the hard drive.
Note: Some hard drives once shucked from external cases won't turn on when attached to regular SATA power adapters due to mismatches in power connector specification. Easily solved by using a SATA power to Molex adapter, and connecting that to a Molex to SATA power adapter (SATA->Molex->SATA) as Molex does not support the 3.3V pin which is the issue.
Also, what shitty phone takes 108MP photos? That's guaranteed to be some stupid Android phone gimmick. There's no way having that many pixels with a teeny optical path and a teeny sensor is useful. I'd only want above 100MP on a medium format sensor.
And if the answer is to buy 2 NASes, what happens if something physically destroys the location. Fire & theft isn't all that uncommon over a lifetime. Or if you (probably a more likely) accidentally format or delete its contents.
I like to backup my UHD blurays losslessly and stream them from my PC via plex (~100mbps). I'm thinking local storage via a NAS is probably the cheapest option here, right?
https://github.com/danielgtaylor/jpeg-archive
Tried a lot of combinations of quality, and settled for -q medium --max 75% --accurate
Anything higher makes no sense/difference, even if you print it at A3. And my images are 8-18-24Mpix.
i'm not a pro-foto but somewhat of an old pre-print school - Your level of pickiness may vary :/
Effectively it is like 4-8 times down - Total went from 150Gb into 30Gb - that's about 40,000 images.
Which still doesn't solve the OP's or anyone's problem of too-easy-producible digital "assets".. but is more manageable.
Easy but your data isn't yours: sync your data to GDrive or Apple or whatever, and sync a NAS to that.
A little harder but still doable: get a Hetzner and set that up as your storage, set up your own access, sync to local NAS. A Hetz is also really useful for running a load of other services, so for 50 bucks or so it seems pretty reasonable.
You could just buy a huge disk and run a server in your house, but it gets annoying in various ways. Kids unplug the power, it creates heat, maybe noise, multiple disks end up needing management, that kind of thing.
So we buy sensors with more megapixels that just save more noise. Then we need to buy more storage capacity, more bandwidth, more processing, more battery etc.
Since the cameras have been hitting closer and closer to physical limits but something still needs to be sold...
It's like people buying clothes from the mall and taking them straight to a storage building. They don't have enough space at home. And they won't have time to wear so many clothes anyway.
A good starting point for me is to not keep everything I shoot. I have often noticed that every 50-100 photos I shoot I only want to keep 5-10. Rest are bad/weird shots, pics I’m not interested in etc.
I never faced this when I used film.
I haven't had to use it yet, but I check the files from time to time, and so far I'm satisfied with the service.
B2 worked great and the pricing is indeed excellent, but the volume of raw data to backup over broadband is just not realistic.
I think it would be better to physically mail a drive for archiving raw footage, or use tap drives.
The newer phones always have a higher resolution camera so file sizes keep increasing.
I would love a simple easy to use tape drive that just works. Something I could backup files to for 20-25 years.
You don't keep the RAW files from your DSLRs? That'll bring the total size up pretty quick.
Check out storj.io
Take whatever the backups are, say 1TB, split it into 12 pieces, add as many pieces as the user wants (say 6 on average), then you can recover with any 12. When you add another 1TB, add another 18 peers.
Monitor your peers, only trust the ones with a track record of successful challenges, and of course let you white list peers you trust (like friends and family). Of course the perfect challenge is just a restore, but some checksome of a range of a blob could be useful as well, and consume less bandwidth.
Drop peers that are unreliable, or ask for too much bandwidth for their restores.
Encrypt the files before adding reed solomon, use a unique encryption key for each peer that stores 1/Nth of your backup set.
That way you can "pay" for your storage by just adding that much more local disk space to trade with your peers.
1. Everything is harder to work with: you have to deal with less reliable networks and storage, computers which aren't on all of the time, very slow uplinks, etc. Since these aren't professionally managed systems, too, you have less visibility — did that node just drop offline because the hard drive failed, losing everything, Comcast is having a bad day, the owner just bought a new one and wiped the old one without unenrolling it, or because it rebooted and is almost back up?
2. People are selfish: I don't want Netflix getting slow because you decided to retrieve your data, I'll complain if I hit a storage limit on my computer due to your stored backup data, etc. This forces you do deal with things like traffic shaping and storage rebalancing more aggressively and those are hard problems to get a popular balance on. Consider, for example, what happens when someone uses your service and it goes well but then they run out of space and need to clear some up in a hurry (_especially_ if they put on their cowboy hat and just delete a bunch of large files because they know they aren't the only ones with a copy).
3. The solutions to the previous problems make the cost problem worse: storing more copies can avoid some of the problems but then you need to figure out how to get the network to support, say, 5 copies instead of 2-3.
4. Consider what happens the first time the police bust someone for a major crime and their data is backed up on your computer. Not many people are enthusiastic about going into court to prove the negative assertion that they didn't have a decryption key.
Only peering with trusted systems avoids some of these issues but not all and the _big_ problem is that the upper bound for how much this service is worth is basically the cost of iCloud/Dropbox/S3/Backblaze/etc. The savings you can get between the fixed operational costs and what those services charge is probably not enough to support development.
No, hoarding is a behavior which may or may not, in any particular case, be a symptom of one or another mental illness.