71 TiB DIY NAS Based on ZFS on Linux
louwrentius.com
louwrentius.com
I know the allure of raid 7 is strong, however if a disk were to fail on a LUM thats 24 disks wide, the chances of a rebuild is pretty low. (I know you've since changed that) the rebuild time is around 60-90 hours with no load.
(we use 14 disk raid6 LUNs for the balance of performance vs safety, however we've come pretty close to loosing it all.)
You'll get better performance if you have many smaller vdevs. You'll also compartmentalize the risk of multiple disk failure.
a raid 0+6 of four LUNs of 5 disks will have much greater performance, and will rebuild much faster, without risking too much. However you will loose more space
[nsivo@news.ycombinator.com ~]$ zpool status -v
pool: arc
state: ONLINE
scan: scrub repaired 0 in 0h43m with 0 errors on Wed Aug 13 20:14:53 2014
config:
NAME STATE READ WRITE CKSUM
arc ONLINE 0 0 0
mirror-0 ONLINE 0 0 0
gpt/arc0 ONLINE 0 0 0
gpt/arc1 ONLINE 0 0 0
errors: No known data errors
pool: tank
state: ONLINE
status: One or more devices has experienced an error resulting in data
corruption. Applications may be affected.
action: Restore the file in question if possible. Otherwise restore the
entire pool from backup.
see: http://illumos.org/msg/ZFS-8000-8A
scan: scrub repaired 0 in 5h3m with 0 errors on Tue Aug 26 00:19:05 2014
config:
NAME STATE READ WRITE CKSUM
tank ONLINE 0 0 0
raidz2-0 ONLINE 0 0 0
gpt/tank0 ONLINE 0 0 0
gpt/tank1 ONLINE 0 0 0
gpt/tank2 ONLINE 0 0 0
gpt/tank3 ONLINE 0 0 0
logs
mirror-1 ONLINE 0 0 0
aacd5p1 ONLINE 0 0 0
aacd6p1 ONLINE 0 0 0
errors: Permanent errors have been detected in the following files:
/usr/nginx/var/log/nginx-access-appengine.log.0
I can't delete that file or recover it from backup, and don't really trust the pool or our disk controller anymore. So we're moving to a new server, again. Good times.I imagine this would be disastrous for someone who's filled a 24 disk array with data, and has to purchase a duplicate to safely restore to.
Do you have snapshots? Can you read that file in a snapshot?
Corruption happens. A good filesystem needs to cope.
You very well may have some flaky hardware which is causing the issue, in which case, it probably needs to be fixed sooner than later. And maybe with some magic, you can get the pool into a happy state without rebuilding.
As mentioned in my other response I had random bit flips in memory. After I replaced my broken NIC crashes stopped. But I got a nasty surprise when I removed a filesystem - kernel panic due to failed assertion, and then kernel panic for the same assertion every single time I tried to import the pool.
I suspect it possibly could be resolved somehow without losing data, but decided to simply restore it since my backup was relatively fresh.
I suspect you might have something similar to what I experienced, and even though you use ECC memory I doubt it would help much when corruption is happening on a bus.
zpool clear tank
ZFS does actually come with a decent set of tools for fixing and clearing faults - which I'm starting to realise isn't as widely known (possibly because people rarely run into a situation where they need it?)Anecdotally speaking; I've ran the same pool for years on consumer hardware and that pool has even survived a failing SATA controller that was randomly dropping disks during heavy IO and causing kernel panics. So I've lost count of the number times ZFS has saved my storage pool from complete data loss where other file systems would have just given up.
In my case the problem turned out to be a broken WiFi card (PCI). After I replaced it no issues for 2 years now.
zpool clear tank
Source: http://docs.oracle.com/cd/E19253-01/816-5166/6mbb1kqod/index... (unfortunately there's no tag ID's to hash to, so you'll have to [ctrl]+[F] to find the specific section on 'clear')There is very rarely a good reason for doing asymmetrical vdevs in a single ZFS pool.
I approve of his choice of HBA (based on the SAS2008 chip). For those looking for even bigger setups, I maintain a list of controllers at http://blog.zorinaq.com/?e=10 where you can even find 16-, 24- and 32-port controllers. 8-port ones are the most cost effective as of today.
I have been a happy ZFS user for 9 years, both personally and professionally. Using snapshots, running periodical scrubs, having dealt with a dozen of drive failures over the years. Everything works beautifully - it's a pleasure to admin ZFS.
I've been eying this - http://www.supermicro.com/products/motherboard/Atom/X10/A1SA..., but the project is firmly in the "one day when I have a spare $5k" basket.
ZFS is production ready on Linux now?
Illumos is the de-facto upstream of open-zfs and has many features not found in oracle solaris. [1]
So it is well worth checking out at least one of those:
- SmartOS, which is a fantastic Cloud-Hypervisor [2]
- OmniOS, a classic server distribution, probably best for what OP is building. [3]
- NexentaStor, a storage appliance - but i have not tried it yet [4]
I can definitely recommend using ZFS on a Illumos system as it has much better integration with things like kstats and fault management architecture that I would miss on any other OS.
As yet I don't have any hands-on experience with OmniOS, but I'd tend to agree that it seems the best direction to go looking in, for a conventional OpenSolaris-derived OS.
OmniOS on the other hand is pretty solid and with pkgsrc (default package system on SmartOS and NetBSD) you get nearly 13k packages of open-source software.
My List is also missing a few other commercial players that use and contribute to Illumos:
- Delphix, uses ZFS snapshots to do fancy things with big databases
- Pluribus has "Netvisor OS" that is some SDN Router/Networking appliance [3]
And then there is the long tail of small hobby distributions:
- Tribblix, "Retro with modern features" [4]
- DilOS, which has a focus on Xen and Sparc [5]
and probably more that I forgot and don't know.
It is a very nice community and lots of interesting technology gets developed. Our company migrated nearly all servers from Linux to SmartOS in the last 2 years and we could not be happier.
[1] http://wiki.openindiana.org/oi/Hipster
[2] http://pkgsrc.joyent.com/installing.html#installing-on-illum...
Technically it seems quite impressive, and Illumos being more or less the ZFS upstream is appealing. The distribution situation has been more confusing, though. OpenIndiana seemed to have a community, which is why I was first looking into that option, but it does indeed not seem to be very active. That's one attraction of FreeBSD & Debian, that they're community supported, and I can with fairly high confidence expect they'll be around in 5 or 10 years supporting the distribution. It's less clear to me how to wade into Illumos in a way that mitigates risk of the upstream going away. I think Illumos itself fits that description, with a community that will be around, but does any individual distribution? There are clearly resources behind OmniOS and SmartOS, which gives some confidence, but they are also quite heavily concentrated resources (one company drives each, and it's not impossible that they could change focus and deemphasize development of the public distribution).
On the other hand they have been - and still are - very good at building a community around it. There are multiple people and organisations building their own SmartOS images/derivates ([1], [2]). Pull-requests on github come in from a diverse enough group that I think it would survive even without joyent.
For OmniOS I don't have enough insight to make that judgement. They have a very clear and detailed release plan and recently hired some good people just to work on OmniOS. These things definitly add some confidence but of course is not a guarantee forever. But if that would be needed one can always buy support from them.
I recommend it, but only after you've done enough research to feel comfortable with all of ways to configure it and understand the internals from a high level. For a personal setup though, you can just install it and forget it.
Also I invested in 10GbE, at first was only server to server using some SFP+ NICs and DAC leads I picked up off eBay for between £10 and £80. But have since bought a couple of Dell 5524 switches and a PCIe enclosure for my Mac and having 650MB/s read/write to my server from the desktop certainly reduces the waiting around :)
Two 2.5" 160GB internal SSDs for ZIL and l2arc.
12x4TB drives.
$2000 + Chelsio NIC ($230) plus 12 x 4TB WD Red ($2000)
So for $4250. I get 48TB and a better performing system in 2U.
His system costs 6012 EUR, or about 7900 USD at today's exchange rate. For $8500 I can fit 96TB in the same space, have 192GB of RAM, 16 cores @ 2.4GHz, 740GB of SSD for ZIL and l2arc.
And redundant power.
I'm not sure what a "RAID-like cache" is. I know what redundancy is, and I know how ZFS works.
16GB is likely adequate for his storage needs. He doesn't have that many clients accessing the storage.
and I could keep two complete copies of the data on two machines, each with redundant power, etc.
ZFS uses integrated software RAID (in the zpool layer) for technical reasons, not merely to "shift administration into the zfs command set". Resilvering, checksumming, scrubs are all a unified concept that are file aware, not merely block aware as nearly every other RAID. The implications of this are massive and if you don't understand, please don't write an article on file systems.
For various reasons, the snapshots and volume management are more usable due to the integrated design and CoW, and also as pillars for ZFS send/receive.
The "write hole" piece is bullshit. ZFS is an atomic file system. It has no vulnerability to a write hole.
The "file system check" piece is bullshit. Again, ZFS is an atomic file system. The ZIL is played back on a crash to catch up sync writes. A scrub is not necessary after a hard crash.
Quite frankly, for any modern RAID you probably should be using ZFS unless you are a skilled systems engineer and are balancing some kind of trade off (stacking on higher level FS like gluster/ceph, fail in place, object storage, etc). You should even use ZFS on single drive systems for checksumming and CoW, and the new possibilities for system management with concepts like boot environments that let you roll back failed upgrades.
Hardware RAID controller quality isn't spectacular, and the author clearly has not looked at the drivers to dish out such bad advice. You want FOSS kernel programmers handling as much of your data integrity as possible, not corporate firmware and driver developers that cycle entirely every 2 years (LSI/Avago). And there's effectively one vendor left, LSI/Avago, that makes the RAID controllers used in the enterprise.
ZFS is production grade on Linux. Btrfs will be ready in 2 years, said everyone since it's inception and every 2 years thereafter. It's a pretty risky option right now, but when it's ready it delivers the same features the author tries bizarrely to dismiss in his article. ZFS is the best and safest route for RAID storage in 2014 and will remain such for at least "2 years".
That's a catastrophe waiting to happen. EMC recommends no more than 8 disks in a parity array group (ZFS recommends no more than 9), specifically because larger arrays cause longer rebuilds, which are more likely to trigger a URE, in which case your SOL.
Isn't there still a danger of typing something silly, an 'rsync --delete' type of command that bricks the whole thing? Rather than the one Borg Cube I think I would prefer to add disks/machines with disks on an as required basis, or keep the last system in the loft 'just in case'.
I am still amazed though at how much disk people 'need'. Even in uncompressed 4K resolution it would take a long time to watch 71TiB of video.
Another cool feature: you can do a send a diff between 2 snapshots as a file stream (over e.g. SSH) and replay it on another machine. So that's like using rsync except it's actually a stable snapshot (rsync can't fully handle files that changes while it's copying them) and you don't have to re-checksum every part of every changed file. So you can use that for DR.
Transparent compression is also great (as is deduping -- if you have tons of spare memory!)
BTRFS will be able to do the same thing... one day.
It already can. btrfs sub snap does snapshots (and they work the same way), btrfs send/receive do the diffs, supports lzo/zlib compression and deduplication
Aside all of the positive things already mentioned about ZFS, one of the biggest selling points for me is just how easy it is to administrate. The zfs and zpool command line utilities are child's play to use where as btrfs felt a little less intuitive (to me at least).
I really wanted to like Btrfs as it would have been nice to have a native Linux solution that would work on root partitions (my theory being that Btrfs snapshots could provide restore points before running system updates on ArchLinux) but sadly it proved to be more hassle than benefit for me.
As a plug, I've been using CrashPlanPro for offsite backup, I've been very happy with their service.
# zpool list
NAME SIZE ALLOC FREE CAP DEDUP HEALTH ALTROOT
lanayru 73.3T 53.7T 19.6T 73% 1.00x ONLINE -
# zpool status
pool: lanayru
state: ONLINE
status: The pool is formatted using a legacy on-disk format. The pool can
still be used, but some features are unavailable.
action: Upgrade the pool using 'zpool upgrade'. Once this is done, the
pool will no longer be accessible on software that does not support
feature flags.
scan: scrub repaired 1.58M in 201h47m with 0 errors on Fri Feb 21 13:37:24 2014
config:
NAME STATE READ WRITE CKSUM
lanayru ONLINE 0 0 0
raidz2-0 ONLINE 0 0 0
ata-SAMSUNG_HD204UI_______________-part1 ONLINE 0 0 0
ata-SAMSUNG_HD204UI_______________-part1 ONLINE 0 0 0
ata-SAMSUNG_HD204UI_______________-part1 ONLINE 0 0 0
ata-SAMSUNG_HD204UI_______________-part1 ONLINE 0 0 0
ata-SAMSUNG_HD204UI_______________-part1 ONLINE 0 0 0
raidz2-1 ONLINE 0 0 0
ata-SAMSUNG_HD204UI_______________-part1 ONLINE 0 0 0
ata-SAMSUNG_HD204UI_______________-part1 ONLINE 0 0 0
ata-SAMSUNG_HD204UI_______________-part1 ONLINE 0 0 0
ata-SAMSUNG_HD204UI_______________-part1 ONLINE 0 0 0
ata-SAMSUNG_HD204UI_______________-part1 ONLINE 0 0 0
raidz2-2 ONLINE 0 0 0
ata-ST2000DL004_HD204UI_______________-part1 ONLINE 0 0 0
ata-ST2000DL004_HD204UI_______________-part1 ONLINE 0 0 0
ata-ST2000DL004_HD204UI_______________-part1 ONLINE 0 0 0
ata-ST2000DL004_HD204UI_______________-part1 ONLINE 0 0 0
ata-ST2000DL004_HD204UI_______________-part1 ONLINE 0 0 0
raidz2-3 ONLINE 0 0 0
ata-WDC_WD30EFRX-68AX9N0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68AX9N0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68AX9N0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68AX9N0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68AX9N0_WD-____________-part1 ONLINE 0 0 0
raidz2-5 ONLINE 0 0 0
ata-WDC_WD30EFRX-68EUZN0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68EUZN0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68EUZN0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68EUZN0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68EUZN0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68EUZN0_WD-____________-part1 ONLINE 0 0 0
raidz2-6 ONLINE 0 0 0
ata-WDC_WD30EFRX-68EUZN0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68EUZN0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68EUZN0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68EUZN0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68EUZN0_WD-____________-part1 ONLINE 0 0 0
ata-WDC_WD30EFRX-68EUZN0_WD-____________-part1 ONLINE 0 0 0
logs
mirror-4 ONLINE 0 0 0
ata-KINGSTON_SV300S37A60G_________________-part1 ONLINE 0 0 0
ata-KINGSTON_SV300S37A60G_________________-part1 ONLINE 0 0 0
cache
ata-KINGSTON_SV300S37A60G_________________-part2 ONLINE 0 0 0
ata-KINGSTON_SV300S37A60G_________________-part2 ONLINE 0 0 0
errors: No known data errors
If anyone has any questions, I'd be happy to try to answer.How much does it cost in terms of power to have something of that size running 24/7 (I'm assuming it is?)?
Now you've set it up, does it require much of your time to manage monthly?
What sorts of things are you using the space for?
Lastly, how much did it cost in parts/time to get it all up and running, does it work out cheaper rolling your own or would you have been cheaper using a cloud solution instead, or does that just not meet your use-case/needs?
After setting it up, I wouldn't say that it requires any time to manage. Getting it all set up just right, with SMART alerts and capacity warnings and backups and snapshots, all of which I roll myself with various shell scripts, took a long time. Besides that initial investment, the only "management" I have to do is respond to any SMART alerts, add more vdevs as the pool fills up, and manage my files as I would on any other filesystem.
I use the space for just about everything. Lots of backups. I use the local storage on all of my workstations as sort of "scratch" space and anything that matters in the long term is stored on the server. The highest density stuff is, of course, media: I have high definition video and photographs from my digital SLRs, I have tons of DV video obtained as part of an analog video digitization workflow, media rips, and downloaded videos. (I even have a couple of terabytes used up by a youtube-dl script that downloads all of my channel subscriptions for offline viewing, and that's something I doubt anyone would do unless they had so many terabytes free.) I keep copies of large datasets that seem important (old Usenet archives, IMDB, full Wikipedia dumps). I keep copies of old software to go with my old hardware collection. I have almost every file I've ever created on any computer in my entire life, with the exception of a handful of Zip and floppy disks that rotted away before I could preserve them, but that is only a few hundred gigabytes. I scan every paper document that I can (my largest format scanner is 12x18 inches, so anything larger than that is waiting for new hardware to arrive someday), so almost all of my mail and legal documents are on there too.
(I had a dream the other night that someone got access to this machine and deleted everything. Worst. Nightmare. Ever.)
A cloud solution would not have met my use case, since one of the primary needs I have is to be self-sufficient in terms of storing my own data, and I also want immediate local access to a lot of the things on there. I do use various cloud solutions, but only for backup, never as primary storage.
Rolling it myself was definitely cheaper than any out-of-the-box hardware solution I've seen. The computer itself is a Supermicro board with some Xeon middle-of-the-range chip and a ton of RAM, and an LSI SAS card. Connected to the SAS card are two 24-bay SAS expander chassis, which contain the drives, which are all SATA.
I'd say that building something like this would cost you maybe about 4000USD, not counting the cost of the drives. The drives were all between $90 and $120 when I bought them, but of course capacity eventually started going up for the same price over time, so let's say another 3500USD for the drives.
I kind of want to set something like this up while spending the least amount of money. I am comfortable enough with Debian/Linux to do most things, but I have never managed anything like this. In the end I want to end up with somewhere relatively safe to store data pretty much in the same way you are, I just do not need 70TiB, and I have no experience with ZFS/hardware stuff/storage.
As far as ZFS on Linux, it still has its wrinkles. I use it because, like you, I'm comfortable with Debian, and I didn't want to maintain a foreign system just for my data storage, and I still wanted to use the machine for other things too. (I actually started with zfs-fuse, before ZFS on Linux was an option.)
So, I don't know. If you just want a box to store stuff on, you might want to just look into FreeNAS, which is a FreeBSD distribution that makes it very easy to set up a storage appliance based on ZFS. FreeBSD's ZFS implementation is generally considered production-ready, so you avoid some ZFS on Linux wrinkles, too.
So, I'd recommend checking out the FreeNAS website, and maybe also http://www.reddit.com/r/datahoarder/ for ideas/other opinions. I do a lot of things in weird idiosyncratic ways, so I'm not sure I'd recommend anyone do it exactly how I have. :)
Thanks for the link, I will take a look there.
Plus FreeBSD has a lot of good documentation (and the forums have proven a good resource in the past too) - so you're never going it alone (and obviously you have the usual mailing groups and IRC channels on Freenode).
While I do run quite a few Debian (amongst other Linux) I honestly find my FreeBSD server to be the most enjoyable / least painful platform to administrate. Obviously that's just personal preference, but I would definitely recommend trying FreeBSD to anyone considering ZFS.
I can understand your preference though. I'm not a fan of apt much personally, but pacman is one of the reasons I've stuck with ArchLinux over the years - despite it's faults :)
You don't need ZFS for this, as cool as it is.
This is probably the number one thing that excites people about ZFS over any other solution, and it's something that isn't really easily implemented on a standard RAID + standard filesystem arrangement, since this sort of functionality depends on the filesystem knowing about the underlying disk arrangement.
"ZFS uses its end-to-end checksums to detect and correct silent data corruption. If a disk returns bad data transiently, ZFS will detect it and retry the read. If the disk is part of a mirror or RAID-Z group, ZFS will both detect and correct the error: it will use the checksum to determine which copy is correct, provide good data to the application, and repair the damaged copy."
Also it may have copies of the data eg raid 1 but does not know which is correct if they differ.
First, in between your monthly/weekly scrubs your disks/controllers will be silently corrupting data, possibly generating errors on multiple devices resulting in data loss depending on raid type. ZFS detects corruption much more quickly.
Second, your traditional raid recovery is to rewriting the entire device to fix a single block. Let's say you're using RAID5 and you're rewriting parity. You get another block error. Oops, now you've lost everything. Since disks have an uncorrectible block error rate of 1 in 10^15 bits, you only need a moderately sized array to almost guarantee data loss. ZFS rewrites corrupt data on the fly.
I'm not trying to argue that mdadm is better than ZFS, just that in this case they pretty much compare the same.
I think the pertinent question is: when the filesystem goes to read a 4K block, and one drive's copy of this block in the RAID-1 set is different to its counterpart 4K block on another disk, which one wins?
Honest question: how often to drives silently fail? Drives contain per-sector checksums these days, explicitly to prevent this problem.
When I started out, I began with a single 5-disk vdev using a SATA port multiplier enclosure to connect the drives. Over time, I bought more SATA port multipliers, but eventually ran into tons of problems with the SATA PM technology. I do not recommend anyone use SATA PM for anything that really matters, and if you must use it, do not run more than one port multiplier on a single system at a time. So, I had tons of failures where the drives would just drop out due to their enclosure, at least once a month.
After switching to SAS and SAS expanders for drive connectivity, about a year ago, I have had NO problems at all. Rock solid.
Edit: I think I've lucked out by choosing very good, slow, low-temperature drives, first the HD204UI by Samsung, then the WD Red series. With my air conditioning, and the airflow in my rack, the drives average around 32 degrees C (a little colder than Google's report would suggest is best, but close). I would get very anxious if I were running with faster/hotter drives, or "green" drives not designed for 24/7 use.
Also, why haven't you upgraded your ZFS file system?
I haven't upgraded because, well, I haven't really seen a need to? The original reason is that this pool began under zfs-fuse, and when I switched to ZoL I kept the version at the last version supported by zfs-fuse so that I could switch back if needed. I doubt I'll ever switch back, but I do like the idea of maintaining compatibility with other ZFS implementations in case of any problems. I suppose when the OpenZFS unification stuff actually finishes, I'll be happy to upgrade to the latest version?
Another question, have you checked your block size is configured correctly? I hadn't even realised that mine were wrong until I'd upgraded to the newer versions of ZFS, which throw the following helpful message:
pool: primus
state: ONLINE
status: One or more devices are configured to use a non-native block size.
Expect reduced performance.
action: Replace affected devices with devices that support the
configured block size, or migrate data to a properly configured
pool.
scan: scrub repaired 0 in 16h44m with 0 errors on Tue Aug 26 17:54:48 2014
config:
NAME STATE READ WRITE CKSUM
primus ONLINE 0 0 0
raidz1-0 ONLINE 0 0 0
ada0 ONLINE 0 0 0 block size: 512B configured, 4096B native
ada2 ONLINE 0 0 0 block size: 512B configured, 4096B native
ada3 ONLINE 0 0 0 block size: 512B configured, 4096B native
raidz1-1 ONLINE 0 0 0
ada7 ONLINE 0 0 0
ada5 ONLINE 0 0 0
ada6 ONLINE 0 0 0
cache
ada1 ONLINE 0 0 0
errors: No known data errors
Just in case you were curious, this is what my pool looks like: NAME USED AVAIL REFER MOUNTPOINT
primus 7.00T 115G 287G /primus
primus/audio 774G 115G 774G /primus/audio
primus/devel 17.2M 115G 10.9M /primus/devel
primus/documents 6.13G 115G 6.08G /primus/documents
primus/downloads 275G 115G 275G /primus/downloads
primus/git 325M 115G 316M /git
primus/jails 10.7G 115G 49.3K /jails
primus/jails/alphatrion 212M 115G 752M /jails/alphatrion
primus/jails/cybertron 91.6M 115G 729M /jails/cybertron
primus/jails/elitaone 370M 115G 940M /jails/elitaone
primus/jails/galvatron 750M 115G 1.25G /jails/galvatron
primus/jails/megatron 937M 115G 1.31G /jails/megatron
primus/jails/template 2.60G 115G 691M /jails/template
primus/jails/unicron 5.81G 115G 5.33G /jails/unicron
primus/pictures 11.3G 115G 11.3G /primus/pictures
primus/videos 5.66T 115G 5.66T /primus/videos
zroot 12.7G 94.6G 6.56G /
(in case you weren't aware; jails are FreeBSD containers - the FreeBSD equivalent of LXC / OpenVZ)If it is possible, yeah, I'll definitely replace the older vdevs entirely with new ones that have better ashift. And, while I'm at it, I'll probably switch to all 6-disk raidz2s, since that's another thing that I only learned too late, that raidz2 works best with an even number of disks in the vdev...
However I've only done minimal investigation into this issue so if you do find a safe way to upgrade the block size then I would love to know (I'm in a similar situation to yourself in that regard)
I have a much smaller zfs setup at home, as well as a few at work, and I would be super concerned if the scrubs were repairing data. Am I overreacting?
See http://docs.oracle.com/cd/E23823_01/html/819-5461/gbbwa.html... for more information.
http://en.wikipedia.org/wiki/Data_corruption#Silent_data_cor...
So, yes, I think even with the best hardware, and proper maintenance, seeing some data repaired in a scrub is to be expected.
(That said, I did spend way too long using technology [SATA PM] that often failed and made it impossible for me to run a scrub. It's very possible that normal error rates are more like what I'm seeing now, a single byte every month or so, and that the megabyte figure is representative of errors from the days of my arrays dropping out unexpectedly.)
-Cam
edit: never mind, just saw someone else asked about that and your response since I opened the page earlier today.
Are they totally wrong, do they have some kind of special implementation that requires lots of ram not to fail, or are you just not worried about it?
ref: http://doc.freenas.org/index.php/Hardware_Recommendations
So, yes, you definitely need a lot of RAM. Consider cache devices if you can't do the GB/TB ratio.
Do you have millions++ of files or just a lot of very large files?
I'm asking because I feel that not enough is published on the subject of personal filing/archiving systems, whereas it's something we all do and there's a lot of best practice sitting out there uncaptured.
A lot of my older files, sadly, are stored in "SORT/Sort Me/To be sorted/Old computer/Sort again/Miscellaneous..." and the like. My server has an mlocate index, so I'll use mlocate, and I'll use find sometimes. I make sure to preserve metadata like last-modified/created dates, so I can use that to narrow things down.
Newer stuff, I try to keep a bit more organized, but I still have lots of unmanaged stuff floating around. For big projects, or big files, that's easy enough; my photos are sorted into a Y/M/D hierarchy, my VHS digitization projects are fairly well organized, some other things have their own structure. For my scanned documents, I just dump them all into a mess of folders, but then have a custom Django app with a management command that indexes them and gives me a nice "document management" website, and then I just search based on OCR'd text or title or date.
I really hate hierarchical filesystems. After using computers for this long, I'm convinced that hierarchy-optional, metadata-driven stuff is the only future I'll be happy in. I long for the ability to save things without really having to say anything about where it's saved, and still be able to find it... So, sorry, I don't think I have a satisfactory answer for you, as I don't think there's a good solution to this problem as long as we have filesystems where the organizational primative is a hierarchy. Even with tag-based systems that build on top of that, it's usually clunky and you still fundamentally have to figure out where to save something "first", even if you plan to access it via tag/metadata later. Such a pain.
Or at least, that's the plan. I've only implemented it as far as photo storage is concerned. Haven't yet figured out if Dropbox can be part of the general solution - security and privacy concerns.
I wrote up a spec very similar to this (though I just used the hash itself for the folders, as in HA/HAS/HASH structure [there's probably a name for that scheme]), but haven't gotten around to implementing it. My main problem with actually implementing such a system is that I don't really like depending on Django or web-based interfaces; I'm a huge fan of files, and UNIX-style tools that operate on them, I just don't like the hierarchical filesystem. I've considered that a FUSE frontend to such a system would probably address most of my concerns, but at that point it's still a big huge abstraction layer that I start to feel uncomfortable for nebulous reasons aside.
But, very nice. It's nice to hear that I'm not the only person driven to such extremes. :)
o_O
I am also ripping all my DVDs onto my home server. That chews through a bit of storage, but it's also a lot more convenient (and, with kids, less likely to end in tears of sctrached or bejammed DVDs). BluRays will be next.
It's easy to use a vast amount of storage.