FreeNAS: Open Source Storage Operating System
freenas.org
freenas.org
Always remember to ask about the "HN Readers Discount".
[1] http://www.rsync.net/products/zfsintro.html
[2] http://arstechnica.com/information-technology/2015/12/rsync-...
I wouldn't recommend using LUKS, though - leave it to the built-in support.
This is unrelated to ZFS send/recv, other than the fact that we are also the only cloud storage platform to offer borg/attic support.
How does integrity protection work in practice? I know that the bits on the HDD itself do not map one to one to bits you can actually use and that it uses that to protect us from flipped bits.
I know that beyond that RAID/RAIDZ is supposed to help. I guess redundancy alone does not help. Two copies would not help to determine which bit is the correct one, you would need three. Once you have these, I guess you would be able to detect corruption. Am I guessing correct that duplication beyond 3 times would just help with read speed and in the case when a drive fails while one is already dead (e.g. when you are rebuilding)?
What then? Does RAID/ZFS automatically fix it? I imagine at some point the drive would have to be replaced, do you have to check that manually or does it detect that by itself? I imagine after that you would have to run a rebuild.
So I guess my question would be: Could I put a FreeNas in a corner a room and leave it there? Would it just blink when I need to put a new drive in or would it need more maintenance? Of course talking about worst case scenarios here, so say I want this to live for 10 years sitting in the corner.
You're confusing two topics. Integrity vs. availability. 3x copies, RAID4 XOR, RAIDZ (a form of RAID6), etc are about increasing the availability of your data in the face of device failure/partitions.
The integrity of you data is computed different ways in different systems, but you can safely think of it in your head as a pile of checksums on top of eachother. There is a half decent blog entry[0] on Oracle's website, but you can mentally model it as having checksum on disk for each block, which then is combined with other blocks into some larger object that is also checksummed, and on and on until you have "end-to-end integrity". It is a feature of all serious storage systems.
I'd love if someone who was more familiar with ZFS could respond with how ZFS's internals work for you.. I worked at a direct competitor of Sun's that sued eachother over this stuff, so I actively tried to not learn about ZFS. I regret this.
[0]: https://blogs.oracle.com/bonwick/entry/zfs_end_to_end_data
As for your ZFS question, Bonwick's blog is certainly a good source -- though if one is looking for a thorough treatment, I might recommend "Reliability Analysis of ZFS" by Asim Kadav and Abhishek Rajimwale.[2]
As far as NetApp v. Sun, it was very unfortunate. I could rant for days about how crappy the business reasons were for NetApp going after Sun (zomg coraid!).
I use your quote about Oracle being a lawn mower from your usenix presentation all the time. Thank you for your reply.
ZFS does protect against undetected (at the disk level) errors. To a first approximation, it does this by keeping block checksums alongside pointers so that it can verify that a block has not been changed since it was referenced, by keeping multiple checksummed root blocks, and by having enough redundancy to reconstruct blocks that fail their checksums. Naturally, there are information-theoretical limits to the number of corruptions that may occur for detection/correction to be guaranteed.
You should refer to the relevant wikipedia pages if you want more detail than that :p
edit: To answer your actual question, yes, you should be able to leave a ZFS box in the corner unattended for years with reasonable confidence that your data is safe and uncorrupted (tune redundancy settings to your taste on the space efficiency vs. error tolerance trade-off). Two caveats: 1) the machine must have an effective means of communicating a catastrophic disk failure for you to resolve (this should hold for any made-for-purpose NAS device, but you'll need to do some work if you're DIY). 2) ZFS does not actively patrol for corruptions, it fixes problems when it encounters them. If your data is not being periodically accessed, there will need to be some provision for walking the filesystem and correcting any accumulated errors (the ZFS scrub util exists for this purpose, but it has to be used)
To my practical point: Does it tell me, when it approaches that limit or do I have to put in more maintenance? Can it be fixed by swapping one of the drives?
You still need to worry about failing drives and about the integrity of your raidz arrays (or whatever), but that has nothing to do with the flipping bits.
That being said, you can see statistics about error corrections (which should typically be near-zero) and if you see a lot of them it might be advanced warning of a drive dying. But the actual bit errors themselves would not be a problem and you would not need to take any action specifically related to them.
It can be configured with a hot spare so the "resilver" starts immediately on failure detection, not "when you happen to fix".
We lost everything on a FreeNAS system once because we were alerted to a failed-disk error, but decided to wait until the next week to replace it. But then a second disk failed, and we lost the storage pool. Lesson learned!
But this was an operator error, not a system problem. FreeNAS has been tremendously stable and reliable for us.
We buy a batch of 20 drives all at the same time, and they are all the same manufacturer, model, size etc... Possibly even from the same batch or date of manufacture. Then we put them in continuous use in the same room, at the same temperature, in the same chassis. Finally they have an almost identical amount of reads/writes.
Then we act shocked that two drives fail within a short interval of each other :)
Now its 18 disks (3x6 raidz2) from 3 different manufacturers and every vdev has 2 of each. And the vdevs are physically evenly spread throughout the case.
I sleep so much better. It was kind of a miracle the first setup survived the 4.5 years it did.
The nice zfs built-in autoreplace functionality doesn't work on Linux or FreeNAS. You need some scripting/external tooling to do the equivalent. See
https://github.com/zfsonlinux/zfs/issues/2449
https://forums.servethehome.com/index.php?threads/zol-hotspa...
I've had a few drives fail on my ZFS for Linux fileserver and wondered why my hot spares weren't automatically kicking in, and this is why.
On Linux, if you don't use the zed script that's referenced in that Github issue above and just replace a failing drive manually, a hot spare is worse than useless, because you need to remove the hot spare from the array before you can use it with a manual replace operation.
As for "automatically fix it", the short answer is there is a lot of stuff that automatically fixes problems, but it is a leaky abstraction. RAID-5 rebuilds are notoriously terrible for performance, and often it is easier to have logic for dealing with failures & redundancy at the application layer.
FreeNAS and similar projects are definitely intended to be turnkey storage solutions. They have their strengths & weaknesses, but the notion that you just plug it in and go isn't too far off. Usually you don't go with the blinky light, but with an alerting mechanism (e-mail, SMS, whatever) that you integrate with it for notifications about problems. In principle, it is a fire & forget kind of thing.
ECC is always written alongside the actual block and the overhead for ECC is the reason for the move from 512b sectors to 4Kb sectors in HDDs. For SSDs the data is already written in different block sizes depending on the NAND and the internal representation and ECC is done for larger than 512b units.
The probability of failure during rebuild is not really directly linked to drive size, the usual interpretation of the drive BER is wrong (media BER is stated across a large population of drives rather than just one drive).
Yeah, I expressed that badly. I always forget about the ECC bytes that are in the firmware.
> The probability of failure during rebuild is not really directly linked to drive size, the usual interpretation of the drive BER is wrong (media BER is stated across a large population of drives rather than just one drive).
Regardless of interpretations of BER, I can't agree about drive sizes. The phenomenon of failures during rebuild is well documented and the driving principle behind double-parity RAID. Adam Leventhal (who ought to know this stuff better than either of us) wrote a paper several years back on the need for triple-parity RAID, and it was entirely driven by increased drive densities: http://queue.acm.org/detail.cfm?id=1670144
The reality is the higher drive densities mean you lose more bytes at a time when you have a drive failure, and that means more bytes you want to have "recovered".
I do assume though that the RAID array does media scrub (BMS) periodically, if you don't you are at risk anyway and I'd call that negligence in maintaining your RAID.
If you do use scrubbing the risk that another drive has a bad media spot is low as it must have developed in the time from the last scan and that is a bounded time (week, two, four) so the risk of two drives having a bad sector is now even lower (though never non-zero, backups are still a thing). If you couple that also with TLER and proper media handling instead of dropping the disk on media error the risk to the data becomes very low since there isn't a very high likelyhood that two disks will have a bad sector in the same stripe.
I've been working with HDDs and SSDs and developing software for enterprise storage systems for a number of years now, I've worked in XIV and for all the thousands of systems and hundred of thousand disks of many models I've never seen two disks fail to read the same stripe and RAID recovery was always possible. Other problems are more likely sooner than an actual RAID failure (technician shutting down the system by pressing the UPS buttons or a software bug).
I did learn of several failure modes that can increase the risk but they depend on specific workloads that are not generally applicable. One of those is that if you write to one track all too often you may affect nearby tracks and if the workload is high enough you don't give the disk the time to fix this (the disk tracks this failure mode and will workaround it in idle time). In such a case same stripe can be affected in multiple disks and the time to develop may (or may not) be shorter than the media scrub time. And even then a thin provisioned raid structure would reduce the risk of this failure mode and giving disks some rest time (even just a few seconds) would allow the drive to fix this and other problems it knows about.
All in all, RAID is not dead (yet).
Some systems mitigate that by having a more evenly distributed RAID such that the rebuild time doesn't increase that much as the drive size increases and is actually rather low. XIV systems are like that.
That was exactly what I was saying. Given the same sustained transfer rate, the bigger the drive, the longer the rebuild time, hence the greater the chance you'll have a failure during the rebuild. While you might think they would, throughput increases have not grown to match increases in bit density & storage capacity.
SSD's have helped a bit in this area because their failure profiles are different and under the covers they are kind of like a giant array of mini devices, but AFAIK they still present challenges during RAID rebuilds.
Checksums.
If you get big enough checksums (aka: Reed-Solomon Codes), you can not only detect errors but correct them as well. See https://en.wikipedia.org/wiki/Reed%E2%80%93Solomon_error_cor... for the math.
Now that you have this "error-correcting checksum", where do you put it? Raid5 means you place the error-correcting checksum on different disks.
If you have three disks: A, B, and C. You'll put the data on A & B, then the checksum on C. This is RAID4 (which is never used).
RAID5 is much like RAID4, except you also cycle between the drives. So the checksum information is stored on A, B, or C. Sometimes the data is on A&B, sometimes is B&C, and sometimes its on A&C.
-----------------------
> Could I put a FreeNas in a corner a room and leave it there? Would it just blink when I need to put a new drive in or would it need more maintenance? Of course talking about worst case scenarios here, so say I want this to live for 10 years sitting in the corner.
Maybe, maybe not. If all the drives fail in those 10 years, of course not (Hard Drive arms may lose lubricant. If the arms stop moving, you won't be able to read the data).
"Good practice" means that you want to boot up the ZFS box and run a "scrub" on it every few months, to ensure all the hard drives are actually working. If one fails, you replace it and rebuild all the checksums (or the data from the checksums).
RAID / ZFS isn't a magic bullet. They just buy you additional time: time where your rig begins to break but is still in a repairable position.
ZFS has more checksums everywhere to check for a few more cases than simple RAID5. But otherwise, the fundamentals remain the same. You need to regularly check for broken hard drives and then replace them before too many hard drives break.
---------
This also means that no single box can protect you from a natural disaster: fires, floods, earthquakes... these can destroy your backup all at once. If all hard drives fail at the same time, you lose your data.
I run it on all my hdd's and its kept them alive and running smoothly for years at a time.
When I built my home NAS there wasn't an off the shelf FreeNAS option and it was definitely a "research all the things, build your own system" with the huge caveat of "Did you put enough RAM in that?".
The 8-bay one is particularly good value, rivalling similar systems by QNAP, and I personally do have a QNAP and if I were to buy again today it would definitely be one of these FreeNAS boxes instead.
One limitation is the 16GB max ram, but given that it's for home use with (presumably) large files and a small number of simultaneous clients, it shouldn't be a problem - even if you use 4x 8 TB disks.
The key is to populate it with ECC ram which should be an absolute requirement for any ZFS system.
[0]: http://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-yo...
I've used both FreeNAS and OMV, and find the OMV community more welcoming and helpful than the FreeNAS one. The FreeNAS community has organized itself with the intent of focusing exclusively on very specific very corporate use case. If you don't fit they're use case, you are bad and should feel bad. OMV on the other hand is much more open and willing to help people use adapt it to their own needs.
With the notable exception of ZFS vs EXT4, they underlying software appears to be the same, which makes the community distinction every more cogent.
Huh no. OMV is Linux + PHP.
I keep a copy of the FreeNAS's storage on a second DIY NAS running NAS4free, which is the original base from which FreeNAS was forked some years ago.
One popular thing I don't do with my NAS boxes is run any non-NAS services e.g. media streaming, etc. For me it feels like an unnecessary risk to important data.
Works well, with some limitations. I dont think their automatic port forwarding service (quickconnect) works well - since that service checks for a valid Synology MAC address.
Additionally - you can't do the automatic system updates as it might break certain portions that XPEnology overrides.
If it isn't Open Source this is off-topic since there is a large number of other commercial NAS OS e.g. Thecus, SoftNAS, QNAP, Napp-It... Most of them tend to have an "app store" and some people building more or less working apps out of Open Source and free projects.
I think it's in there, their account's got a lot of files so it might be elsewhere. Latest version isn't published yet.
This thread seems to confirm this notion: https://forum.synology.com/enu/viewtopic.php?t=120535
Here is the one for Thecus: http://www.thecus.com/Downloads/GPL/ QNAP: https://sourceforge.net/projects/qosgpl/
And since they haven't published the latest versions they are in violation of the GPL.
There's also an independent community that facilitate installing and using it on your own hardware including the latest version that's still waiting on the source dump.
As an example for their relationship with upstream, they market the btrfs feature (https://www.synology.com/en-global/dsm/Btrfs ) but don't have anyone working on btrfs (and btrfs could surely need the help). I can see one patch set from them which hasn't made it into mainline ( https://patchwork.kernel.org/project/linux-btrfs/list/?submi... ).
There are a few ways that they can comply with the GPL without posting the source on the web.
This is not the same thing as "The full source code for all components that ship with Synology NAS"
xpenology is an illegal mashup of some open source components with a large amount of binary blobs installed on top of it.
https://www.cachem.fr/xpenology-mot-fin/
It's fine if you want to use it, but don't kid yourself that you are using open source software.
To save effort, the 1.5G braswell-source.txz file, contains:
attr-2.4.44 tzdata linux-3.10.x u-boot-mv-3.5.9 gvfs-1.x freetds-0.x wvstreams-4.6.x systemd faac-1.28 upstart-1.x e2fsprogs-1.42-virtual-glibc freetype-2.x bridge-utils-1.4 bluez-4.x php-apcu-4.x ipmitool-1.8.x libpng-1.2.x libharu-2.x libcap-2.x libffi-3.x libusb-0.1.12 lnxlibnet gnupg-2.x synodb sysfsutils-2.1.0 nut-2.6 compat-wireless taglib-1.9.x imagemagick-6.9.x iptables-1.4.x readline-6.x apache-2.2.x-virtual-npn postgresql-9.3.x iproute2-2.6.31 zlib-1.x gzip-1.x libsynosmtp libnih-1.x libsoup-2.x sqlite-3.8.x imap-2007f netatalk-3.x u-boot-mv-3.6.0 hostapd-1 open-iscsi-2.0-871 mdadm ntfs-3g krb5-1.12.x gmp-6.x libevent-2.x apparmor-2.9.x procps-3.2.6 libxmltok-1.2 quota-tools-3.17 libical-0.x irqbalance backports libsynosdk-virtual-gpl libproxy-0.4.x Freescale_QorIQ_boot util-linux-2.x miniupnp-1.x iproute2-3.2.0 alsa-driver parted mhash-0.x bdb-5.x gnutls-3.2.x libdaemon-0.14 xtables-addons-1.x php-5.5.x ctdb-2.5.x coreutils-8.x e2fsprogs-1.42 ncurses-5.x-virtual-64 lame-398 nbnsd libgphoto2-2.1.99 alsa-lib-1.0.25 apache-2.2.x mDNSResponder-258.13 mod-fastcgi-2.x cups-1.5.x libexif-0.6.x tar-1.x debsig-verify-0.x php-5.5.x-virtual-module urftopdf kmod-16 openssh-6.x usbip-0.1.7 libpng-1.6.x curl-7.x pth-2.x xz-5.x gd-2.x libassuan-2.x neon-0.x libgcrypt-1.6.x rsync-3.x bdb.1.85 mbedtls-1.x bzip2-1.x wireless-tools.29 nettle-2.x uthash-1.9.x nss-pam-ldapd-0.8.x ssmtp-2.x wvdial-1.x usb-modeswitch-1.2.x libsynocore-virtual-gpl samba-4.x openssl-fips-2.0.x libdbi-0.9.x uclibc0929 cups-bjnp pcre-8.x u-boot-mv-3.5.3 sg3-utils ppp-2.4.x libaio-0.x libnl-2.x vmtouch u-boot-armada-2011.12 popt-1.x fuse-2.x libxml2-2.x syslog-ng-3.5.x libdbi-drivers-0.9.x u-boot-mv-3.4.4 eventlog-0.2.x libjpeg-turbo-1.x dosfstools28 u-boot-1.3.3 wget-1.x libksba-1.x intelce-utilities ndisc6-1.x x264 python-2.7.x opencore-amr-0.1.2 boost-1.x linux-firmware openssl-1.0.x wireless-iw cifs-utils-5.x libgpg-error-1.x iscsitarget-0.4.17 rp-pppoe-3.x busybox-1.16.1 json-c-0.x libmcrypt-2.5.8 logrotate-3.x dbus-glib-0.x openldap-2.4.x flex-2.5.4 u-boot-mv-3.4.27 util-linux-2.x-virtual-64 avahi-0.6.x gdbm-1.x glib-networking-2.x net-snmp-5.x ncurses-5.x ntp-4.2.8 flashcache jsoncpp expat-2.x libssh2-1.4.x lftp-4.x glib-2.x dnsmasq-2.x icu-53.x cyrus-sasl-2.1.22 ffmpeg-2.0.x linux-pam dbus-1.6.x suphp-0.7.x
other than
synodb libsynosmtp libsynosdk-virtual-gpl libsynocore-virtual-gpl
which are likely related to their SDK, is all standard linux packages.
I went back to Synology after my uper-nas crashed and burned (really bad run of 3TB Seagate drives, 9 out of 12 dead in under 3 years)... Now debating between a 4x8TB or 5x6tb in a future nas box.
Some retailers have even started to sell the 3TB Seagate drives at almost the same price as the 2TB ones. Probably the only way to get rid of them.
Also, back blaze statistics only apply if you're operating at back blaze's scale. If you're making a home NAS you would have to be _incredibly_ unlucky to hit a bad drive as quickly as back blaze did. I
I absolutely agree though, the 4TB are a great improvement, you can't beat them for reliability unless you're really willing to pay for it. I've chosen them for my personal NAS and some servers even with great results so far.
Yeah, I've recovered several failed drives without issue, even on LVM systems where some partitions failed because they were striped but others worked because they were mirrored. No real issues overall.
The only case where you might run into problems is when you get conflicting data between 2 drives, but I've never seen that happen in the real world. Some people will scare you into thinking that happens often but it hasn't been my experience. Often one drive just goes kaput completely or has an entire section which is completely unreadable or extremely slow, for which recovery is just a matter of marking that drive offline, swapping a new drive in and telling it to rebuild. The only case in which you lose everything is if you use RAID0. If you can do RAID10 that's the way to go IMO.
RAID5 can be a little riskier because if one drive is damaged and one fails then you can wind up losing the bunch, but I use it on my personal NAS because it's mostly just TV and the important stuff is syncthing'd to my other machines. Regular scrubbing prevents these issues usually.
RAID5 can be pretty risky with large drives that are the same type.
>conflicting data between 2 drives, but I've never seen that happen in the real world
Heh, run enough servers and you'll see everything eventually. Had a really fun one where a hardware raid card was telling us that writes were successful, but when data was coming back corrupt. It was writing bad data to two of the drives in the array.
First, buy good drives. Buying desktop drives that suck are a great way to end up with a corrupt SAN. In general you want drives that support TLER so if something goes wrong, you hear about it right away, rather than the drive trying to fix the problem silently while causing the disk to have long delays. You don't have to go super expensive enterprise, but buying from the bottom of the barrel, or is a 'green' or efficient drive is a great way to lose data.
Second, no matter if you are using hardware or software raid, set up the monitoring utilities properly. You need an immediate alert if a drive is going bad or has failed. SMART should be enabled and when it alerts, replace the drive. The software should also do a verify, patrol read, or scrub (different terms for the same thing) that occasionally check the surface media of the disk and the validity of the data. If anything comes back with an error, you replace the disk right then.
Lastly, one of the big issues why people lose raids all at once is they go out and buy a huge stack of the same kind of disk, made on the same day, with the same firmware, and possibly the same inattentive QA person. Sometimes when problems crop up, it can be a systemic problem with that model, and when all you have is one model, bye bye everything.
if you want to learn more about this topic, come on over to www.reddit.com/r/datahoarder and read and ask about it. We're friendly.
I've always thought that was essential, it bemuses me that there aren't many solutions out there that help you do this. I guess everyone must buy have loads of disposable income to buy 8 drives at once.
You can expand the capacity of the pool by replacing each drive, changing them one by one and re-synchronizing the array between replacements.
It's a KVM hypervisor and I haven't played online, why would they ban you for using a VM?
- rpi 3 (running debian testing) - 2x1Tb disks (with usb docking) - zfs (raidz) on the disks - ssh through tor for tunneling purposes
I can backup my personal computers and run some extra stuff on the rpi. Yes, it's only USB2, but I don't care. I mostly send incremental snapshots from my computer and it's fast enough. (e.g. zfs send -i yesterday home/user@today | ssh rpi zfs recv home/user)
I'm tempted on building a small recv server that can allow partial transfers, but so far it works for me. Also, as soon as encryption is enabled on zfs-linux, I'm turning that on :)
Total budget: ~$180 (the most expensive part were obviously the HDDs)
Where they fall down for me as a home user is in their inability to offer other services. For example, photo sharing. I have a fast home internet connection and all my photos are already on the NAS. But the photo sharing capabilities are very, very weak; so inevitably I have to use another service (ie Google Photos, Dropbox, Flickr) if I want to make those photos available to myself and others while away.
The same goes for other forms of media sharing, self-hosted email, etc.. The boxes offer these things - and the appeal is obvious to the end-user - but the potency is so weak as to be essentially unusable.
Given that the OS needs to run somewhere, and one would rather use the 4 bays for drives in the zpool, I used the recommended method of running the OS from a USB drive. Sounded strange to me, but it's worked a treat. I saw that there are kits out there for replacing the empty CD slot in the Microserver with an SSD for the same purpose, but I'd rather not touch something that's working now.
I went for RAIDZ1, despite a lot of internet protestations about it, as I basically only store movies/music/Time Machine backup there anyway.
Been thinking about backup solutions for everything, but it all seems fairly expensive for what is essentially just backing up media.
To people recommending Plex for media streaming, I would actually recommend Serviio over it. It took a little extra setup in a jail, requiring some extra port downloads, but it works extremely well in it's DLNA streaming to my PS3. Where it shone brighter than Plex for me was it's support for transcoding and subtitles - it transcodes and shows all of my H264 video files with ASS subtitles with all of the style intact. It also supported PGS subtitles, but my piddling processor didn't quite match up to it.
Plex has done transcoding, including subtitle burn in, for a LONG time. It's a modified (I think) ffmpeg they ship with the server.
All the major solutions have their drawbacks. For example, btrfs' present day raid code is in such bad shape that apparently it needs a total rewrite. zfs' volumes can't grow without a resliver, which is not something a home user should ever have to do.
From a design perspective, the internals of both btrfs and zfs are only accessible to experts and totally opaque to their users. I've long thought someone could come along and write a FUSE-based distributed storage layer that sat on top of existing file system abstractions like ext4, etc. And that a system built that way could just crush both of those solutions in almost every conceivable way. I've toyed with writing it myself.
This approach might not be the fastest, but it would be the most flexible, possibly very reliable, and as a bonus it would be really easy for average users to grok and maximize their data recovery from all but the most catastrophic events.
Hyperbole, yes, but the point is that the devil is in the details, and there's a hell of a lot of caveats to the term "reliable data storage". For example, in the modern era, if someone had a way to make ye olde floppy disks 100% reliable for 10 decades, they still wouldn't be used by that 99.9%, simply due to the low data capacity.
Downside is, you lose a significant chunk of storage (the author measured > 20%), but you can add new storage to a pool online.
I've gone ahead and actually created a new array using btrfs raid10 instead (with an eye towards a RAIDZ2/3 like configuration whenever the btrfs guys get around to fixing it); btrfs allows for online reshaping of the mirroring configuration, so it doesn't have the limitations of ZFS in this instance.
I don't have experience with FreeNAS unfortunately.
Other raid levels (raid0, raid1, raid10) work fine and are considered stable -- at least as stable as the filesystem itself.
Either way, if you care about the data, you must backup, no matter what fs you use.
$ zpool status
pool: BoxODisks
state: ONLINE
scan: scrub repaired 0 in 9h20m with 0 errors on Sun Jul 31 01:31:31 2016
config:
NAME STATE READ WRITE CKSUM
BoxODisks ONLINE 0 0 0
mirror-0 ONLINE 0 0 0
gptid/18059e22-b6a4-11e5-9cca-0cc47a6bbf34 ONLINE 0 0 0
gptid/194649b1-b6a4-11e5-9cca-0cc47a6bbf34 ONLINE 0 0 0
mirror-1 ONLINE 0 0 0
gptid/1a86b3cc-b6a4-11e5-9cca-0cc47a6bbf34 ONLINE 0 0 0
gptid/1bcd3ad6-b6a4-11e5-9cca-0cc47a6bbf34 ONLINE 0 0 0
cache
diskid/DISK-S24ZNWAG903847Lp1 ONLINE 0 0 0
I think I have mirrored vdevs.What does that mean for replacing failed disk/s? What does that mean when I want to upgrade to 8TB drives? The biggest "mistake" I made was buying a very expensive 4 bay enclosure - 4 drives is not enough.
True 'dat!
I even run some ZFS without ECC memory, because if hardware goes bad and programmatically corrupts it isn't the end of the world.
Don't concentrate your resources into a single "invincible" box.
Is there a uniform resource identifier where I could read more about it? Or perhaps some book? Research paper?
It synchronizes directory trees on arbitrary filesystems, so I think the answer to your question would be "asynchronous".
Most individuals don't have TBs of working set continually being updated and requiring a central authoritative copy. They have a small working set and a long-term archive they'd just never want to lose.
Unison's model isn't perfect (eg having to create a star topology, lack of built-in inotify). But it's been around forever, is written in a sane language, and is rock solid.
... I was able to recover, but, I remember how many hours I spent on it.
Please tell me it is better now?
you can lose the system disk (usb key) that holds your settings, but the data you actually care about will be safe.
Even if it's ZFS it can get quite messed up if you set it up based on sda, sdb, etc like naming conventions. In that case entire zpools can fail to import in case of failure.
Therefore the best practice is to use device-by-id, so that if a device like sda falls out, the rest of your array should still resolve correctly.
And how the array is created is entirely a FreeNAS thing. Hopefully one which have, like OP says, improved.
As soon as I get home, I'm going to swap some of my HDD cables and see what happens... I've basically been assuming that NAS4Free uses device-by-id like any sane OS should.
That said, ZFS pools under FreeBSD and derivatives, just as under linux, are imported from info stored on the drives themselves. They should be able to reboot after scrambling the cables. It would take more nerve, or craziness, than I've got to PROVE that using any pools I care about.
However, I have completely screwed a ZFS replace command in FreeBSD following a bad drive in a RAID-Z3 pool. It dutifully replaced a perfectly good drive with my new drive. But no harm was done to the pool. After it finished, I ran another zfs replace, replacing the bad drive with the "mistake" old drive. It all worked!
ZFS is so brilliant, be careful or it will put your eyes out :-)
Now I'm back to HW RAID, LVM, and ext4, which works a treat, and I'll be using that until I have more faith in ZFS.
Nothing is safe.
A split caused by Olivier waking up one morning and announcing that freenas is dead because he wants to rewrite 0.8 on Linux because reasons. I switched back across to freebsd right about then as I lost all trust in the project (now NAS4Free).
Currently running FreeNAS for the last 2 years since last machine upgrade.
Upon seeing FreeNAS Mini, I have to say that I'm really interested. I've become increasingly uncomfortable with the idea of keeping all of my data "in the cloud". Being able to keep a local version of all my media would help give me peace of mind.
Is maintaining your own "personal cloud", if you will, a reasonable use-case for this? What are the proper expectations from a maintenance perspective? i.e. How often should I expect things to break or require tuning / fixing?
This is not the only copy of any of this stuff, though. The photos get synced to S3 and backed up on a second local hard drive. All of the movies I have are rips of DVDs and Blu-Rays I still own, and can use if needed. And my music is all in Apple Music/iTunes Match (and on my computer, backed up via both Time Machine and Crashplan).
So, I'm using the server really as convenient local storage for stuff, not as an alternative to the cloud.
It has a good cross section of features from 'home' to 'pro', AFP/NFS/CIFS/iSCSI, directory integration, plugins (more aimed at home, with things like news downloaders, BT clients).
Can anyone explain the reason for this? What makes the specs so hardcore?
Kind of like when you don't need a full blown CRM to manage two customers.
Having said that, RAM is pretty cheap. 8GiB isn't a huge ask for any machine in 2016. Any system with less than 4GiB, is it really worth bothering with redundancy/ZFS? UFS (which is far less RAM hungry) + nightly backups might be good enough. (Edit: of course, you miss out on the awesome ZFS features with UFS. It's never easy, is it?)
I found Ceph. Also 100% opensource. But needs like 3 physical machines in production (at least) in order to have some kind of built-in redundancy.
EDIT: Let me explain a bit more: once you start running a distributed system you stop caring about OS, filesystem and disk level issues, because you have redundancy on another level. And it makes all the difference. You don't worry anymore, you can always just reboot, hard-reset or take a node out to investigate. Suddenly you realize that it's not a big deal even if some node starts freezing or some process starts OOMing and crashing, you don't care, you just let them.
I see this line of thought that somehow self-hosted "clouds"/clusters look after themselves, but thats usually not the case.
You'd come close by buying QNAP or Synology hardware, they provide software updates (including OS). Even buying software solutions like Unraid. I dont know how maintainable FreeNAS is for someone who doesnt want to worry about OS tho..
They don't really look after themselves, but do handle failures on the highest possible level. Which makes it unnecessary to keep each OS in check. What's critical for a single NAS box is critical for a distributed storage only if all boxes of some replica have the same problem at the exact same time, otherwise you just reboot and move on and it doesn't matter if it happens again, it doesn't cause any downtime.
I actually speak from my own experience. I run and maintain a distributed key-value storage for many years. Although I designed and implemented it myself (and redesigned a bunch of times), I don't see how experience with those other distributed storages, like Swift, would be any different.
In most modern systems the rate at which the CPU can compress data is many, many times faster than the disk can write data. In general compression greatly increases write speeds because less disk IO is needed. The vast majority of benchmarks show that to be the case too.
TL:DR, if you notice a performance decrease after enabling compression, something is wrong with your NAS/SAN.
That said, if you are just storing large encrypted files on your ZFS then compression isn't needed.
On modern CPU LZ4 should enable about 400MB/s current PCU write compression https://github.com/Cyan4973/lz4
https://blog.voltagex.org/2016/08/26/benchmarking-compressio...
Fun fact: while researching this I found yet another article that contradicts the information I had when I built my zpool, and apparently I've done it all wrong (again).
Rookie mistake, but a simple patch solves it.
Start your flamewars nerds!
--------------
I actually have only used Nas4Free. I know that FreeNAS is technically the fork (despite having the original name), but the technology IMO is quite solid.
Its important to know all of the competition however. Here's a basic overview of the technologies:
* Nas4Free -- ZFS-based. Free as in beer and Free as in OSS. More barebones and simple than FreeNAS. As it is based on FreeBSD, you need to be somewhat careful about hardware choices, although in my experience FreeBSD seems to support hardware that I'm interested in.
* FreeNAS -- ZFS-based. Forked from Nas4Free and name-shenanigans happened. Newer web-gui and more plugins. Can't speak too much about it, since I haven't played with it.
-------
* Windows Storage Spaces -- Windows8 and up have ReFS + Storage Spaces as their ZFS-competitor. Runs a daemon in the background to automatically check for bitrot (unlike NAS4Free / FreeNAS where you need to schedule a cronjob). Comes as part of Windows, if building a dedicated system you need to pay the $100 Windows Tax. Best hardware compatibility available. No head-scratching about random AMD A10 / FM2+ motherboards with obscure drivers (FM2+ compatibility is not listed on FreeBSD yet)... you know everything has Windows compatibility.
Windows Storage Spaces are superior technologically to ZFS IMO. You can extend a storage space after building it, while ZFS volumes are locked to a specific size. (You can add more drives to a ZFS mirror, but this only increases reliability). Start with 2-hard drives and then extend the storage space to 6-drives later.
You can stripe data to increase a ZFS pool size, but this doesn't keep the same level of reliability. The Windows Storage Space methodology where you overpromise on storage size (and then later build out capacity) just seems to be an easier methodology to work with in the long term.
--------
ZFS has more features, but nobody uses them. ZFS supports dedup, but all documentation I've seen says its not worth it.
I guess one important feature ZFS has is that it supports L2ARC / ARC caching for SSD Acceleration.
Windows ReFS does not. ReFS is also not a complete solution. Parity is implemented at the "Storage Space" level, not at the filesystem level. I don't think this is a major downside, but it is important to note that ReFS + Storage Spaces is the complete solution. (Whereas ZFS stands alone)
Note: Snapshots (Called 'Shadow Copies' in Windows land) exist on NTFS.
--------
* Synology -- Out-of-the box systems, usually built on Intel Atom. I find that the 2-disk options are cheap, but the 4-bay or 6-bay options are outrageously expensive. I can definitely build a cheaper WINDOWS system than most 4-bay Synology Box.
Synology is mostly a soft-RAID setup. I don't see much on bitrot or other storage issues. I hope they handle it? But I'm not 100% sure.
know that FreeNAS is technically the fork
FreeNAS as it is today is not a fork, it's a re-write.Which can massively accelerate small file load times. Using this on a VM pool greatly speeds things up.
>Also ZFS has compression which can drastically reduce data usage.
Which in some virtual environments decreased our data usage by 50 times or more.
Windows storage spaces suck balls on speed. Using it in the Microsoft recommend methods to avoid data loss or corruption make it even slower.
Oh, and just throwing files on ReFS and using it directly with services and such is a great way to get weird issues if you don't understand the filesystem are different.
>while ZFS volumes are locked to a specific size
Specific number of disks. You increase the disks size and you can grow you raid.
I did some research, and it appears that Microsoft has added "SSD Tiers" to Storage Spaces. I didn't realize Microsoft had that feature!
> Specific number of disks. You increase the disks size and you can grow you raid.
Yeah. Except my disks don't grow in size. I can have a Storage Space composed of 1TB drives. And then add 2TB drives next year.
Its a nifty trick, and definitely is a smoother transition than ZFS "replace everything" methodology.
I've also heard that it is a black hole from which CPU time and RAM enter, and do not emerge.
Is this true, and if so, to what degree?