ZFS on a single core RISC-V hardware with 512MB
andreas.welcomes-you.com
andreas.welcomes-you.com
ADD: Geekbench 5.4.1 on RISC-V
- under QEMU/Ryzen 9 3900XT: 82
- under QEMU/M1: 76
- Native D1: 32 (https://browser.geekbench.com/v5/cpu/13259016)
The M1 result is skewed because for some reason AES emulation is much faster on Ryzen. The rest of the integer stuff is faster on the M1, up to 30% faster.
PiBox is the only contender I am aware of.
The application profiles (like RVA22) have a bunch of requirements on hardware that must be present, the boot process and interface the firmware offers to the OS (i.e. opensbi).
I just checked the Ubuntu RISC-V download page and it has 2 different ISOs for 2 boards from the same vendor! And apparently those are the only 2 boards with ISOs available 0_o
Yes, you will be able to, at some point.
>I just checked the Ubuntu RISC-V download page and it has 2 different ISOs for 2 boards from the same vendor!
No board out there is RVA22 compliant, as RVA22 itself isn't done and closed. It is expected to be this spring.
By the time large scale production of boards happen and they ship to the masses, this will be a solved problem.
I do not expect this to happen this year. Maybe the next it'll start to pick up.
Meanwhile in ARM/RISC-V embedded land, every chip and every board is its own special snowflake. With no ACPI/UEFI, someone’s gotta hardcode the config of every device on every board, including the ones inside the SoC. Naturally the communities around these boards are even more fragmented than Linux distros already are.
Well, you will be able to. Coming soon.
For ARM, rk3568 is available now and has 2 lanes of PCIe 3.0 and up to 3 SATA-3 ports. As opposed to the rpi4 which has a single PCIe 2.1 lane.
Pine64's Rockpro64 has a regular x4 slot for expansion cards.
I've seen lots of upcoming Pi CM4 options, but don't know what's out yet.
No idea what caused it, I suspect excessive writes and wearing the flash storage or corrupting the bootloader and not being able to recognize the boot media. The "wear" I mentioned could be something else of course, but except the normal OS writes, everything went to the SATA drives.
It sounded like a solid piece of hardware with the 4 SATA-ports-hat and good CPU, but at the end it turned out to be un-reliable as at some point the OS hang, couldn't boot from the OS storage and the bootloader wasn't seeing the partitions via UART debug session.
I other words, cheap unreliable plastic-boxes.
You are much better of with TrueNAS and a HP MicroServer Gen10 Plus (with xeon and ecc)
The world has become better, life has become more interesting.
Nobody ever said well I plugged an 20TB external hard drive so I better plug in a few more sticks of RAM so that works.
Dedup needs RAM in proportion to storage because for each duplicate block it maintains an entry in an in memory table.
All file systems have metadata which is good to keep in memory. Building several 50+TB NAS boxes recently, it isn’t just ZFS either. And it isn’t some sort of linear performance penalties sometimes when you don’t have enough RAM. It can be kernel panics, exponential decay in performance, etc.
Can you quantify what you are saying. What OS/filesystem? What minimum RAM requirements for what amount of storage?
seems like no one is building these larger systems on boxes small enough for it to matter, or at least google isn’t finding it.
However specific implementations can indeed have memory requirements that scale in relation to storage capacity. For example, if the implementation keeps the bitmap of free space in memory, then more storage = larger bitmap = more memory required.
There's been several attempts in ZFS to reduce memory overhead. I'm pretty sure that if you took a decade old version of ZFS you'd struggle to run it on a system with 512MB RAM.
I will agree that ZFS should handle large pools once you clear the ~fixed minimum memory requirement.
ZFS does complain that the minimum recommended memory is 512MB and that I can expect unstable behavior. However basic file I/O seems to work, I copied some multi-GB files around and such without issues.
So seems the bare minimum was lower than I recalled, at least on a plain system.
A proper test would include heavily fragmenting the pool, and preferably with more vdevs. But it's something.
I don't follow FreeBSD closely, but if memory serves there was a concern, particularly in the early days of the ZFS port, that their implementation couldn't be counted on to release RAM fast enough, if the system suddenly came under significant memory pressure. Hence the advice was to always run with more-than-sufficient RAM to minimize the likelihood of getting into low memory situations. I think this is a significant part of why FreeNAS considers 8GB to be the minimum supported configuration.
So it seems to me that this isn't really about ZFS's RAM requirement, rather it's about hedging against the volatility of the RAM requirements of other software on the same box, in case ZFS can't back off fast enough.
These days I run ZFS on Linux, and I remember about 4 years back spinning up some bulk data processing job that was configured to use 14GB of RAM, on a 16GB box, and watching ZFS's ARC RAM use drop in a single second from 5.5GB to 0.5GB. So I'm satisfied that for my purposes, on ZoL in recent times, this isn't an issue I need to worry about.
ZFS releases ARC memory as the memory pressure from applications running on the system increases; it's been that way since day 1.
> Unfortunately OpenZFS seems not to support (yet) cross-compiling for the RISC-V platform, hence you have to build the kernel on the RISC-V board.
Both of these seem like low hanging fruit, yes? If the upstream code supports it I'm surprised Debian doesn't already have packages, and cross-compiling isn't that special.
will it take a decade ? less?
Outside the discount pricing, Intel has promised to tape out SiFive's P650. Revos, Tenstorrent, and others are also working on fast cores, but it'll be at least 2-3 years before they hit the market if at all.
So far SiFive's dual issue in-order core (~ 40 GeekBench 5.4.1) (like on now-cancelled BeagleV) is the fastest chip you can buy as a lay person. The D1 (~ 32 GB 5.4.1) is cheaper but less powerful.
The actual chip designs, I think, are already there in terms of getting a high-performance risc-v chip built, but currently the market and tech stack is still getting they so we are still at the scaled-down-test phase in the high end market.
FWIW though I don't see much reason to care about the ISA of the CPU beyond it being RISC, chances are it'll still be full of all kinds of closed source crap like the Pi.
RISC-V's application profiles (like RVA22) do have requirements regarding some hardware that must be present (serial port with a specific interface), boot process and firmware interfaces.
These are there to prevent an ARM-like messy situation.
I still have some boards with ~512mb RAM lying around (an UltraSPARC for example) that I'd love to re-purpose to a cheap NAS, just for the heck of doing it on a non-x86 platform....
ZFS dedup is block based, and actual block size varies depending on data feed rate for most workloads (zfs queues up async writes and merges them), so in practice once a file gets some non-zero block offset somewhere which happens all the time, even identical files don’t dedup.
My setup tries to get the absolute highest bandwidth and uses NVMe sticks in a stripe (I get my redundancy elsewhere), no compression, no dedup and yet can only hit ~ 3.5 GB/s reads (TrueNAS Core, EPYC 7443P, Samsung 980PRO, 256 GiB). I hope TrueNAS SCALE will perform better.
Though there is no clean way to disable either. Compression can be removed from files by rewriting them, but removing deduplication requires copying over all data to a fresh pool.
https://www.truenas.com/community/threads/zfs-dedup-disable-...
chicken:~# zpool list
NAME SIZE ALLOC FREE CKPOINT EXPANDSZ FRAG CAP DEDUP
HEALTH ALTROOT
chicken_test 15G 144M 14.9G - - 0% 0% 16.00x
ONLINE -
chicken:~# zfs send chicken_test/dedup_source@send | zfs recv -o dedup=off
chicken_test/nodedup_dest
chicken:~# zpool list
NAME SIZE ALLOC FREE CKPOINT EXPANDSZ FRAG CAP DEDUP
HEALTH ALTROOT
chicken_test 15G 2.29G 12.7G - - 0% 15% 16.00x
ONLINE -
chicken:~# zfs get dedup chicken_test/nodedup_dest
NAME PROPERTY VALUE SOURCE
chicken_test/nodedup_dest dedup off local
chicken:~# zfs destroy -r chicken_test/dedup_source
chicken:~# zpool list
NAME SIZE ALLOC FREE CKPOINT EXPANDSZ FRAG CAP DEDUP
HEALTH ALTROOT
chicken_test 15G 2.29G 12.7G - - 0% 15% 1.00x
ONLINE -One issue I had is that due to what I eventually tracked down as power issues, I had some corrupted data written to disk under my zfs pool (at the media write later), and I had dedup on.
So dedup, unfortunately, actually made it REALLY suck to fix, because I couldn’t even copy a new version of the file to the same pool! It kept nuking the duplication, and keep the old bad data and I then couldn’t read the copy. :s
It even did this after I deleted everything, because prune couldn’t remove the bad underlying entries because it was having a media failure.
So delete files, scrub, put new files on resulted in them having the exact same failure.
When I nuked the pool and recreated it, it was all fine though.
So yeah, be careful with dedup.
At the time, it was the difference between slow and impossible: I couldn't afford another 2x of disks.
These days, the pool could fit on a portable SSD that would fit in my pocket.
Careful, file-based dedup on top of ZFS might be more effective.
Small changes to single, large files see some advantage with block based deduplication. You see this in collections disk images for virtual machines.
You might see that in database applications, depending on log structure. I don't know, I don't have that experience.
For most of us, file-based deduplication might work out better, and is almost certainly easier to understand. You can come with a mental model of what you're working with, dealing with successive collections of files.
Even though files are just another abstraction over blocks, it's an abstraction that leaks less without the deduplication.
I haven't used a combination of encryption and deduplication. That was Really Hard for ZFS to implement, and I'm not sure how meaningful such a combination is in practice.
Hmmm, that 3.5GB/s sound low. From rough memory of doing initial storage benchmarking of our "new" Hetzner dedicated boxes a few months ago (AX51-NVMe, https://www.hetzner.com/dedicated-rootserver/ax51-nvme), they were giving about 10GB/s with mirrored NVMe drives.
Just logged into one of those boxes now, and it's running 2x 1TB Samsung PM9A1 drives (https://semiconductor.samsung.com/ssd/pc-ssd/pm9a1/), compression is on (lz4), and dedup is off.
(Didn't do any real tuning at the time, as these specs were already far in excess of what's needed for these servers.)
Would enabling lz4 compression be useful for your use case?
My understanding is that the OpenZFS project devs feels that the answer is an emphatic no, it should not exist, but they're committed to backwards compatibility so they won't drop it. (Take with a grain of salt; that's an old memory and I can't seem to find a source in 30s of searching.)
They discussed it, along with some options for a background-scanning dedup service (trying to find potential files to dedup), in the February leadership meeting[3].
[1]: https://openzfs.org/wiki/OpenZFS_Developer_Summit_2020_talks...
Why is dedup even present when the primary use case (backups) is better served in every way by snapshots?
In theory, it could also really help with virtualized disk workloads where there may be a lot of duplicated data from the base OS, but you can't use a snapshot (easily) because windows won't run from a zfs filesystem. You could maybe do snapshotting on zfs volumes, but that's not as flexible as a dedupe that worked as imagined.
Personally, I think online dedupe ends up being too expensive in memory and computation and ends up missing things because of divergent block sizes or offsets as another poster mentioned. ZFS doesn't support an offline dedupe, but I think btrfs does. That might be more interesting. It's still expensive to find duplicates, but it's possible, and it'd be neat to be able to rewrite the metadata to refer to a single copy and free some space, maybe.