Comparing Ceph, Linstor, Mayastor and Vitastor Storage Performance in Kubernetes
blog.flant.com
blog.flant.com
IME Ceph really starts to shine on a) slightly bigger setups and b) maintenance like upgrading nodes and swapping faulty disks, that's just a bliss if one doesn't violate a few basic rules (like maybe don't use 2/1 replica (two copies but return already after only one got written)). But also smaller cluster can work fine and can be quite performant, well at least if the interconnect bandwidth isn't one that was specified at the end of the 90s..
Also, ZFS 0.8.6 is from 2020, lots of stuff happened since then, a bit odd to post a new benchmark using almost two-year-old software.
[0]: https://www.proxmox.com/en/downloads/item/proxmox-ve-ceph-be...
That's probably a side effect of using Ubuntu 20.04, just like having a 5.4 kernel (which is likewise getting old). Although yes, that probably isn't helping, and it would be interesting to do the same test and change nothing but moving to Ubuntu 22.04 and the corresponding kernel and ZFS versions, both of which should help.
I don't think ceph ever does that.
IOW. 2/1 will never ever make sense, it's just crying for data loss, if you want to save on data usage, and if you also don't mind the (slightly) higher CPU load, you can use erasure coding instead:
https://pve.proxmox.com/pve-docs/chapter-pveceph.html#pve_ce... https://docs.ceph.com/en/latest/rados/operations/erasure-cod...
> Sets the minimum number of written replicas for objects in the pool in order
> to acknowledge an I/O operation to the client. If minimum is not met, Ceph
> will not acknowledge the I/O to the client, which may result in data loss.
> This setting ensures a minimum number of replicas when operating in degraded
> mode.Creating on-premise infrastructure a year ago we went for 2x25Gbps network. This or even 100Gbps seems to be current "mainstream" and 10GbE is definitely not enough for NVMe speeds available today.
A couple of years ago I could see volumes on Linstor getting completely stuck and unrecoverable whenever the network was getting busy or unstable. Nodes reboot were a nightmare too.
Have a setup now with their Piraeus operator[1], Kubernetes >= 1.20, rancher and calico, and it seems to be very stable. XFS have been giving better results too. Still, better not to try too many reboot loops on the nodes.
A lot of the corner cases nowadays are related to not being able to scale the cluster, odd behavior with CRD's of users and other components that require you to run ceph admin commands directly through a proxy. The annoying thing about a lot of the "Cloud Native" projects is that you will spend a lot of time going through github issue lists.
There is a Taiwanese company that provides extremely cheap OEM ceph racks though. I've noticed that in terms of maintenance and performance something like that or a cheap TrueNAS almost certainly ends up cheaper than production workloads, unless you don't care about the data integrity of your ceph workloads.
So the answer to your question is that your question was wrong, you are comparing different layers.
you can actually access ceph over nfs, however usually you will try to use a ceph aware access method, ether the ceph filesystem driver, or the S3-like library depending on what your application wants.
And you will need one. permamently. because this stuff breaks in so many spectacular ways... i.e splitbrain. a joy to resolve.