PostgreSQL Benchmark on FreeBSD, CentOS, Ubuntu Debian and OpenSUSE
redbyte.eu
redbyte.eu
A few thoughts:
* You weren't testing OSes which the subject implied, you were testing Linux kernel variants and their stock OS configurations/kernel scheduler setups, and FreeBSD was tossed into the mix. Whether you are running Ubuntu, CentOS, Debian or whatever you should have the Linux kernel tuned to perform well, so adding the distribution as a variable is just a red herring. I'd be more interested in removing that variable and comparing different storage configurations (such as XFS, and LVM).
* Clients connecting over the network adds a huge variable at play (the network) -- ideally you would want to remove this.
* I may have missed it, but it wasn't clear if you had a warmup period to your benchmarks. Especially with a system like ZFS which has COW, you need to do a few benchmarks on the same blocks first, to break past the cache.
As a counter-example, I could easily cherry pick versions, tunables and patch sets to make the numbers go whichever way I want so these types of comparisons aren't that useful unless someone is dropping a big delta on the floor with out of the box settings vs another.
As a FreeBSD developer, I will actually tell you that Linux could be selected to graph massive wins by cherry picking hardware with very high core count and several NUMA domains. But even then, by selecting kernel features (which could be innocuously hidden in a version number/vendor patch set) you can cherry pick large swings. That said, FreeBSD+ZFS+PGSQL (https://www.slideshare.net/SeanChittenden/postgresql-zfs-bes...) is a joy to administer, and is unlikely to be the weak link in a production setup if you stick to a two socket system. There is a lot of work going on in HEAD that is relevant to this workload in the memory management subsystem, including NUMA support. And some TCP accept locking changes that'd be relevant for TCP connection turnover.
I would bet this would wipe out the SUSE advantage:
# cat /proc/version
Linux version 4.1.12-112.14.1.el7uek.x86_64 (mockbuild@) (gcc version 4.8.5 20150623 (Red Hat 4.8.5-11) (GCC) ) #2 SMP Fri Dec 8 18:37:23 PST 2017
The "RedHat-compatible kernel" should have identical performance to CentOS.
# rpm -qa | grep ^kernel | sort
kernel-3.10.0-693.11.1.el7.x86_64
...
kernel-uek-4.1.12-112.14.1.el7uek.x86_64
...
I realize that many people don't like Oracle Linux due to their unauthorized appropriation and support of their RedHat clone. It does bring new functionality to the table, however (primarily Ksplice) and has great support for the eponymous database.
When RedHat 9 support ended, it never touched my personal systems again. Even my CentOS habit has quelled.
It also seems unfair to compare Ext4fs on Linux with ZFS on FreeBSD. Ubuntu ships with ZFS included. There's also Btrfs.
Also, its known that Netflix uses FreeBSD internally, and I'm curious why.
I think pacman can get you their PostgreSQL package easily enough.
I have screenshots at the end of this (unpublished) article:
For a lot of people who don't have time/understanding to play with things much beyond using stock versions and configuration, this could still produce a useful benchmark. People who have the knowledge, confidence, and time, probably won't be reading the article at all as they'll have already performed their own less artificial tests (i.e. benchmarking with their own application using live-like data and load patterns).
Though that is my argument against any benchmark like this: it is at best an indicator of peak activity under very specific conditions, a starting point but it doesn't really represent my application with any precision nor accuracy.
> it wasn't clear if you had a warmup period to your benchmarks
I would agree that is a significant point.
I don't think that a Linux distribution is just a variable and the only thing that differs is the kernel version. Each distro made its own choices, for better or worse...
As for the clients connecting over the network - that was exactly my point. My idea was to benchmark in conditions similar to production deployment. I doubt that many production systems connect over unix socket.
And for the warmup period, as you can see in the benchmarking script there is a 30 min warmup period before I start to record the results.
It is from the “Transaction Processing Performance Council”, correct? At least, that’s what they call themselves at tpc.org.
Otherwise, interesting results that I think need further examination.
Are you also setting a max limit for the ARC? You don't want postgres and the zfs ARC to compete for memory. I wonder if this impacts FreeBSD's poor performance in the read intensive tests.
The great thing about PostgreSQL is using tablespaces, so you can put tables and indexes on different ZFS filesystems, with different hardware(disk vs ssd), recordsizes and compression options.
I wonder what it will take to fix that in Linux?
> 200GiB read only: read only test, the dataset fits into the PostgreSQL cache
How come a larger dataset fit into the PostgreSQL cache and not the smaller one ?
The only difference was the amount of RAM PostgreSQL could use - as specified in postgresql.conf.
So, 74GB database does not fit in PostgreSQL cache for 32GB instance and fits for 200GB instance.
Does anyone who knows FreeBSD and Postgres have any idea why it was so much slower for the read only tests?
so is ext4..
> and because it can be grown online
How this is different from how resize2fs can resize a mounted ext4 filesystem?
xfs may still be better than ext4 for some workloads, but "it's there" is not a reason why someone should use it.
In Linux one would need to benchmark BTRFS with LZO compression to benchmark compete against FreeBSD ZFS with compression.
It's also generically an unwise assumption because you have to know how fast the compression is compared to disk I/O. If, like almost all servers these days, you have more CPU than I/O capacity the compression overhead will often be buried in the I/O latency and if the data compresses well it it's easy to have it be faster because the I/O savings is greater than the CPU.
Since LZ4 was designed to be very fast, that seems like a reasonable bar to hit — I see single-core performance on old desktops in the 1.8+GB/s range.
If you want the ability to scrub (in such a way that detects errors over the whole storage stack), you must sacrifice hardware raid. You must also pay the fletcher/sha256 tax.
ZFS enthusiasts advocate that critical data absolutely requires those sacrifices. I'm inclined to agree.
Anyone from FreeBSD devs have any thoughts on why we're so far behind?