For what it's worth, one of the biggest "containerization" recommendations is to not run your database (example: Postgres) in a container, correct? Due to I/O performance decrease?
For what it's worth, one of the biggest "containerization" recommendations is to not run your database (example: Postgres) in a container, correct? Due to I/O performance decrease?
docker run --rm \
--name postgresql \
-e POSTGRES_PASSWORD=postgres \
-d \
-p 5432:5432 \
--cpuset-cpus="0-1" \
--cpus 2.0 \
-m=1024m \
--mount type=bind,source=$HOME/docker/volumes/postgres,target=/var/lib/postgresql/data \
postgres:12.2
native results on a digitalocean VM: $ pgbench -c 100 -j 2 -T 60 postgres
starting vacuum...end.
transaction type: <builtin: TPC-B (sort of)>
scaling factor: 1
query mode: simple
number of clients: 100
number of threads: 2
duration: 60 s
number of transactions actually processed: 17660
latency average = 342.761 ms
tps = 291.748482 (including connections establishing)
tps = 291.791293 (excluding connections establishing)
in a docker container: starting vacuum...end.
transaction type: <builtin: TPC-B (sort of)>
scaling factor: 1
query mode: simple
number of clients: 100
number of threads: 2
duration: 60 s
number of transactions actually processed: 13014
latency average = 466.928 ms
tps = 214.165822 (including connections establishing)
tps = 214.199201 (excluding connections establishing)
214tps in docker, 291tps outside of docker26% decrease, with a bind mount on ubuntu 18.04
I ran some casual tests using and found out there is a performance hit in using db binaries inside a Docker container due to Docker networking (different for different types of networking).
You generally are trusting a database to keep your data safe, so those things will contribute to data loss.
Remember the freakout about PostgreSQL not handling sync() correctly on Linux due to ambiguity in the man page? Having a networked filesystem + additional abstractions (like layers) etc only reduces data durability.
looks like docker with a bind mount has a 26% decrease in performance versus native
We do run databases that do 100k`s ops/s in containers but we don't run them in kubernetes. We just mount the VM hard drive in it.
With a healthy understanding of how the individual storage engines commit to disk, upgrading, backing up, etc. can be done in parallel and without impact to a running production system thanks to the power of overlayfs.