6,287 karma · joined April 14, 2015
andres [at] hn dot anarazel dot de
Because that'll be too much of a scalability limitation. Rails etc are rather CPU heavy. In cloud environments it's also typically much more feasible to scale the stateless parts up and down than the database.
> Even if you have it on another machine, it's on the same switch, right (also an architectural choice)?
Yep, was on the same switch in my example.
> So latency should be sub millisecond, and irrelevant for a single request.
In my testcase it was well below a millisecond (2.8k QPS on a single non-pipelined connection would not be possible, it implies a RTT <= 0.35ms), but I don't at all agree that that makes it irrelevant for a single request.
Just as an example, here's the results of pgbench -S of a single client, with pgbench running on different servers (a single pkey lookup):
pgbench on database server: 36233 QPS
pgbench on a different system (local 10gbit network): 2860 QPS
It's rare to actually have that low latency to the database in the real world IME.
If the same pkey lookup is executed utilizing pipelining, I get ~145k QPS from both locally and remotely.
It can: https://www.postgresql.org/docs/current/logicaldecoding.html
It's not entirely from the WAL though, some catalog accesses are necessary for metadata (shape and name of tables etc).
You're right of course that there are lots of cases (missing metadata, synchronous operation like extending files, ...) where it's all offloaded to the wq.
It's definitely not as fast to start postgres as it is to start sqlite. Pretty much inherently - postgres has to fork a bunch of processes, establishes network connectivity etc. And running trivial queries will always be faster with sqlite, because executing queries via postgres will require intra-process context switches.
That's not to say postgres is bad (I've worked on it for most of my career), but there just are inherent advantages and disadvantages of in-process databases vs out-of-process databases. And lower startup time and lower "dispatch" overhead are advantages of in-process databases.
FWIW, since 15 postgres you can influence that behaviour with NULLS [NOT] DISTINCT for constraints and unique indexes.
https://www.postgresql.org/docs/devel/sql-createtable.html#S...
EDIT: Added link
Personally I'm not sure it's the right call, as it comes with some increased write overhead.
FWIW, nothing forces an extension to do so. I'm pretty sure there are several that do DML using lower level primitives.
I don't think security plays a huge role for shared memory in postgres - if an attacker gains arbitrary code execution in one postgres backend, the installation is hosed. No need to go through SHM to escalate to other backends, there's easier ways.
WRT capable: There's definitely substantial costs due to using inter-process shared memory. But it's more an architectural cost, rather than something that everyone has to bear while just doing mostly unrelated hacking. Some features get harder, less flexible and require more code.
FWIW, there's some work towards moving towards a threaded connection model. Still some way to go, but I expect it to happen eventually.
What you'd want is the ability to control the buffer for the "raw network side", so that asynchronous network IO can be performed without having to copy between a raw network buffer and buffers owned by the TLS library.
It also would really help if TLS libraries supported processing multiple TLS records in a batched fashion. Doing roundtrips between app <-> tls library <-> userspace network buffer <-> kernel <-> HW for every 16kB isn't exactly efficient.
Definition of macro: https://sourceware.org/git/?p=glibc.git;a=blob;f=include/lib...
Use: https://sourceware.org/git/?p=glibc.git;a=blob;f=string/memm...
For a bunch of other places -fno-builtin-* seems to be used.
https://github.com/postgres/postgres/blob/master/src/backend...
which uses
https://github.com/postgres/postgres/blob/master/src/backend...
If you want to be scared: Until not too long ago postgres' supervisor process would start some types of subprocesses from within a signal handler... Not entirely surprisingly, that found bugs in various debugging tools (IIRC at least valgrind, rr, one of the sanitizer libs).
> Re. SIGFPE, to be fair, it feels a bit like the "asynchronous vs. synchronous abort¹" thing on CPUs; synchronous aborts are reasonably doable while on asynchronous aborts you're pretty much left with torching things down far and wide.
Agreed, I think it's quite reasonable to use signals + longjmp() for the FP error case. In fact, I think we should do so more widely - we loose a fair bit of performance due to all kinds of floating point error checking that we could set up to instead signal.
Uh, huh. Details please?
Postgres doesn't quite have the span of the error, just a single location :). We can of course measure the length of the token at the error point, but that isn't quite the same, as the cause of the error does not have to be a single token.
postgres[122581][1]=# SELECT * FROM kdjfkdj;
ERROR: 42P01: relation "kdjfkdj" does not exist
LINE 1: SELECT * FROM kdjfkdj;
^> In order to do that, they purposefully don't explore the true planning space...they might explore 3-10 alternative ways of executing, whereas there might be hundreds or thousands of ways to do the same thing.
FWIW, often postgres' planner explores many more plan shapes than that (although not as complete plans, different subproblems are compared on a cost basis).
> While Postgres has explicitly chosen to not implement planning pragmas to override planner behavior, it would be really cool if you could have multiple planners optimized for different types of workloads,
FWIW, it's fully customizable by extensions. There's a hook to take over planning, and that can still invoke postgres' normal planner if the query isn't applicable. Obviously that's not the same as actually providing pragmas.