How Linux 3.6 Nearly Broke PostgreSQL
lwn.net
lwn.net
Every time I read or hear about unintended-consequence incidents like this one, I'm reminded me of Jean-Baptiste Queru's essay, "Dizzying but Invisible Depth" -- highly recommended if you haven't read it.[1]
--
[1] https://plus.google.com/u/0/112218872649456413744/posts/dfyd...
You just arched an eyebrow. Simple, isn't it? Well, when you know a bit about biology...
http://www.spinics.net/lists/linux-fsdevel/msg58725.html
It has gnulib (not a test suite, but a very comprehensive POSIX API test).
It has the poor saps who have to use it.
However none of these things gate commits to the kernel.
I don't see how this is so bad, it seems like the best solution too me. If you're writing a specialized high performance piece of software I feel like the application developer should be the one tasked with making sure the kernel knows certain things about it's application. It's pretty clear a project like postgres is doing all sorts of tricks and optimizations already, I don't see how this would be any more or less burdensome.
Overall I feel like it's a fair trade-off to have kernel be told specific things by the application so it can make the better scheduling decisions vs it having to guess and potentially make poor decisions at the expense of most common applications.
After all the purpose of the operating systems is to solve the hard nuts for you, isn't it?
First, there's no POSIX standard way to communicate this scheduling request to the kernel. So you either have to add a new API or add some "secret knock" to an existing API that will trigger the desired behavior. Neither of these encourage portable code, and neither help existing, deployed applications, which would have to be modified to get the desired result.
Second, it introduces a new axis for regressions. With today's scheduler, maybe setting this "I'm a control process" flag makes your application faster. But maybe the new scheduler implementation in the next major kernel release actually causes your application to run slower with this flag on than off. Some other applications might see the reverse.
Linux doesn't fully comply with posix and has many of it's own apis. Cross platform applications either abstract this away themselves or libraries do it for them. I doubt a database doesn't use a lot of os specific features.
I'm even more surprised by "some benchmarks show it's faster, let's merge it".
Maybe they could try something larger than subsets of 2 CPUs?
I really prefer Linus suggestion, even if it's the hardest.
But I agree with you about the Linux patch. Considering the performance impact, it seems irresponsible not to implement the best solution. The kernel isn't some web app where everyone gets the latest version automatically. Bad code in the kernel sticks around for a long time.
Here are two really simple mutexes: a wrapper around the pthread functions, and a no-thought-required spinlock:
class Spinlock {
private:
bool state;
public:
Spinlock() { state = false; }
void lock() {
while(!__sync_bool_compare_and_swap(&state, false, true));
}
void unlock() {
state = false;
__sync_synchronize();
}
};
class PthreadMutex {
private:
pthread_mutex_t mutex;
public:
PthreadMutex() {
pthread_mutex_init(&mutex, NULL);
}
void lock() {
pthread_mutex_lock(&mutex);
}
void unlock() {
pthread_mutex_unlock(&mutex);
}
~PthreadMutex() {
pthread_mutex_destroy(&mutex);
}
};
I used these to protect a not-optimized-at-all dispatch queue built on top of std::deque (single reader, single writer, reader busy-waits until work is available, work is an empty function). I did this test on Mac OS X 10.6.8 (with Apple
s gcc 4.2.1 build).A run using PthreadMutex:
0,5944405
1,1627173
2,135598
3,363801
4,3360217
5,5773374
6,5727638
7,5953505
8,4567648
9,5106300
Using Spinlock: 0,154817
1,187507
2,145171
3,102454
4,180604
5,155360
6,128448
7,82418
8,49538
9,39560
All times are in microseconds to run 1,000,000 iterations of the test. Only time to enqueue is measured.I've seen similar results with a very similar test on FreeBSD, though I don't have a box to retest on at the moment.
I can only conclude it's not at all unreasonable for PostgreSQL to use its own spinlocks.
http://msdn.microsoft.com/en-us/library/aa175385(v=sql.80).a...
Keen observers of database history may remember Sybase. Sybase made a similar decision about doing their own scheduling, rather than relying on the operating system. Oracle at that time let the OS do the scheduling. The former turned out to be a strategic mistake.
That Sybase implemented its own threading was actually quite an astute decision at this time. We're talking early to middle nineties when threading provided by the operating systems was in a, charitably put, pretty unstable state. We're talking of a time when a gig of physical memory was not that bad for a database server.
The reason why Sybase failed (relatively to Oracle at least) had nothing to do with threading, but with the fact that Sybase did not support row level locking and adamantly refused to implement it.
From a purely theoretical position Sybase was right. If your database application is well designed, page level locking has a number of advantages. Namely in much less resource usages and less requirements for internal housekeeping tasks. While a page level managed database is not quite maintenance free it is much more so then when the data is row organized.
The problem, however, was that the real world is not a theoretical thing and the various application suites (SAP R/3 PeopleSoft, etc), which boomed around that time, absolutely required row level locking. PeopleSoft actually did run on Sybase, but performance was, again charitably put, difficult.
It didn't help that Sybase, at that time, released Sybase 10, which was an dreadful product, quality wise. From what I heard (and yes, this is hearsay) was that engineering implored on senior management to give it six month more time, which they refused to do.
While I never heard about data corruption on Sybase 10 databases, the quality was quite horrible. Couple this with Sybase' arrogance as a high flyer at that time chiding their customers for their own quality issues was not a smart move.
But the main issue was row level locking and certainly not the threading architecture.
Two additional points: The new Sybase kernel (15.7) actually uses OS threads as a default. You can still use the internal scheduler, but according to Sybase most installations should profit from the new kernel.
The other thing is that it's rather ironic that SAP bought Sybase (it's marketed now as SAP / Sybase), which somehow brings the whole story full circle.
I work with Sybase products since 4.2, which is early nineties and I worked for Sybase Professional Services from 95 - 99. Which makes me believe I'm somewhat qualified to comment on the issue.
I recall Philip Greenspun dedicated a substantial portion of his late 90's database-backed website book to that topic.
Sybase SQLServer (now ASE) was designed from the ground up as an OLTP database. It required that you keep your write and update transactions really short.
I've seen - and worked on projects - that had an amazing throughput. They where, however, designed from the ground up, to perform well on the underlying database.
Where Sybase' concurrency turned to dreadful, was when you ran chained transactions on isolation level 3. All I can say is: don't try this at home, folks.
Also, Oracle's locking mechanism didn't come free. I remember (and my Oracle knowledge is really minimal) the dreadful, overflowing rollback log, whoms sizing was a science of its own.
I'm not saying one is better then the other, but it points out quite nicely the impact of desing decisions. And how they always come with a price.
Postgres as a project is very keen to offload tasks to the operating system when operating systems at large are not unacceptably slow or broken, and sometimes even when they are, but nobody has the resources to do anything about it. That's why it doesn't schedule its own writes using O_DIRECT, taking a hit in buffer copying from shared memory, unlike virtually every proprietary database.
I did not see anyone in the lkml discussion blaming PostgreSQL for having an own spinlock implementation.
It's essential if you have N processes contending for a single mutex which they will hold for very short periods of time. Asking the kernel to put you to sleep until the mutex is available means progress is limited by the rate at which the OS can wake up processes. If the mutex is only going to be held for a few dozen cycles (say, to increment the heads of a few queues) then the throughput cost could be considerable over simply spinning a few nanoseconds in user mode until the mutex is available.
And yes, the need becomes more acute if you want to be sure you'll get reasonable performance across a broad range of platforms and their corresponding scheduling policies.
FWIW, the discussion in kernelland was based on the assumption that this regression was the kernel's problem. Nobody there suggested telling the PostgreSQL developers to come up with a new locking scheme.
OS scheduler optimizations are difficult. Often what makes the desktop nice and snappy makes background stuff slower. There are always trade-offs. Its also allows vendors to sell expensive versions of linux with different schedulers (redhat mrg...cough..) The Completely Fair Scheduler with its tree of process seems to work quite well though.
It seems like they were trying to optimize for specific hardware (the link to "scheduling domains" was interesting) when cpu swapping. (2 cores vs 2 sepearate cpus...) good intentions, but..
Sometimes its useful to let users explicitly control which cpus processes can run on (process affinity). On the HPUX variant we used they let us set up groups of cpus and then map processes run on those cpu sets. you could also select scheduling of each process startup. It was a pain to get things running, but in the end it worked great. Manually selecting the wrong scheduler and process priority could result in some processes running terribly however.