MySQL vs PostgreSQL
wikivs.com
wikivs.com
Just about every single claim provides a reference. For example, the sentence MySQL 5.1 natively supports 9 storage engines links directly to the MySQL documentation where they are listed.
Pretty slick.
You can get Galera integrated with an otherwise vanilla MySQL from http://codership.com/products/mysql_galera or integrated with Percona's XtraDB fron http://www.percona.com/software/percona-xtradb-cluster/
PostgreSQL replcation alone isn't comparable to MySQL with Galera. I don't know enough about the various extensions to PostgreSQL to know which ones could give you
* Synchronous replication * Active-active multi-master topology * Read and write to any cluster node * Automatic membership control, failed nodes drop from the cluster * Automatic node joining * True parallel replication, on row level * Direct client connections
in a reliable and perfomant manner. Keep in mind I'm just atarting to play with Galera, I haven't used it in production yet, but it's making my planned architecure much simpler to manage than the traditional MySQL or PostgreSQL replication approaches.
Why is that ?
When you move off the web site/content management side, PostgreSQL has long been the open source DB for complex business tools.
You also have Bucardo which can do master-master replication between two nodes, and a few other solutions out there.
One that doesn't randomly truncate your data, or allow insertion of nulls into not null columns.
SET GLOBAL sql_mode='TRADITIONAL';
?So still no guarantees and you more or less have to audit every app connecting to your db to make sure it isn't tampering with the sql_mode if you value data integrity.
The GLP-not-LGPL license sounds like a booby trap gladly left in place by Oracle from before they got it. I can't see Oracle wanting to do anything other than make it light and fast (at expense of correctness), either, since if you want "data integrity", they have a solution for you.
In all seriousness, Facebook's requirements are for high availability, with data integrity only mattering to the point at which it affects the usability of Facebook.
In other words, Facebook can tolerate a high degree of shit in the system before it becomes a problem. That isn't bad in and of itself, but you need to investigate your own needs before using something because Facebook/Google use it.
And I would imagine that most of those wouldn't tolerate any issues with data integrity.
The way I look at it is this:
MySQL is at its roots an SQL-like database specializing in content management. The sorts of data integrity problems that can occur in MySQL are entirely tolerable in a content management environment for the most part. It was designed to be a fast backend for web sites with an SQL interface. A lot of the issues it has are entirely due to that legacy, but those issues don't matter at all when you are using MySQL for content management.
Every single one of those examples are probably using MySQL for some sort of content management.
So it isn't just whether they can tolerate some crap in their systems. It is what sort of crap they can tolerate. If the MySQL gotchas don't result in intolerable crap, then it doesn't matter.
Now, personally I wouldn't run an accounting system on MySQL particularly if I was expecting many apps to access the same db. This is because the sorts of issues that MySQL has are real show-stoppers in these environments. But content management? Why not? Heck many of these data integrity problems may be features in these environments.
For example, suppose MySQL@Twitter truncates your tweets to the maximum length silently without issuing an error. Bug or feature? Suppose it truncates numbers in your accounting software? Those are two completely different cases and while they may in theory be comparable, in practice they are not.
I am sorry but are you that deluded as to think so many of the world's leading IT companies are going to choose a database that silently loses data ? Do you really think they are that stupid ? I mean come on.
Let me give you an example. I have a customer that uses MySQL for some processing on their web site and all data gets batched up nightly and entered into the PostgreSQL-based accounting system. The MySQL db is accessed only by the app which does the processing and the connectors to the PostgreSQL db. The PostgreSQL db has a lot more logic in it and is a lot more complex.
In their case, their environment is not subject to any of the MySQL gotchas. They can make sure the credit card processing software runs in an acceptable SQL mode, and if something goes wrong between the PostgreSQL export of reports and the MySQL public side (which has happened for reasons other than MySQL's fault), they can track down and fix the problem. That's what loosely coupled systems are about.
IT in this case is about risk management. How can you guarantee that specific important data will not be lost. You can do this with MySQL for some set of environments (either because the data truncation isn't a big deal, regarding your tweets, or because it can be retrieved from another source, or because there is only one app accessing the database so you don't have to worry about apps turning off strict mode), but you can do it with PostgreSQL for a much larger set of environments and that's what the complaint is.
> MyISAM was the default storage engine for the MySQL relational database management system versions prior to 5.5
This is what the anti-MySQL crowd is talking about - the dark days when MySQL lacked transactions, referential integrity, and concurrency. It was basically the SQLite of its day.
There's nothing wrong with all that if you just want a key-value store, which is what most web apps are.
And it generally doesn't "silently lose data" unless you are touching that data. It tends to just silently do stuff that doesn't quite make sense. Like if you post in a thread, and your "post count" goes up but your post itself fails, it's not a huge problem. Unlike in an accounting app, where you really don't want a transaction to partially work (e.g. money falling through the cracks).
When MySQL broke on a website, I think most people just assumed it was internet gremlins.
But PostgreSQL folks look at MySQL from a different perspective. It's a db-centric rather than app-centric perspective. In this perspective your db needs to guarantee that declarative constraints will be followed. In this perspective the db is the center of the environment, not the bottom tier of the app stack. In this perspective you could potentially have dozens of apps using a single db.
It isn't a matter of MySQL having grown up a bit (and it has). It is a matter of it not having outgrown the content management and/or single app per db environment in which it arose.
For every one of those sites, I suspect the database integrity plan is "meh, restore from a backup."
Look, I used to have to support mysql. Every couple months we had to increment our minimum required version because we found yet another query that didn't work right.
http://www.youtube.com/watch?v=Zofzid6xIZ4#t=04m10s
Click on that link and listen to Mark Callaghan, an engineer at Facebook, talk about their query workload. From a slide:
The Workload
- OLTP
- Fast Queries
- Simple Joins
- Fast Transactions
- Secondary indexes critical to performance
- Some rows very hotAnyway, I can see how you might have come away with the notion that MySQL is a glorified key value store by watching the first video which only briefly touches on their MySQL usage.
I believe that MySQL has become better since Oracle got a hold of it, their stewardship I consider to be much better than Sun's. Since acquiring MySQL, Oracle has put out an extremely solid release, MySQL 5.5, and are making steady progress on 5.6. If you lived through the early releases of MySQL 5.1 you have an idea of what a botched MySQL release is like.
I don't believe you do anyone on HN a service by spreading your gut feeling FUD about Oracle and MySQL.
However, you need to reword that slightly. MySQL does have such a working insert statement if you set strict mode. The problem is that apps can unset strict mode. Until that changes.....
So really you should word it as:
"One that can be guaranteed not to randomly truncate your data, or allow insertion of nulls into not null columns."
MySQL inserts in strict mode don't do these things. MySQL inserts cannot be guaranteed not to do these things however. Therefore this relegates MySQL, in my view, to a one-app-per-db environment because you cannot prove that your db constraints will be properly enforced and therefore have to independently verify this aspect in every app that connects.
Maybe they changed it in later versions of MySQL, but adding a column to a table become so lengthy for some of our projects that we switched for that reason alone.
"pt-online-schema-change modifies data and structures. You should be careful with it, and test it before using it in production. You should also ensure that you have recoverable backups before using this tool."
With that being the case what im doing is fine I guess.
"INSERT modifies data and structures. You should be careful with it, and test it before using it in production. You should also ensure that you have recoverable backups before using this tool."
"ALTER modifies data and structures. You should be careful with it, and test it before using it in production. You should also ensure that you have recoverable backups before using this tool."
It's just good advice to test stuff before you use it in production. Discarding an awesome product simply because the authors (responsibly) mention that you should probably test it out seems overly paranoid and would severely limit your choices. Percona Toolkit is the most well known and reliable set of tools out there in the MySQL ecosystem.
I started with MySQL, then used MySQL and PostgreSQL for a while. Then, when I started to do more "real" projects, I just got so frustrated with MySQL in so many ways at once. It wasn't that MySQL couldn't do it, it's that it was frustrating at every turn.
In my case, what caused me to drop MySQL almost entirely (around 2003 or 2004) was doing a few simple reports involving dates. Then I started using postgres and it was refreshingly consistent and flexible without so many caveats. And now I develop for postgresql, and the code is similarly consistent and flexible (and just all-around nice).
I have had a long string of positive experiences with postgres. It's hard to wrap them up into a feature checklist.
If something held you back from using postgres in the past, it's a good idea to watch the release notes to see if something new might solve that problem (or better yet, discuss on the lists so maybe it will be solved faster). But I tend to think that looking at long lists of features is a distraction.
re: replication, slony is horrible yet they focus on that. The slony author says you can daisy chain things, but that's a setup nightmare. Also, slony's n^2 communication gives you consistency guarantees, something I'm pretty sure mysql can't do, but most people don't need that. I much prefer bucardo. It's simple, easier to configure, fewer guarantees, but replicates much faster. I just wouldn't run a bank on that. However, how many people design bank software.
PostgresSQL has feature X, Y, Z
MySQL had X, recently introduced Y, and does not always have Z.
It's always MySQL catching up to postgres. And lots of it's db engines are not ACID. Its not the DB I thought it was.
Previous discussion: http://news.ycombinator.com/item?id=328257
Anything else?
What are you comparing it to?
It died, mostly because SQL was used by Oracle and IBM DB2 which were better marketed than Ingres. The DB world ended up standardizing on SQL because of this.
Worse is better, once again.
It looks almost like SQL; the main philosophical differences:
a) it embraces order (that is, every query result has an implicit "running index" field (called i, starting at 0).
This single handedly solves a lot of inconsistencies in practical SQL having to do with order, which is abhorred by the relational data model, thus not a first class concept, but is often required in practice, and thus inconsistently bolted on.
b) columns can nest, and do not have to have the same type - which means that aggregations like count, sum, max, min and distinct are not special in any way.
c) there's a simple underlying programming language, so if you have an intermediate select result that you need twice, you just give it a name by prepending a "name: " to the select.
temp: select from grades where age>10;
b1: select from temp where eyecolor=`blue;
b2: select from temp where eyecolor=`brown;
(compare to the mess that is correlated sub-queries, or alternatively, horrible "create temporary table x as ..."I'm not sure I fully understand (b) - do you have a link so I can learn more?
As you can probably guess, I'm very interested in this stuff!
First of all, there's some tautology here - a "relational database" (and similarly, relational calculus, relational algebra, etc), BY DEFINITION deal with sets of tuples ("relations") which are (again, by definition) unordered. So let's just ignore the word "relational" in this discussion.
> I think most databases do support the ROW_NUMBER() function now though.
Many do. But do compare (sql 2005):
SELECT * FROM
( SELECT
ROW_NUMBER() OVER (ORDER BY sort_key ASC) AS ROW_NUMBER,
COLUMNS
FROM tablename
) foo
WHERE ROW_NUMBER <= 10
To (kdb+ / q): select from tablename orderby asc sort_key where i<10;
And usefulness of order goes why beyond time-series data (and sorting): let's say you have a tiered pricing scheme for widgets you sell: order_size | per_widget_cost
----------------------------
1-9 | $2
10-49 | $1.9
50-199 | $1.8
Without embracing order, you:a) duplicate the range data (have a "from-count", "to-count" fields for each record, risking that you might have holes or overlaps)
b) not duplicate data, but have crazy subselects (among all with count > from_count, select the one with maximum to_count) or stuff like that.
When you actually have order, you have operator that embrace that order - e.g. kdb+/q's "bin" which finds the "bin" (as in "bucket", not as in binray) something fits in:
select unit_price[from_count bin order_count] from table
There are many other use cases involving running sums (e.g, you have a table of weights and priority; select the list of highest priority items whose weight sums to 100lbs or less).Order is really really missing in SQL, but it's one of those things people are not aware of because they've never used anything that does support it properly.
nested columns just means what it sounds like: that you can put anything in a cell (including lists, tables, lists of tables, lists of lists of lists of tables). Many one-to-many relationships in sql that need additional tables can just be done within the same table in kdb+/q
regarding Common Table Expressions - I wasn't aware of them, they do help a lot. The syntax is horrible, but I guess they do work...
edit: more info on kdb+/q can be found in http://kx.com/q/d/kdb+1.htm - you have to get to section 8 before they start discussing the query language, but it's short and to the point. There's a lot more in http://kx.com/q/d/ if you are interested.
a) duplicate the range data (have a "from-count", "to-count" fields for each record, risking that you might have holes or overlaps)
That's never how I have done it. Since this is a tiered pricing model, you just:
select * from widget_price where min_size < ? ORDER BY min_size desc limit 1;
We actually do something almost identical for sales tax rate changes in LedgerSMB. You look to the date of the transaction and take the most recent rate. No range types, etc.
Where range types are handy (and why I am looking forward to them in 9.2) is where you have to deal with things like a financial transaction that should be amortized over a period of time. This would allow you to adjust the transaction incrementally as a reporting, rather than an accounting, function. But I wouldn't use them for cases where you don't want overlaps or gaps. The best thing there is to just put in floor values and select the next highest floor.
There you go. Again order by and limit/offset do pretty much what you want without doing crazy inline views like you are doing in your example. BTW, I see inline views as an antipattern in SQL best avoided if you can.
nested columns just means what it sounds like: that you can put anything in a cell (including lists, tables, lists of tables, lists of lists of lists of tables).
PostgreSQL supports this btw although back in 7.3 or so if I remember, I found that tables with tuples in columns were write-only but that was fixed pretty quickly (first with a "don't do that" check and then with a real fix).
The point of CTE's is to give you a stable intermediate result set.
ORDER was there from the beginning, but CTEs, nested columns, window functions, connect-by recursive selects and similar stuff is being added because SQL and the relational model are actually quite limited when it comes to real world problems.
Of course, SQL will keep getting extended to solve real world problems; however, that does not mean SQL is "the best solution out there" (or "the worst solution except all others that were tried").
Suppose we have a relation R. We may order this relation physically in order to help the computer retrieve data faster, clustering on an index for example. However, clustering on an index does not mean that we are guaranteed to get the same order back when we do a select * from.... (we probably will but we aren't guaranteed to). And if we add a join or other relational transformation, the order will probably not be the same. In other words, ordering is outside the scope of relational math per se.
However that doesn't mean you can't have pre-ordered relations. It just means the ordering is meaningless as far as the math goes. The ordering may however be of great practical importance as the computer goes about grabbing the relevant tuples from the relation.
Similarly I don't see a reason why ordering can't happen after the relational math is done either, in this case for humans.
What this tells me is that the relational model is not entirely complete in itself in that ordering is an orthogonal consideration largely ignored which means there are certain questions you cannot answer directly with relational math (such select from R a relation L such that it includes the tuples with the five highest values of R(2) lower than 25. I think that's mostly what you are getting at. But that's a matter of relational math being incomplete for real-world scenarios, not SQL (since SQL implementations do provide for ordering).
a is accomplished by windowing functions. These can also do running totals among other things which is really helpful in accounting environments.
b has been supported since at least 8.3. I think a column can actually be an array of complex types in 8.4 and higher (it can be a tuple in 8.0-8.3 at least though I reported a bug in this in 7.3 which resulted in a "don't do that" check).
c is handled using common table expressions.
Examples for b and c:
CREATE TABLE foo (id int, value text);
CREATE TABLE bar (id int, values foo[]);
INSERT INTO bar (values) values ({row(1, 'test')}); -- not sure if this is quite the syntax. Might take some playing around with.
For C look at examples at http://ledgersmbdev.blogspot.com/2012/07/ctes-and-ledgersmb....
We use this extensively for things like relation to tree generation.
Real world usage makes SQL vendors extend SQL to make it less sucky; some of these extensions were later encoded into standard, and some are still proprietary.
Windowing functions are nice and all, but are a complex solution to a problem that would hardly exist if you actually embraced order as fundamental.
Ok, how about the most useful kdb+ extension (which I forgot about earlier): foreign key chasing: if table t has field a which has a foreign key reference to table s (which has field b which has a foreign key reference to table r (which has field c which has a foreign key ...)
in kdb+, you do:
select a.b.c from t
Does pgsql have something similar? Or do you have to spell out all the joins?You can build something to do this in PostgreSQL using stored procedures and a (a.b).c syntax but that's kind of advanced stuff. To do this you have to create a b function such that b(a) returns tuple of type b which has column or function c.
Example:
create table address (...)
create table employee (...., address_id);
create function address(employee) returns address as $$...$$;
select (employee.address).country from employee; will then return the country field from the address returned by address(employee).
So yeah, kinda, if you build your own.
edit: I would be willing to bet you could make an implicit join operator of this sort also but I haven't done so. I don't know what the performance ramifications would be of throwing this into the column list.
Yes. The only requirement for foreign key chasing to work is that it uniquely identifies one record in the foreign table. Whether that key is atomic or composite is of no consequence.
(internally kdb+ stores a pointer to the foreign record when it verifies the existence of said record on insert, so it doesn't have to do a join query - it always knows exactly which record to bring in. So in practice, it is very efficient regardless of what kind of indexes you might have in place, the size or the composition of the foreign key field)
> So yeah, kinda, if you build your own. I would be willing to bet you could make an implicit join operator of this sort also but I haven't done so.
pgsql is a wonderful beast. I really like it. And I would be even happier if they adopted some kdb+/q syntax and semantics, though I don't think that's likely to happen.
SQL is comparable to assembly language. Most people don't need it and wouldn't know how to use it properly anyway. These are the sort of people who use PHP and MySQL.
Nope. "Relational Algebra" / "Relational Calculus" / "The Relational Model" is about sets.
SQL is about bags (orderless like sets, but each item might be repeated multiple times). It's also about order ("ORDER BY" clause) in a horrible inconsistent way.
> SQL is comparable to assembly language. Most people don't need it and wouldn't know how to use it properly anyway
No, SQL is not comparable to assembly in any meaningful way (you could replace "assembly language" with "danish" in your statement and would be equally true)
While assembly language is more verbose, it is more fundamental than everything else in the sense that eventually everything must be expressed in assembly language (machine code, actually, which is equivalent to a proper subset of assembly language) to be executed. Thus, going down to assembly language might be more up-front work, but it is guaranteed that you can match or improve on run-time results from any other language.
SQL is an inconsistent abstraction that makes some things simple, some things hard, and some things essentially impossible -- and many of the things it does do, it does in a way that's inherently inefficient. (And don't tell me about the possible smart query optimizer - it doesn't really exist any more than Intel's Itanium optimizer that makes code properly utilize the VLIW; or a Unicorn).
edit: add the note about machine code.
* why bags?
* where's my closure under composition?
* THAT SYNTAX OH GOD THAT SYNTAX
I'd be much happier with a more mathematical language, rather than the godawful "english-lite" 3GL misery * Unknown value
* Not applicable value
* Value does not exist
This is a big issue, because you would expect operators to treat these cases differently. known || unknown is obviously unknown, but known || not_applicable should probably be known, and known || does_not_exist should be known. In sane RDBMS's there is a possibility of magic values which provides a sane way to handle not_applicable (for example an empty string as distinct from NULL and yes I am calling into question the sanity of Oracle). However, you still have the fact that the first and third cases are ambiguous although you hope not in any given query (the third case implies a missing value from an outer join), the ambiguity could in fact happen.This is a fundamentally broken aspect of SQL. The problem with ambiguity is that if your data is ambiguous mathematically, then it cannot be reliably transformed using math.
If you're typing raw SQL for getting reports out of a database then you're probably fine, but for web apps you're not typing queries, you're constructing them as strings using another language.
I've always hated the idea of writing one language in another, it feels like a giant eval() in JavaScript/PHP/etc. Not to mention it opens you up to injection attacks.
I like programatic access like MongoDB has, it certainly has its downsides but I prefer talking to a database via an API.
I've used ORMs and I've had to write my own once or twice, I really don't like them.
Even for things as simple as integers, do you have unlimited precision, unsigned value support and null values?
No other app code since we started this (at least code in the new framework) includes any SQL. All the SQL stuff is done by one simple function. The real programming is in the database for this interface. Our approach isn't fully developed. I expect we will be working on an object-oriented interface inside PostgreSQL soon which will make the queries look like:
SELECT (f).* FROM (select entity(?, ?, ?, ?, ?).save) f;
save(entity) will then handle actually saving the data.