Can anybody explain why I should go with this instead of Pg?
To the submitter or anybody else sufficiently knowledgeable: When will ORMs get support for this? I use SQLA but could port Django or ActiveRecord code.
Can anybody explain why I should go with this instead of Pg?
To the submitter or anybody else sufficiently knowledgeable: When will ORMs get support for this? I use SQLA but could port Django or ActiveRecord code.
If I understand correctly, their SQL engine can scale through the use of a proprietary peer to peer technology. This would be new and would indeed make it possible to solve bottlenecks and scale the SQL database to a very large number of nodes.
However, I think it misses the point where SQL is simply not always needed. They will always be slower than a NoSQL engine that scales well.
Your example of ORMs is on spot: do you really need SQL as the back end of an ORM? If not, why pay the relational tax?
No silver bullet.
There is no intrinsic reason that ACID transactions, with or without SQL, can't scale, just that up until now, it hasn't.
And there's a reason that until now, it hasn't. From the beginning of time, academic computer "scientists" have confused the terms serializability and consistency. In short, serializability is a sufficient condition for consistency, but it isn't a necessary condition. If you design a system that enforces consistency without requiring serializability, it scales. Period. Legacy RDMSes don't work that way, but that's their problem.
If I may reduce to a concrete example:
Posit shared variables x and y, both zero, and simultaneously do (on threads T1 and T2)
T1: atomic { y = 1; return x; }
T2: atomic { x = 1; return y; }
where 'atomic' indicates a transactional operation.Traditional ACID rules would state that either T1 takes effect before T2 (and so (T1, T2) return (0,1)) or T2 before T1 (so they return (1,0)). Other interleavings are not permitted -- returning (0,0) or (1,1) would violate consistency.
What it seems to me that you are saying is that both (0,0) and (1,1) might be OK, depending on your application domain -- maybe the app can specify a consistency rule that says that (1,1) isn't OK. Then you can design databases that follow application-directed consistency rules.
At first blush this seems impractical, so I am hoping you can elaborate on the point you were trying to make.
This would imply read-scaling, but I don't see how this makes writes go any faster without reducing consistency. What am I missing?
Let me give a simpler example: Database with one table of one field, a number. One transaction: count the number of records and store that number.
A serializable system will force one zero, one one, one two, etc. But a consistent system can have two zeros and no ones. Why? Because that's what each concurrent transaction saw? Nothing wrong with that. But if application semantics dictate that each value must be distinct, then put a unique index on the number and the system will enforce uniqueness. Automatically enforcing "auto-magic" constraints that nobody cares about is why serializability destroys scalability of distributed system.
In NuoDB all messaging is asynchronous and batched, making it very fast and efficient.
For more than you want to know, see http://www.gbcacm.org/sites/www.gbcacm.org/files/slides/Spec...
Someplace there's even an audio recording which I recommend if you're into self-abuse.
If I start with a value of 0, then run my transaction twice, can both transactions return 5? Or is it guaranteed that one will return 10?
note that it's only mutated values that are checked for conflicts, and only against other mutations (this is why it is efficient - the number of checks required is small). so you can get weird behaviour when multiple values are read while different transactions change each - there's a good example in the link above. this is called "write skew".
the whole approach is, in a sense, exploiting poor phrasing of the ansi sql-92 standard, which doesn't actually require serialisation even though that is the most natural way to interpret it (as far as i understand things). so you can think of MVCC as "exploiting a loophole" that leads to a more efficient system, but one that is less intuitive. on the other hand, this is not new - it's already the standard behaviour for postgres, oracle, sql server, etc.
If I have two transactions and one of them is aborted because of the changes done by the other transaction what am I as a developer supposed to do? Retry? I hope not, because a retry is in other words serializing the execution. One after another.
So, if I have a system with a lot of concurrency (i.e. bank accounts and transfers) I'd better not use NuoDB because I'd get a lot of transfers aborted, not good. If I have a system with a very little write conflicts I'd go for serializability because that gives 100% consistency and will be fast anyway due to very little conflicts.
From the wikipedia link you posted, it is fairly clear that the Snapshot isolation is good when you don't need consistency. For that you'd have to either abort every time there is a conflict or introduce write-write conflict (ie. serializing).
(2) you would not get "lots" of aborts with bank accounts because each transaction is, typically, to a different account. you're only going to get a problem when two processes try to change the same person's account at the same time.
(3) what this provides is standard-compliant ACID sql, the same as postgres, oracle, sql server, etc. if you use any of those and don't have retry in your code for when transactions fail then you're already in a mess.
i am not associated with this project, but what you're saying doesn't really make sense. as far as i can see, you're criticising it for being the same as everyone else in the standards compliant, sql world.
re 1) and 2). I think they are related. The serialization can be done on account level. There is no need to have a single global write lock. This also implies that if the bank account transfers are typicaly to different accounts the serialization will be as often as the aborts. Hence very rare and the system will perform equally good in both cases. However, the serialization does not require further retries and doesn't force the application programmers to workaround the problems.
update COUNTER set VAL = VAL + 1
Let's assume the initial state is one entry with VAL = 0. If I run the update statement in two concurrent transactions I'd expect the result to be 2 when both transactions complete. However, I don't understand how that can be achieved without serializing those transactions.
These are the same MVCC semantics I used in Rdb/ELN (1984), Interbase (1986), Firebird (1999), and MySQL Falcon (2006). The implementation, however, is wildly different.
What is the benefit here? As an application developer I need to restart the transaction and hope it's not going to fail again - hence I'm serializing the execution on my own...
On top of all this, when it comes to constraints in relational model they can be complex and the probability of failing transactions is just going to grow. I can see this working only if I start relaxing on my constraints and redirecting transactions in such a way that conflicting transactions are coming to the database already serialized.
What about restarting transactions within NuoDB transparently as soon as an update conflict has occurred?
It's not clear where we're misunderstanding each other, but for clarity NuoDB manages such things as atomic operations and update conflicts without reference to the application. It's an ACID database.
Give us a call or attend a webinar (eg tomorrow) if you want to take a deeper look at it.
Barry Morris, NuoDB
That would be because "ORM" stands for "Object Relational Mapper".
I guess they billed it as "NewSQL" not "NoSQL" which is pretty absurd.
I'd vote that you go with Pg, because it's Awesome MostAwesomeDude, at least until this has been around for a while.