Databases: The Einstein Hypothesis - Distributed Transactions Cannot Exist
nerds-central.blogspot.com
nerds-central.blogspot.com
There is even a paper about using Paxos specifically for distributed commit: http://research.microsoft.com/apps/pubs/default.aspx?id=6463...
In practice, it is not di±cult to construct an algorithm that, except dur- ing rare periods of network instability, selects a suitable unique leader among a majority of nonfaulty acceptors. Transient failure of the leader-selection algorithm is harmless, violating neither safety nor eventual progress. One algorithm for leader selection is presented by Aguilera et al. [1] "
" The algorithm satisfies Stability because once an RM receives a decision from a leader, it never changes its view of what value has been chosen "
Fundamentally, the idea is that the leader is the transactional arbitrator using the RM as the recorder of that transaction.
[0] http://en.wikipedia.org/wiki/Consensus_(computer_science)
Of course distributed systems can be built, of course there are mechanisms to reach consensus between nodes.
However such a system needs to be designed from ground up to cope with partitioning. And one MUST accept that there will be conflicts in the system - and that some of them won't be resolvable automatically or in timely fashion. Its a trade off that cannot be avoided.
And I have sat in more than one meeting where customers and my bosses demanded that I violate laws of physics and provide them with a distributed solution that will have no errors whatsoever.
Edit: And work faster than a single node solution to top it off.
Here's the link to the proof that no deterministic protocol achieves distributed certainty in presence of network failures, even for two parties only: http://en.wikipedia.org/wiki/Two_Generals%27_Problem#For_det...
And yeah, the misconceptions around that are huge. I once had to expose a totally-not-database-related API as a JDBC-compliant driver because someone insisted it would provide transactional consistency.
No. The messages could as well be delivered instantaneously, the impossibility proof does not take advantage of non-finite delivery time.
Oddly - the project failed...
Just because a large company says they do something - does not mean they actually do. Distributed Transcation as said by RDBMS companies actually is 'Distributed as long as nothing goes wrong'.
Of course, that's the point of ACID. When something does go wrong, your database will rollback the distributed transaction and throw a scary ORA error rather than commit inconsistent data across the cluster. I've never known a properly configured Oracle system to break ACID.
I build these systems for a living, and I am quite certain that they can be made to work. Oracle is hard to set up, no doubt, but I would never blame it for the failure of a project...