Why some NoSQL DBs only let you perform transactions on a single data item
dbmsmusings.blogspot.com
dbmsmusings.blogspot.com
Joe Hellerstein gave a wonderful keynote at ACM SoCC last year that is a great read for all who are interested in this subject: http://db.cs.berkeley.edu/jmh/talks/SoCC14-keynote.pdf
MarkLogic has had transactions since version 1.0. The same mechanisms that enable multi-statement transactions are also critical for overall reliability. A recent Gartner study ranked MarkLogic as #3 for reliability (not NoSQL databases, all databases).
Many on HN may have not heard of MarkLogic as we are not open source. We are, however, the largest NoSQL ISV by employee count (and probably other measures also.)
I am VP, Product Strategy for MarkLogic (ExIngres, ExCohera)
The closest we have to a "comparison" is http://www.marklogic.com/what-is-marklogic/features/ those pages do actually end drilling down to documentation.
I'm not sure the buckets provide much guidance in real life. One can configure a system, for an insert only batch load, to prioritize T without giving up FI for the data and transactions that aren't part of that load. If an FI system can load faster than an IT system, this distinction ends up not actually mattering. Plenty of folks who need FI are using products that don't provide it because of culture + budget.
Of course, the next question is whether you actually need scaling. Outgrowing a database is a "problem of success"; if you have that problem you can throw money at the problem. It's much more important to concentrate your efforts on achieving success; choose your database focusing on your business problem, not on hopeful future potential problems.
RavenDB is a nice middle ground: eventual consistency for queries for speed, but ACID for create, update, and load (one or more items by ID).
It uses transactions throughout, so a failure in the midst of 10 writes will rollback all of them, as one would expect in a traditional relational database.
[0]: http://ravendb.net/
[1]: http://ravendb.net/docs/article-page/3.0/csharp/start/gettin...
For commercial enterprise, Raven is effectively $788/core/year [0]. Quite reasonable, IMO.
Contrast this with MS SQL Server Enterprise, which appears to be $14,000/core one-time cost [1].
(Disclaimer: I'm a part-time employee for RavenDB. But I loved and used Raven on my own projects before becoming an employee.)
[0]: http://ravendb.net/buy [1]: http://www.microsoft.com/en-us/server-cloud/products/sql-ser...
Complex queries are eventually consistent.
Hope this helps. I'm @judahgabriel if you have more questions.
designers haven't solved all problem around transaction and performance, but that doesn't mean they have stopped trying
"As various NoSQL databases matured, a curious thing happened to their APIs: they started looking more like SQL. This is because SQL is a pretty direct implementation of relational set theory, and math is hard to fool."[1]
To handle massive loads, you can trade away transactional integrity and get increased performance AND a whole new range of complicated problems that was more or less solved for you, that you now have to take care of yourself.
For some applications, the tradeoff is not a problem - and in others it requires a lot of work, up to more or less your own implementation of distributed transactions.
I have this feeling that I can't shake off , that a lot of people think relational (transactional) databases are complicated, and fail to see why they actually are complicated. This probably is not the category of people that truly need to solve their performance problems though.
But a lot of services would also get by with a distributed DB with weak integrity guarantees. Does it really matter if a few upvotes get lost on a busy post? By the time you notice sporadic integrity issues in your DB, the startup will probably have failed or raised enough money to not care.
Unfortunately this application is new, so we haven't yet seen the benefit of schemaless data changes.
I'm not saying it's a bad decision or that MongoDB is bad in general. I'm saying it's a decision that should not be taken lightly. If you've done the analysis and feel strongly that MongoDB's the perfect fit, then cool.
However, I think people routinely discount the operational overhead that comes with using something like MongoDB. For example, they don't think about how it's much harder to find good devops people with experience maintaining a MongoDB sharded cluster, or they don't think about what happens when the CFO wants to hook the DB up to Excel and run some queries. And I'm not even talking about the additional development complexity due to implementing integrity checks and manual transaction rollback/compensation in your application code.
Those problems can be all be solved though. It just takes more effort/money. So, you need to be sure MongoDB offers advantages to offset that.
In my opinion, schemaless data changes are not even close to offsetting the overhead, and concerns about scaling beyond MySQL or PostgreSQL are probably premature optimization. Not that those are your reasons – just reasons I often hear.
And of course, I'm making a lot of assumptions, for example that the team has experience writing & operating systems on top of an RDBMS. If the rest of the team are MongoDB wizards and you're the only one new to it, then that's a different situation.
The E/R information between tables remains exposed and enforced, but the content details can be made fungible where desired. There is an extension to include JSON fields in queries, as well, though it won't be as fast as using a real column with an index on it.
If I was leaning towards NoSQL I would probably choose something rock solid like DynamoDB on AWS. That thing (mostly) scales by simply throwing money at it, on my last project that had tens of millions of users it worked very well, but it's definitely not cheap.
For my personal project I ended up going with PostgreSQL 9.4. The initial user experience is still as bad as I remember from years ago (after installing PostgreSQL on Ubuntu, neither the logged in user nor root have a default DB, can't even run psql, cannot create new DB's out of the box, and the tutorial(!) waxes on about "Architectural Fundamentals" instead of "Do x,y,z to install and insert a few rows of data").
With Postgres I'm able to have JSON type columns, where my unstructured documents are stored, as well as regular foreign keys, transactions and joins at the table level. So far I'm very pleased with how it's working out, I think it's a good middle ground between the flexibility of NoSQL and the reliability + querying power of SQL.
Note that this is a mainly packaging and entirely the maintainer's fault: They typically run the database under a "postgres" user, so you'll have to do "sudo -u postgres createdb" or whatever. PostgreSQL itself doesn't care; if you install it from source and start the postmaster process as yourself, then you'll have full access right away.
That said, Debian/Ubuntu packages typically don't set up a default environment for you as a convenience. It's debatable what it should do. Create a user and database for root? Postgres maps, by default, POSIX users to database users. If you run "apt-get install postgresql-server-9.4", who should it create database users for?
Also, I'm not sure you're fair to the documentation you're referencing, which seems entirely reasoanble to me. Section 1.2 of the 94 manual's tutorial, "Architural Fundamentals", is pretty important, and it's short. But if you skip it, you'll get to the next section, "Creating a database", which tells you what need, and the next one goes into the SQL stuff. Both sections work if you've installed it from source (or Homebrew, which also runs Postgres as yourself). But it would be silly to expect this documentation to tell you how to specifically do it on Ubuntu.
Which is fine, if you want to keep a high bar, but don't be surprised when people still think the product is "harder to use" than MongoDB or whatever in 2015. Blaming maintainers does nothing for the users, take ownership of the experience! If you want to be seen as approachable and easy to use, you have to actually target people new to PostgreSQL and not write the basic guide assuming you're compiling from source on Solaris or have a dedicated site DB administrator. The magic phrase "sudo -u postgres createdb" does not appear anywhere on postgresql.org that I can see, and as a experienced user you know that's not the end of it (can I use psql as my regular user after running it? No, of course not. There's many more entirely undocumented arcane incantations left!). I claim, it is in fact impossible to get up an running with no previous knowledge if all you have access to is the postgres site. Thankfully there's Stackoverflow, so at least some users are still getting through the gauntlet.
Sorry for the rant. I get a bit frustrated when I see a great project that I love dropping users through a bad onboarding experience (I know, contributions welcome...)
Just for reference, check out: https://docs.mongodb.org/manual/tutorial/install-mongodb-on-... https://docs.mongodb.org/getting-started/node/ In fact, you get different sets of documentations for the whole matrix of OS * ENV (Shell, Node, Python, C++, Java, C#). When the docs have their target audience down to a T ("Why yes, I AM a C++ developer on OS X! This looks like the perfect fit!"), it's easy to see why they're so popular.
Kidding aside, I suspect with a little discipline, one could design a PG DB that could be ported to Oracle (Exadata???) should the need, and budget ($$$$$), for that sort of scaling arise. Cheap and safe now, Expensive and safe later (if needed)
In theory there's no reason why we shouldn't be able to support a nice transactions API in the client drivers too, I just haven't gotten around to it yet.
In the ideal MongoDB implementation, an "order" would be a document of its own, and all the other parts would be attributes or nested structures. So you only really need a transaction for the document itself. Other documents could be inserted/updated at the same time without creating a conflict. That's the theory, at least. Martin Fowler has a good Youtube on "Introduction to NoSQL" that I sometimes show to my students: https://youtu.be/qI_g07C_Q5I
The other thing about NoSQL is that it's often used for write-once, then read-only applications. Twitter, Facebook, or Instagram would be good examples. When 1% of your queries are INSERTs and 99% are SELECTS (with UPDATE and DELETE almost never occurring), you really don't run as much risk of anomalies. I would not recommend NoSQL for something like a banking application where you're doing lots of updates and need to guarantee consistency.
Simply choose the most suitable solution, not the most hyped one. A mistake would be to sacrifise transactions via using some "general NoSQL database". Today, for real, you can have a single node capable of millions RPS on real-world scenarios with ACID transactions of arbitrary complexity, choosing solution like http://starcounter.io/. Please, please don't just take yet another no-transactions DB, which is no-transactions even on a single machine, having 4 db nodes on 4 cores completely separate as if they've been 4 different machine. Then you observe bad performance and start to scale things up, paying more and more for the cloud. Not the best idea to spend time and money.
And, if you DO really need distributed transactions, then it mostly means they'd be driven by a logic of your subject domain. E.g., you might have one department in Sweden, one in the USA, then you need to manage distributed accounts in the right way, where "right" is up to your banking policy. However, if you need to scale reads, there is just no problem of doing so within "no-distributed-ACID" solution. The same time, if you need to scale writes, doing distributed transactions isn't a good idea either, as you've seen from the topic starter article.
So, right tool for the right job.
Outside of the brackets I keep a topic of fighting with latency via distributed transactions. I mean those things around caching, CDN and async replication. Distributed transaction isn't a remedy there at all, since it doesn't patch speed of light by any means.
[].slice.call(document.querySelectorAll('#post-body-2680112629272717029 *')).forEach(ele => ele.style.color = 'black')And before us FoundationDB delivered the same feature (although implemented a bit differently).
The truth is that there is a market for databases with "unreliable" writes and there are easier to design. That's why you see a lot of them.
Given that distributed transactions necessitate distributed coordination, it would seem that there is a fundamental tradeoff between scalable performance and support for distributed transactions. Indeed, many practitioners assume that this is the case. When they set out to build a scalable system, they immediately assume that they will not be able to support distributed atomic transactions without severe performance degradation.
This is in fact completely false. It is very much possible for a scalable system to support performant distributed atomic transactions.
In a recent paper [http://cs-www.cs.yale.edu/homes/dna/papers/fit.pdf], we published a new representation of the tradeoffs involved in supporting atomic transactions in scalable systems.
Do you use 2PC or something like Raft to handle distributed transactions?
From the docs it sounds like everytime a node joins/leaves the cluster goes into a (brief?) unstable mode and failures can happen, that doesn't sound great.
We indeed stopped using LevelDB. It has been used for a while but we couldn't work around some speed issues. LevelDB is good for many scenarii when properly configured, though.
The database is multilayered, LevelDB was the last layer, most parallel operation were managed by a massively parallel in-memory database which used LevelDB as a backend.
We do 2PC to handle distributed transactions. Raft breaks our consistency model.
When a node joins/leaves the cluster you might have some transient failures which are generally absorbed by the protocol.
I hope I answered your questions and feel free to give it a spin next week, we're about to release a major update!
Yes, "better" is subjective, but rethink has a good page detailing the differences (between rethink and mongo):
http://www.rethinkdb.com/docs/rethinkdb-vs-mongodb/
Feature-list type breakdown:
http://www.rethinkdb.com/docs/comparison-tables/
And pay attention to the guarantees that it provides/doesn't provide when you pick. I think this is a big part of what separates novice developers form middle-tier developers. There are always tons of choices, for just about everything in software these days, and your job is to figure out how best to solve which problem you're facing, in a way that future generations won't hate you for.