MongoDB CTO: How our new storage engine will earn its stripes
zdnet.com
zdnet.com
At this point why not just:
apt-get install postgres psql < "CREATE TABLE mongodb (key varchar[1024] NOT NULL PRIMARY KEY,value jsonb);"
then you get the best of both worlds.
It's like saying it's more important that the car drives reliably rather than everyone that needs to travel fitting into the car in the first place.
Naturally, there are plenty of people who do need to scale a lot, too, but presumably by the time they need to, they have a good idea of the problem and exactly how best to scale it.
When this is the case, you probably have enough money to hire someone who knows what they're doing with databases in order to set up a migration properly - you don't really need the ease of use configuration features.
SQL databases give data durability and transactions, integrity constraints, joins etc, which makes writing software a simpler experience (with a smaller chance of corrupting your data etc) which is generally a nice property to have when you're a small company that's moving fast and breaking things.
Replication is not very easy to setup in PGSQL. Sure, I can do it, but our product is an on-prem product that our users are responsible for configuring and maintaining. Having them configure PGSQL replication would be something they're not willing to do. MongoDB's configuration of replica sets is quite trivial compared to PGSQLs.
I bring this up every time someone says "PGSQL > MongoDB", even on PGSQL threads where PGSQL devs participate and they admit it's not quite there yet, but that they're working on it.
And all that said, it is getting demonstrably better to configure replication in PGSQL, so it's going to get there some day and I can't wait.
Do you have a link to such a thread where pgsql devs admit its inferiority?
Also, if you take a look at perf when running PG on for example Azure or EC2, you will realize that IO is pretty slow but nodes are cheap. So you want to scale out early.
Running stuff on a single machine sounds like a perfect single point of failure to me. The actual size of the data does not affect wheter single point of failures are acceptable in a business.
I've seen so many people recommend PGSQL, saying its very simple, but when actually asked about how to set up a simple cluster which fulfills the absolutely basic requirements, then everyone just responds somethin similar to what you wrote. I find it very annoying to be honest.
How many QPS does it perform?
What's your read/write balance?
It's better to say "key text" instead of "key varchar(1024)" and primary key already implies not null.
And I still wouldn't use Mongo!
From the article:
"In addition to compression and record-level locking, WiredTiger also gives MongoDB multi-version concurrency control (MVCC), multi-document transactions, and support for log-structured merge-trees, or LSM trees, for very high insert workloads"
I'd also love to see them add document joins at some point. That's a a big differentiating feature for RethinkDB, although I've found Rethink's syntax to be difficult to reason through.
This result is particularly relevant for MongoDB, since a document store tends to have large records containing multiple fields, as opposed to individual values being stored as separate small records.
And in cases where that applies (large records), nothing comes anywhere close to LMDB's performance.
(A preview of the on-disk test results I mentioned was presented at BuildStuff.LT last week http://symas.com/mdb/20141120-BuildStuff-Lightning.pdf pages 103-on)
[1] https://github.com/couchbaselabs/forestdb/wiki/Performance-R...
It crashed in all of my test invocations though, so I haven't looked at it since.
With perfect candor, I believe forestDB's announcement was extremely premature.
* SQLite - it's everywhere thanks to mobile phones.
* Oracle.
* Microsoft's thing has to have a big following.
* Mysql.
* DB2 is not as big as Oracle, but IBM isn't some lightweight, either.
MongoDB gets talked about a lot, and probably used among people trying new stuff out, but I think a lot of mostly silent people wouldn't touch something like that for their important data with a 10 foot pole.
and browsers...and tons of other pieces of software
SQL Server is definitely one of Microsoft's best products - as well as being an excellent database engine it also is used by SharePoint, the various Dynamics applications (CRM & ERP) as well as being the "natural" choice for application development in .Net
If I had to guess I would suspect that, at least in terms of revenue, SQL Server probably tops Oracle and DB2.
MySQL is most likely the most deployed thanks to the LAMP stack's proliferation across all sorts of small and medium sites. After that it really would be a toss up between MS-SQL who has traction in the SMB space due to it being used for sharepoint and exchange, and PostgreSQL
but lets be honest the most utilized database is Access.
http://db-engines.com/en/ranking_definition
It's kind of how I did http://langpop.com/ back in the day (subsequently sold and it looks like they don't maintain it).
I'd actually say that its feature set relative to other equally stable open-source DBs has gotten worse over the years, not better. Not because its gotten absolutely worse, but because other open source databases have gotten better faster.
As a long time observer of the database market with no skin in this particular game, in terms of regaining market momentum it is difficult to get past the reputation of being technically defective once that perception is broadly in place. To make matters worse, needing to completely replace the storage engine with a proper design will be viewed by some as an admission of that rather than an upgrade per se. It may be too little, too late for its long term prospects. By contrast, PostgreSQL is doing great right now; just about every company I talk to is using it somewhere.
Putting on my database engine designer hat, MongoDB did need to start from scratch with that storage engine; the original one was a naive design with a lot of problems. The sharding/replication design is still quite dodgy, which will hinder practical scalability. A new storage engine is a start but not enough if MongoDB wants to regain their momentum.