MongoDB 2.0 Released
blog.mongodb.org
blog.mongodb.org
so:
8 + 2 = 10, therefore 1.10.
It's best to view the version number as a string rather than a decimal number.
This is a somewhat old-school way of versioning and common in open source software. Personally, I prefer something more intuitive and don't see anything wrong with MongoDB being at version 2.0!
A major version bump should indicate major changes e.g. time to go through any code that talks to the server. If they just want to have a number that grows, they should do something like vRELEASE.MINOR.
Now in Mongo's case they might just be incrementing the major version number for the hell of it. But I'm assuming they actually planned this out.
Isn't it more about personal preference rather than having an "ancient mindset"?
"Please note version 2.0 is a significant new release, but is 2.0 solely because 1.8 + 0.2 = 2.0; for example the upgrade from 1.6 to 1.8 was similar in scope."
OpenBSD 5.0 is in its final stages before release and the 5.0 simply means it came after 4.9.
There are lots of other criticisms, but it's a fine data store, if you're dealing with data that fits it well. The best way to think of it is as a hash store - you can store N-deep hashes in it, index pieces of those hashes so you can find whole records ("documents") quickly, and that sort of thing. You can't perform table joins (or don't have to, depending on who you're talking to), but you mitigate that by designing your data such that you get what you need for a given resource in a single query.
Personally, I'm all on board with it for web apps. I think it fits your typical web app's data requirements far more closely than a traditional RDBMS does, and I'm using it very successfully in multiple production systems.
The biggest remaining fault, in the context of web apps (IMO) is the lack of transactions - if you have an application that requires transactions to ensure proper operation, don't use Mongo. Do recognize, though, that many transactional use cases are replaced with Mongo's atomic operation set.
While not ACID, "tension" documents are a good way to add a level of transactional support.
For example, if you want to update multiple documents, instead of modifying the original documents directly, you create a new document with instructions on how to update the original documents. Then, you roll through the tension documents and apply their changes. Any failures can be dealt with on subsequent passes, only removing the document from the scan once it has been successfully applied.
It is an eventually consistent pattern, but that is a tradeoff you have already accepted by choosing MongoDB in the first place, so it ends up working well for a lot of transaction use cases.
For example, in the canonical transaction example of Alice giving Bob $5 you could insert a document in the transactions collection that says Alice gave Bob $5. When you go to fetch Alice's pending balance you fetch her current document, then you fetch all pending transactions than she is involved in. Same for Bob and all other users. Then each night, you pull that days transactions and apply them to each user and update an "as of" field to make sure no transaction can ever be applied twice.
Now you could argue that this isn't fully consistent because it is possible for Alice to simultaneously give $5 to both Bob and Charlie even if she only has $7 in her account. However she will then have a pending balance of $-3 so it immediately reflects her current situation. This is similar to the real world example of overdrawing you checking account. If you wanted to, you could void one or both of those transactions if you never want to allow a user to go negative.
You could use optimistic concurrency control to ensure that only one tension document is written at a time while ensuring that your constraints are respected:
current_id = atomic increment global state_id
create tension document (but with active bit set to false)
check constraints
atomic:
if global state_id = current_id
set active bit to true
if active bit is false:
delete tension document and report error
This would provide isolation (you don't see tension documents that are not yet validated) and consistency (your balance is respected).The atomic operation in RDBMS is usually implemented in one very simple and fast SQL UPDATE query. I believe mongo must provide something similar too.
Although my bank would allow me to overdraw from my checking account, I have a higher margin/tolerance than my brother, so they must perform some constraints check ;-)
That flexibility comes back to haunt it in larger projects. A typo in a query can trigger a reindex of a huge collection, make moot GBs of index data, or invalidate other queries that work on the same collection. It also demands that the devs do lots of documentation on the data structures as you can't ask mongo to tell you it's schema, it's just a collection of documents.
And if you work with a loosely typed language, be sure to type your inserts as mongo is not a loosely typed datastore (coughPHPcough).
But for internal tools, and smaller projects, the speed of development is great.
Gotta sacrifice some flexibility for a bit of safety.
EDIT: Damn Google, you good: http://www.mongly.com/
Is it pure speed on a single machine? Is it the query interface? Something else? Datacenter awareness? Geographic support?
The key features that made me choose riak: - built in distribution/clustering with homogenous nodes - bitcask had the level of reliability/design that I was looking for. - better impedance match for what I was doing than CouchDB (which was what I looked at before choosing riak, but couch does view generation when data is added and I need to be able to do it more dynamically.)
What sold you on Mongo? What would you most like to improve?
(Please don't let this be a debate, I'm more interested in understanding the NoSQL market, what other developers priorities are, etc.)
Also, I really like the MongoDB docs.
Can you point me to a great getting started page for Riak (setting it up, loading in data, doing advanced queries, Python/Ruby libraries etc)?
It is being updated now for Riak 1.0.
- Well supported and very active community. It's a project that is clearly going to be moving forward for a long time.
- For my needs (structured data) it's super fast.
- The auto-sharding support is nifty.
- Geo-spatial queries out of the box
* Fast asynchronous writes for non-critical data (logs). * Replica sets allow for zero-downtime hot failover to a backup server when the primary is taken down or dies. * Simple queries are crazy fast, assuming your index fits in RAM. * find_and_modify() feature is awesome for some tasks.
But we're supplementing Mongo with PostgreSQL because of the following Mongo weaknesses:
* Concurrency is so-so: the entire database is locked during writes. * Weak aggregation capabilities. * Greatly reduced performance when the index doesn't fit in RAM. * Wasteful when it comes to allocating space (on disk and in RAM)
Thus I don't see how's that related to locking. My understanding is that the reason they still use a global lock are:
a) They're young and haven't gotten around implementing this yet.
b) Since writes are fast, this hasn't been as big of an issue as people might expect (it was for us, though)
And everything works as expected as long as you're fine with your writes being asynchronous. But once you call getLastError() requesting your writes to go through after every insert, then you start seeing the global lock bite.
"Download and unpack the package, run the binary, and you can instantly start playing with the examples from the console" is a hell of a good way to get people interested in your software.
wget http://downloads.basho.com/riak/riak-0.14/riak_0.14.2-1_amd64.deb
dpkg -i riak_0.14.2-1_amd64.deb
Our production configuration and setup/deploy recipes for riak are 241 lines in total, ~150 of which are the stock config files.While it's mostly just from experience I find that virtually no one makes a decision based on how easy it is to get up and running with a piece of software but it is an impression that stick's in people's minds and gets repeated when they talk about it later.
Riak
-> I love the fault tolerance
-> easier to scale with a dreamworld of hash-rings I wish I could have.
-> no indexes
-> no Geo-spatial indexes
-> I don't like link walking... it's not intuitive.
-> I want to use Riak (Lucerne/Solr like) full text search, but the docs are poor and have to read dozens of useless documents to get nowhere very quickly. It's a separate install that requires pre-commits?(which are never explained well enough to start using it).
-> installing on Mac OSX is painful (homebrew works/doesn't work depending on version, 32 bit or 64? how come I don't have include 32bit flags with older repo's, yet newer ones are only 64bit yet don't compile.... errors errors errors...ERLANG which? such a pain in the ass installing Riak unless you do it using specific build tools.
-> Only Joyent provides a hosted solution which is too expensive relative to its offering.
-> accepts any docs just set the content type.. handles original documents without conversions
-> easy to load balance behind NGINX
-> can choose between conflicting writes
MongoDB -> straight forward install.. up and running in no time
-> SQL like queries
-> intuitive indexes
-> Geo-spatial indexes
-> Affordable hosted services with more than one provider.
-> no full text searching over documents
-> have to fiddle with doc handling GRIDFS etc..
-> can't choose between conflicting writes (last one wins)
I wish I could have "RiaMong" -> straight forward install.. up and running in no time
-> fault tolerance
-> easier to scale with a dreamworld of hash-rings.
-> indexes
-> Geo-spatial indexes
-> Affordable hosted services at with more than one provider.
-> full text searching for documents
-> handle original documents without conversions
-> can choose between conflicting writes
Oh well maybe someday :)Just wanted to let you know Basho is addressing some of your wishes. Riak 1.0, which is coming out at the end of this month, has integrated riak search (so it is no longer a separate install) and also has integrated secondary indexes (on numerical values, so not full text search but integrated nicely.) I believe, it may be possible, to do geo-spatial searches using the new index system, but I've only thought about it, haven't tried it yet. (It's not built in, but I think its something one could build, and I'd like to build myself eventually.) They now have binary builds for installing on Mac OS X (IIRC) and I've been able to install via homebrew and source lately, so they might have fixed that. They've also got a new thing called riak_pipe which is really useful for certain classes of problems (its new so not well documented yet.)
Anyway, thanks again for your answer, and just wanted to let you know Basho seems to be addressing the issues you ran into.
That aside, you might like to know that for the 1.0 release Search is bundled along with the standard Riak package (you just need to set a flag to enable it). You can try using the latest pre-release or even build from source on master or the 1.0 branch. If you don't actually need full-text search then you could also checkout Riak's new support for secondary indexes. Please drop a line on the mailing list or IRC if you have questions.
CiuchDB is another DB I like quite a bit, but wrapping your had around writing map reduce views in order to index and execute queries is A LOT different than a select.
We had a 3 node system across a wan, where local clients would read from a local node, and if they needed to write, would write to the master across the wan. The issue we had was there was a significant amount of data being written to the master that one or two of the nodes weren't interested in 100% of the time, but it was still being replicated across the wan (and at cost)
so I investigated how the replication mechanism worked - to see if I could gain greater control over what data was replicated. As it turns out, there isn't, but you can emulate the replication yourself if you're interested. Well this is how we did it:
- master node has a capped collection (we call db.messages) - slave nodes have a bit of mongo console javascript that execute tailed cursor querying for messages being inserted - when a message that matches the query inputs is inserted into the messages collection, it gets inserted into a local db.files collection, which clients then read from
The added bonus is that occasionally we do need to replication additional data to particular nodes, so we just craft up a tailed cursor query that finds any messages we can, and pulls them across the wire
we make a fair bit of use of the adhoc querying in mongodb, so that was a massive selling point for us
speed was also a major issue - we have a average write/very high read requirement, and it's really really fast.
finally the platform was a consideration. We're mostly a windows shop, but we'll use linux where needed, and we were prepared to have riak running on linux if thats what we thought was bets, but it just didn't really fit for us.
the only downside was we use Delphi, and the existing Delphi drivers were.. not good, so we wrote our own, which I'm trying to negotiate with my boss so we can push onto github.
good luck!
"We're mostly a windows shop"
"we use Delphi"
I've worked with Delphi for many years, and its not far from the speed you can achieve with C++. Delphi is one hell of an underrated language for native programming.
Also, Delphi's standard library, the VCL, is great, and you can learn so much reading the source code, which comes with the IDE. Or does it? Delphi 7 was the last version I used.
Yes, Delphi, the language, which uses a static, strongly typed language, reference counting for memory management, and native ahead-of-time compilation can achieve good performance. That's a shocker. The problem is this isn't what it's intended for, and the standard kit is hardly optimized for it.
And the other point you are trying to make, I still don't get it. You imply that the guy is doing it wrong for using Delphi where speed is important, and you come here and say that Delphi is, indeed, a fast language.
What if speed is not what the language is intended for? If the language is intended for other purpose, and speed comes as a bonus, I don't know how that could be a bad thing.
We started out as a TP shop, so there is code dating back to the 1980s that is still used today (=\)
The rumor is Borland/CodeGear/Embarcadero lost the source code to the delphi compiler, which is the reason there had been no additions to the compiler, but who knows if that is legitimate, or should be an entry on snopes
dynamic arrays are reference counted
interfaces are reference counted (mostly for doing com stuff, however I've seen some examples of people trying to use interfaces as a cheap/dirty "i don't have to think about it" memory management tool, badly)
everything else is managed by yours truly
please enlighten us as to what you think Delphi is intended for, and what the standard kit is actually optimised for
Personally, I think Delphi is for cutters.
Of course if anyone has managed to hook into a Riak provider on Heroku I'd love to be proved wrong.
And yes, I could always spin up an EC2 instance and manage a Riak cluster myself, but my time is constrained and I'd rather focus on getting my app out the door than sysadmin tasks.