I'm in an interesting spot myself right now. I've been hacking on Django since 0.96. I love the ORM, and hate everything else. Python on the other hand is a wonderful language. Rails... I've been spending 8 hours a day for the past month building a rails project. I love a lot of things about rails, hate the ORM, and dislike a lot of the magic.
On the flip side: I've been a Javascript/front-end developer since before jQuery existed, and I love coffeescript. I understand callback-style evented programming and therefore node.js apps make perfect sense to me.
My co-founder and I have just finished some heavy duty design/wireframe/conceptual work on a new project of ours. It's time to build the API.
I am COMPLETELY torn with what framework/language/database to use for this project.
At first I thought mongodb and coffeescript+node would be fun and reliable. But the community seems to think mongo is unreliable and almost dangerous. So I thought, hmm, Riak looks like it will let me throw whatever data I want into it (like mongo) but it's rock solid. I later dismissed that since I figured I'd need to do a lot of heavy lifting on my own. Rails keeps poking me in the back of my mind but I honest-to-god hate ActiveRecord and am afraid to use DataMapper for fear it won't play nice with a lot of the popular ruby goodies out there. I've modeled most of the project now in Django and i'm starting to play with Tastypie for the REST component. It feels too kludgy.
I'm a wreck. Mommy?
Thankfully with mongodb-2.2 there is now instead a database lock which mitigates the issue.
Personally I feel redis more closely resembles language primitives with simple key-value pairs, hashes, lists (ordered, repeating values allowed), sets (un-ordered, no repeating values), and sorted sets (more interesting).
All of the options you have presented are "Very Good" - they do so many things right. To steal a pithy quote from my co-worker: "Its a balloon squeeze."
This decision will not determine the success or failure of your business, move along.
I pick Rails & Ruby.
It provides an ActiveModel-compliant interface, so it works fine with Rails. It also has a lot less magic (associations don't use proxy objects, for example) and has a lot of features that ActiveRecord is sorely missing (such as complete support for composite primary keys). I don't know what your other problems with ActiveRecord are, but Sequel is definitely worth a look. I'm very happy with it.
Use this to test and play with and empirically figure out what the best persistence layer is for your app. There's no possible way you know at the beginning of the app what persistence is best for your app, so don't make that decision yet.
I like the fact I can whip something up quickly with MongoDB, get everything working and then once my schemas and services operating thereupon have stabilised then I can start looking at swapping my implementations out for a SQL-based back-end if I think it's worthwhile.
If you like Python and want to use a SQL backend, SQLAlchemy is a good place to start.
Here is my first flask app ever, https://github.com/whalesalad/arbesko-files
No authentication or session management or anything above basic request/response, but it was fun and serves it's purpose inside of Arbesko well.
So the answer to your question is simple. Use what matches your data model. MySQL if your data is more relational as it is proven and well known. MongoDB is perfect when you treat it like a document store. For example I fetch one User object (with all the user's data self contained) from the database and do everything client side. So in my case MongoDB works perfect.
But either way. Do whatever gets you up to speed the fastest.
Not as a criticism of Mongo but the thing about very large companies is that they can devote a team to just making sure some bit of technology works. Almost anything is suitable for a large production environment if you can afford to have a couple people elbows deep in it tuning and troubleshooting.
The idea that a database was shipping with a mode where it would return after a write without waiting for the result to be written to disk. Or a database where silent data corruption is an easy possibility and yet still called itself a database is what is "crazy"
And silent data corruption is an easy possibility ?
I would recommend against DataMapper. It is nicely constructed, but has some weird issues. Also, DataMapper 1 is going to be replaced by DM 2, which is a completely different beast (but better, hopefully).
Don't be fooled by ease of setup of the DB too much. It should be rather straight-forward, but even if a DB can just be downloaded and ran to dabble with it, always remember that this just shifts the moment at which you really have to dig into the details. You don't want to do that 2 days before going to production. Be on the lookout for red flags, though: unreasonable and weird configurations that have to be flipped for no understandable reasons at all before going into production. That shows that the project is not maintained very well or the developers have lost track. Also try to figure out how long it takes you to solve a new problem with the database, with documentation reading and all. Pick the one that you can wrap your head around best.
That's why everyone chooses MySQL as a relational database. It is easy to setup, easy to maintain and EVERY problem has been solved and searchable online.
The underlying issue is that (except for the most trivial cases), investing time in properly operating and using your database is a huge gain that many young companies ignore. 6 Month later, they have a burning datastore at hand that they don't know how to fix.
There are huge benefits that come from being the most widely deployed database.
apt-get install mysql-server redis riak-server mongo-10gen-server
It's actually pretty rad to have every single location I've ever checked-into via Facebook stored in this database. I can say "find all the spots 2 miles from my house" and it's shockingly quick.
I have PostgreSQL on my mac. It's easy to install and use.
PostGIS was the killer. Version 2.0 (homebrew installs postgis2) does not play nice with Django locally (on my mac) and the ticket has been open for a year or something to fix it. That's a huge red flag for me right there. https://code.djangoproject.com/ticket/16455
I guess I could install 1.5 manually.
My problem with trusting Ruby/Rails is that I've only been seriously developing in it for a month.
Lets put it like that: feel free to try some new stuff, but don't start getting all crazy with it. So, if you tried Rails and weren't that convinced by it, maybe default to something that you know, even if it isn't a love relationship. Learning is a great thing, but do it piece by piece, especially if your goal is getting work done.
I've been testing ElasticSearch's geo location search and most queries take 50-150ms.
Standard term searches and filters will take like 5ms, with id GETs taking mere fractions of that still.
Screenshot: http://wsld.me/J5zu
If I hit it repeatedly for a while it comes back in as quick as 0.14, but usually not that fast.
ping to server:
michael at Achilles in ~
○ ping -c 5 apollo
PING apollo (192.168.1.130): 56 data bytes
64 bytes from 192.168.1.130: icmp_seq=0 ttl=64 time=9.497 ms
64 bytes from 192.168.1.130: icmp_seq=1 ttl=64 time=83.113 ms
64 bytes from 192.168.1.130: icmp_seq=2 ttl=64 time=7.335 ms
64 bytes from 192.168.1.130: icmp_seq=3 ttl=64 time=6.998 ms
64 bytes from 192.168.1.130: icmp_seq=4 ttl=64 time=6.393 ms
--- apollo ping statistics ---
5 packets transmitted, 5 packets received, 0.0% packet loss
round-trip min/avg/max/stddev = 6.393/22.667/83.113/30.241 msAs for why node is a bad choice: you can think of 2 categories of languages:
1) Those which allow you to get things done quickly but in which large code bases are hard to maintain because you need to carry too much information in your head, and they possibly don't scale very well because they tend to run in an interpreter (python, ruby, lisp).
2) Those which require some amount of boilerplate to get simple things done, but provide structure and modularity so it's easier to maintain large code bases, and they possibly produce high-performance code because they tend to be compiled or almost-compiled (C#, Java, Go).
There's a trade off between "getting up to speed quickly" and "maintainable performant code".
Node is bad because it fails at getting you up to speed quickly AND it doesn't make your code maintainable at large scale. This is because async-style programming is not natural to how we think.
The "performance" should not be taken as a plus for Node because you can get it in other languages/platforms (Go, for example) without sacrificing maintainability.
The issues with carrying around information in your head is true IFF you have a large monolithic app. SOA, which is advocated with compiled languages as well, helps to avoid this issue because each service is limited in scope. Then again, SOA itself is counter to "hello world" latency...
I think this is important because if you find yourself dealing with a (example) rails app that is experiencing cognitive burden overload issues, splitting it out into a SOA can be a nice transition (eventually replacing parts with other languages if it is required.)
Also, if you care about performance and quick execution and want a scripting environment, luvit is much faster than node, much lower memory usage and IIRC you can use lua continuations instead of async style.
IMHO, node.js as an edge server providing connection handling and templating with backend services giving it JSON seems like a reasonable 3-tier architecture, especially because your edge tier can then be maintained by people with front-end experience.
IE: Encapsulation and isolation, which is something that "scripting" languages generally stink at enforcing (on purpose!)
I have yet to try it, but this seems like a great approach. Manage most of the computing intensive tasks in whatever environment you like/need. Use Node to talk to APIs and render templates. With a well developed client framework you can even make only the first render in Node and the following in the client talking to the same APIs.
This leaves the most complicated back-end problems in the hands of developers that like and understand those problems and creates a friendly environment for the front-end developer who usually focuses on building features. The front-end developer even has already some experience with async code because of events in the browser, so he'll probably get up to speed quickly with Node.
As for other features, I like the REST-ful interface, a continuous changes feed, master to master replication, and a usable interface to see and manage data -- Futon. I haven't yet found a product that comes with all those features.
I've seen this situation quite a bit with traditional SQL databases: somehow you end up with a giant complicated query, and it's becoming a bottle-neck and it's having a terrible impact on performance. Somehow you need to untangle the mess and figure out where the problem is. Sometimes you find out the query is just poorly written. Other times the problem is fixed by simply adding indexes to some table.
This is a terrible situation I don't want to run into.
I have a feeling MongoDB might also suffer from a smilar problem, because it doesn't have cached map-reduce views; it has a query language.
In CouchDB, views are pretty much indexes, and they're pretty much the only way to do a query. I'd argue that because it's explicit like that, it's easy to reason about.
The basic idea is that the intermediate Map function results are persisted, while most other system (read: hadoop) throws them out.
If you're interested in some of the MapReduce history, where it's going, incremental indexing, etc., then you should read this: http://gigaom.com/cloud/why-the-days-are-numbered-for-hadoop...
Cheers.
Cloudant is close, but their BigCouch stuff still isn't merged yet as far as I know, so if you go with them you're still using some weird Couch fork instead of the open source version. Which might be more of a psychological problem than an actual one, but it makes me really uneasy about the Couch ecosystem they're the only thing we have that's close to a gold-standard champion.
I searched Reddit for "couchdb", and there's pretty much no new entry since 2 years ago. I almost thought couch was dead.
And what's the deal with CouchDB vs. CouchBase?
As for CouchDB vs. Couchbase, Couchbase is a fork of CouchDB headed by CouchDB's creators Damien Katz and J. Chris Anderson. Couchbase basically throws a lot of the nice things of CouchDB out the window (such as the HTTP API) in order to integrate Membase feature.
This is worth reading: http://www.couchbase.com/couchdb
As for the leaders, it's a smattering of CouchOne/Couch.io and Membase/Northscale people. Or, as a friend put it, CouchOneBase.io.
Cheers.
The mobile stuff that you're referring to is CouchBase's. It's a completely different product and their work should not be confused with Apache CouchDB's, though it often is.
But discounting all interpreted languages because they are interpreted is crazy biased. I don't at all accept (and it is not self-evident) that a high-boilerplate compiled language is necessary to have a large codebase or acceptable performance. Unless most of your storage is in memcached, the database is far more likely to be a bottleneck than your interpreter (!)
There are a number of compilers for Lisp.
Between these two issues with your post, I would not be convinced to accept your advice regarding node, either.
My point is about comparing dynamic and non-dynamic languages. Dynamic languages tend to have "magic" where complex logic occurs without you explicitly declaring it, at the expense of either a) maintainability or b) performance or c) both
There's a trade off here: varying degrees of magic vs varying degrees of boilerplate.
Boilerplate code can make things very explicit and thus easier to reason about.
Now, NodeJS has the worst of both worlds. The callback mess is not easy to write, not easy to reason about, not easy to maintain, and there's no magic.
That's a lot of sacrifice, it better be for good gain. But exactly do you gain? I don't see any gains. Performance? You can get that with other languages/platforms (like Go for example) without sacrificing code clarity.
tl;dr: Reflect upon your assumptions before making a decision more complicated than it might need to be.
Mongo is an update-in-place structure store with powerful atomic operations. Mongo will not grow storage if you are just changing the value of existing fields.
Couch is a compare-and-swap document snapshot store with multi-master replication. In replicated mode it's eventually-consistent with manual conflict resolution.
Riak is an indexable, eventually consistent key-value store with tuneable tradeoffs for speed versus safety, and no central point of failure.
Cassandra is a horizontally scalable, eventually consistent key-value column hybrid designed for huge data.
And so on... they are not "databases", they are not something where you replace one with the other, they are specific tools for specific tasks, usually much more so than unsexy but adaptable SQL.
Another, real world, object database is gemstone -- and can be used with smalltalk or ruby (among other things)[2].
[1] See eg:
http://www.ibm.com/developerworks/aix/library/au-zodb/
and/or:
http://www.zodb.org/zodbbook/installing.html
and
Besides giving you only the tools you want, it works very well with a driver like PyMongo. If you're starting out small you can also get free developer hosting with MongoHQ to get you up and running quickly.
Oh and you get to keep doing things in Python which is very nice. I've also had no problems with MongoDB for my web needs and I've worked at a few places that use it without the scary data corruption stories told to developers in order to make them keep drinking their Java.