HNHacker News
TopNewBestAskShowJobs

skjhn

102 karma · joined October 9, 2013

submissionscomments
skjhn··on Concurrency Behavior: MongoDB vs. Couchbase
MongoDB returns inconsistent query results. Couchbase avoids this problem with skip lists and snapshots.
skjhn··on MongoDB fails to perform, runs out of gas
Sort of. One of the requirements was replication. The databases had to replicate every document so that it was stored on two different nodes. This ensures availability. The only way to do that with MongoDB is to run secondary nodes too. Now, it's possible to do say 4 primary and 4 secondaries, but performance would go down because only 4 nodes could do the writing. The problem is that it rarely works out. It's fine at first, but down the road...
skjhn··on MongoDB fails to perform, runs out of gas
"Did they test against a single MongoDB server or cluster?"

Couchbase: 9 nodes on 9 AWS instances (1 node per instance). MongoDB: 18 nodes on 9 AWS instances (1 primary and 1 secondary per instance).

"Well if you pick a good shard key it should be able to go to a single node."

Not necessarily. From the docs: In some cases, when the shard key or a prefix of the shard key is a part of the query, the mongos can route the query to a subset of the shards. Otherwise, the mongos must direct the query to all shards that hold documents for that collection.

If you query your data in many ways, you have lots of queries not based on the shard key. Considers users. What do you shard on, email? What if you want to query by location, age, status, and/or preference?

skjhn··on MongoDB queries don’t always return all matching documents
At least you get SQL and JOINs with Couchbase.
skjhn··on High Performance Erlang – Finding Bottlenecks in a CouchDB Cluster
There are a handful of databases that implement map-reduce one way or another - CouchDB, Couchbase, and MongoDB off the top of my head. Views might be a CouchDB/Couchbase concept, but not incremental map-reduce.

In what way is JSON inefficient? Are we talking about size?

GSI indexes may or may not be partitioned. With GSI, depending on the index size and resources available, you would most likely NOT partition the index - that's the recommendation. You can create an index on user_id and column_b, place it on a specific node, and you'd only be hitting that node for a query. Especially if it's a covering index. Again, databases without GSI indexes have an index partition on every single node - that means hitting every single one for every single query. I'm still not sure what you're trying to get at.

I'm guessing you are referring to MongoDB shards and routers. However, that example doesn't make sense. If user_id is the shard key, then yes, the router sends the query to the right node. The same thing happens with Couchbase. Given the key, you get the document straight from the node that has it. However, if you have user_id, why are you querying on column_b too? Now, if user_id is not the shard key, then no, the router does not send the query to the right node, its sends the query to every single node.

I'm generalizing, but key-value databases are best for key-value operations on arbitrary data. Document databases understand JSON and, as such, can provide access via queries. With Couchbase, you can choose from views, N1QL (SQL), geospatial (built on views), or full-text search (preview). Pretty far off from a key-value store.

That, and it already has support for partial updates via N1QL. However, my assumption was that you were talking about partial updates via key-value operations.

skjhn··on High Performance Erlang – Finding Bottlenecks in a CouchDB Cluster
I'm not sure about the format, but it's compressed.
skjhn··on High Performance Erlang – Finding Bottlenecks in a CouchDB Cluster
Let's see if I can help here.

A lot people like async map-reduce. If you need to perform aggregation on a lot of data, its constantly growing, and you need the results to be current, async map-reduce is great. In the best case scenario, the results are precomputed. In a worst case scenario, they are a few seconds out of date. However, you have the option of forcing an update if need be. Either way, it's a hell of lot faster than running the full aggregation every time it's requested.

Redis is great, but a) the memcached protocol is well established and b) Redis is more than a simple cache.

BSON vs. JSON, what's the point here?

A query doesn't have to hit every index node. That doesn't make any sense. In fact, it's quite the opposite. With local indexes, you would, in fact, hit every single node. With global secondary indexes, you hit the index node with the right index.

Are you talking about partial updates? If so, yes, that will be available in the next developer preview. Stay tuned.

skjhn··on MongoDB Is Special, Its Benchmark Proves It
TLDR - MongoDB is faster than any other NoSQL database in a benchmark published without any configuration. The results can't be reproduced let alone validated. However, in a benchmark published with all of the configuration and all of the results...
skjhn··on MongoDB Rules Single Node Deployments, Fails to Scale
Good catch. Fixed.
skjhn··on MongoDB Rules Single Node Deployments, Fails to Scale
TLDR: MongoDB benchmarked Couchbase Server by having it perform CAS operations (read + update) while MongoDB performed basic operations (update), they used an outdated client library that is two years old, and they performed it with single node deployments.
skjhn··on How Wired Is MongoDB and WiredTiger?
It compares 9 nodes of Couchbase Server to 9 nodes of MongoDB. Couchbase Server reads and writes to every node, but MongoDB only reads and writes to primary nodes. That's the problem with active/passive topologies. MongoDB can read from secondary nodes but only if a) you don't require strong consistency or b) you don't mind degrading write performance - it would have to replicate to all secondaries nodes before the write finishes.
skjhn··on How Wired Is MongoDB and WiredTiger?
That may be true, but they both have document models and Couchbase Server will have a complete query language based on SQL, secondary indexes, and support for joins.

http://docs.couchbase.com/developer/n1ql-dp4/n1ql-intro.html http://query.pub.couchbase.com/tutorial/#1

skjhn··on Web, Mobile and IoT distributed database requirements
I think it's safe to say both mobile and IoT network access is anything but reliable. Not sure if is another solution besides an embedded database with automatic sync.
skjhn··on MongoDB and DataStax, In the Rearview Mirror
Network roundtrip included.

Regarding MongoDB, they are. They say MongoDB 2.8 will have document level locking.

No durable configs in this benchmark though. All databases fsync'd after writes. The servers had SSDs.

skjhn··on MongoDB and DataStax, In the Rearview Mirror
Because the data is written to a data structure before the write completed. You can "write" to a cache or to an in-memory database. Databases, relational or NoSQL, write to memory first whether it is to an application managed cache or an OS managed cache (i.e. page cache).
skjhn··on MongoDB and DataStax, In the Rearview Mirror
That's right, and the same is true for MongoDB and Cassandra. They do not fsync on writes and they replicate asynchronously.
skjhn··on MongoDB and DataStax, In the Rearview Mirror
Good point. After all, most interactive applications read and write data from local files via SCP and S3.
skjhn··on MongoDB and DataStax, In the Rearview Mirror
Do you place a maximum amount of data in that 10s window? If not, it too is arbitrary. I do think batching concurrent writes is nice.
skjhn··on MongoDB and DataStax, In the Rearview Mirror
As can Cassandra. It depends on how much data is still in the page cache. I do agree both systems can be configured for immediate durability. However, we went with the default values as most people (and most databases) do not sync on every write. It is too much of a performance cost.
skjhn··on MongoDB and DataStax, In the Rearview Mirror
Most of the data is read from memory, not all of it. All of it is on disk too.
skjhn··on MongoDB and DataStax, In the Rearview Mirror
Performance is a reflection of architecture. Better performance, better architecture. That, and you can never downplay performance.
skjhn··on MongoDB and DataStax, In the Rearview Mirror
Well, writes are not durable until fsync. That's true for MongoDB, Cassandra and Couchbase Server. That being said, Cassandra demonstrated great write latency. The issue was read latency.
skjhn··on MongoDB and DataStax, In the Rearview Mirror
It can and it is, but it might have to be scaled out.

http://www.couchbase.com/liveperson http://www.couchbase.com/paypal

skjhn··on MongoDB and DataStax, In the Rearview Mirror
Good catch. Latency is in milliseconds. I'll have to fix that.
skjhn··on Viber Replaces MongoDB with Couchbase
In what ways do you consider them to be different from the point of view of an architect (other than the different clustering implementations)?

They're both document databases. They're both distributed. They're both supposed to scale.