Show HN: SummitDB – In-Memory NoSQL DB
github.com
github.com
Apparently, in-memory DBs in now. They're fast and all, that's nice, but I'm still running on a machine with 8 Gigs of ram, and when the dataset starts to exceed that, I want an option other than "grab some more DDR3." I remember the fate of TinyMUD: on-disk operation must at least be an option for any DB I'd use for datasets of a non-fixed size.
I store Cap'n Proto objects directly into LMDB, and can read them without any memory allocation or deserialization cost. Freaking rocks.
I run an LMDB dataset many times the size of RAM, works great. Most data isn't recent, so the virtual memory cache works extremely well.
I am using it with Nodejs, on t2.micro, handling 10+ million requests / day
Never had an issue. in production for one year already. Really good :-)
That's LMDB, as others have noted. I have used it in production for years, never once lost data. Insanely fast, low footprint, crash-proof.
(P.s. I have seen SQLite get corrupted regularly on mobile devices, which is why I no longer use it there.)
It's possible my SQLite configuration wasn't as crash-proof as the default install. I moved onto another project before I could investigate further.
it's getting difficult to manage all my different databases.
Anyway, closest to to low-latency part would be NonStop or OpenVMS clusers across redundant, leased lines. They're fast, scale well at HW level, use ACID databases, and high-availability. Commonly used in transaction processing. One VMS cluster survived 9/11 attack without losing a single transaction. NonStop does up to five 9's. Id be interested in the CAP analysis of these older methods.
Quorum across some data centers is fast enough that it's not hard to manage in new applications.
It's similar to FoundationDB (the Apple product you're referring to) in that it's an SQL database layered on top of a K/V store. It's based on some clever technology to accomplish distributed transactions, strong consistency and high availability, and is looking very promising.
MemSQL for blazing fast distributed full SQL database with cross-datacenter replication and in-memory rowstore + disk-based columnstore.
ScyllaDB for Cassandra rewritten in C++ for blazing fast dynamo-style multi-master multi-datacenter disk-based wide-column database.
I'll also throw in AMPS (by 60East) as the best messaging platform that supports innovative SQL and real-time state-of-the-world queries on it's message streams (instead of using that rabbitmq or kafka crap).
Dynomite for production proven, limitless scale for Redis.
Provides ability to Redis for both in-memory and on-disk workloads.
Dynamo-inspired linearly scalable, shared nothing multi-master architecture. Pairs well with Cassandra.
Supports pluggable backends (ex. RocksDB).
Been in production 2+ years. One of the larger clusters handles over 3.6 million sustained ops/sec in production, every day.
The benefit of Dynomite's support for the Redis API + protocol is a.) you can use any Redis client and b.) the same code can be used for standalone Redis on your laptop and on a distributed Dynomite cluster.
Also you were a speaker at Data Layer right?
There are two answers to your first question.
1. Dynomite pairs well with Cassandra. At scale, Cassandra is not a speed deamon when it comes to reads. Dynomite helps to improve Cassandra's read performance.
2. The use case for Dynomite as the primary database is for workloads that require high throughput and low latency. In other words, Dynomite delivers consistently lower latency at any scale.
For open-source fast and clean messaging, I'd recommend NATS, they recently built NATS streaming that builds on top of the pub/sub to include kafka like persistence - http://nats.io/
I don't think nats can be used like I said by looking at the docs.
What are you looking to do? Sharding for what? Data size? Throughput?
You can have (open-source or free) and develop without community (i think like chrome does that).
The question is: is it worth it ? Without vc-money-cash it's a little hard AND time-consuming.
However, I wanted a feature that allowed custom indexes with Javascript, like CouchDB map functions.
The secondary index module will be more focused on automating more traditional indexing of numbers and simple string values. It has an SQL-ish language for WHERE expressions, and then the result of that is piped to any redis command you want it to.
If you're interested, here's a draft of the syntax (it's changed a bit since but you'll get the idea). https://gist.github.com/dvirsky/3ef73143a6d8212f2b50096a8eb6...
The recent additions would also allow more parallelization of the indexes for reads. I could create RW mutexes in the engine and allow multiple clients to work on the same term index. It won't be trivial but it's not super hard as well.
If not, what is the use case? Why would I need all those ACID-y guarantees if my server can fail at any time and all data is gone?
And by the way: a lot of existing databases support this.