Replacing Redis with BoltDB – A Pure Go Key/Value Store
specto.io
specto.io
There seems to be a lot of focus on the performance improvements by commenters. Using a local mmapped KV can remove a lot of network and serialization overhead, however, I would caution others from thinking Bolt is a silver bullet. There are a TON of use cases where Redis would utterly destroy Bolt (e.g. high random write volume).
The two data stores have fundamentally different designs and purposes. For example, Bolt provides MVCC with fully serializable transaction isolation. If you need that then Redis isn't an option. If you don't need that then Bolt may be overkill.
As always, YMMV and be sure to test out different options based on your requirements.
> The two data stores have fundamentally different designs and purposes
With all the datastore hooplah nowadays, it is really nice to see an author come out and say "this isn't for everyone" rather than "we're the fastest for everyone, period." and throwing around a bunch of meaningless benchmarks.
If Bolt supports MVCC using a page log, why does Redis destroy Bolt on random write volume?
Also, I'm seeing wildly different write performance measures. Siddontang on #237 back in 2014 for Ledis suggests 970 Sets per second. This HN article mentions 800 requests per second. Running "bolt bench --count 10000000 --batch-size 10000 --write-mode rnd" gets me 29203 ops per second. Can you shed some light on what random write performance I can expect?
Instead I just use the language's native array and map types to store objects, and write simple functions to persist the data into a transaction log (usually a "newline-separated JSON objects" style file). On application startup, the objects are reconstructed by reading back the log file.
This approach can go surprisingly far. Memory is cheap, and not having a separate database behind a socket eliminates a whole class of potential issues and bottlenecks. The log file is easy to back up. If it grows too big, you can always compact it by dumping the in-memory structure.
(I guess I should add a disclaimer about "toy projects" and "serious production environments", because some people seem to get oddly upset about the notion of not having a "real" database for a web app...)
When it's my code and the persistence layer fits in a few screenfuls in a text editor, at least I know where to place the blame. Debugging database trouble is maybe the worst kind of work I can imagine.
Rolling your own on disk persistence and recovery is like rolling your own crypto. What you don't know can definitely hurt you, you'd better be an expert, and even then you won't likely get it right the first time around. It's just a really hard problem involving complex interactions between your code, the OS, the filesystem, the disk drivers, and your disks all of which can cause data loss in very unexpected ways.
Your method can work, but it almost certainly has issues lurking that can cause data loss. You might be ok with that because it's a toy project, but it's obviously not a good idea for important data.
Here is an MMO in kdb+/q that fits on a single Vim screen: https://github.com/srpeck/kchess
This really helps with testing, because now you have all your initial dataset available without needing to write it after the fact, and the database connection wasn't always an assumed fact.
When people throw everything into a database, in my experience the specific requirements for specific subsets of the data are often lost in the fog of history as people often don't document why they need to be in the database.
Then it turns out the database is used for everything from a short term cache which could just as well have been in memory, to logs which are of limited utility over long term and "just" need to be indexed for short term querying, to vital financial transactional data which needs to be kept for X number of years, to per-user data which is never, ever looked up other than when that specific user is logged in, and so would be ideal to shard, and so on.
Grabbing for the RDBMS right away is then the lazy way out where you can often get away with not seriously evaluating what is best, because it is "good enough". Then it's nice to start out with something that is sufficiently limited that you're forced to take it into consideration sooner rather than later.
- Redis has no dependencies either
- "weird responses" from Redis? Do you change your tools the minute they don't behave the way you expect, without investigation?
- The title "Picking the right tool for the job" is the almost opposite of what's done here. It reads more like "Stumbling on cool Go software and picking what job it can do for me"
- 400 (or 850) requests/second are ridiculously low numbers for an in-memory key/value store. Redis is capable of doing 100x that on small/medium machines.
SET: 109051.26 requests per second GET: 109170.30 requests per second
with 32 byte dataset.
Also, using Redis as a key-value store for its persistence isn't really a great idea.
Could you please elaborate on this point? I'm thinking of using Redis precisely in that manner and would love know about the drawbacks.
So when you server dies you lose the changes since the last flush. I would not use as a primary storage for data you cannot afford to lose.
Have a look at this: http://redis.io/topics/persistence
Edit: you can configure redis to flush on every command that changes data...but you probably wouldn't want to use redis that way :)
Edit: Also, you probably have bigger problems if your database goes down or your cluster falls off the map somehow.
Very few apps can make do with only a key-value store and if you have to throw a relational db in the mix, why introduce more moving parts unless absolutely necessary? ACID, referential integrity, SQL. Sometimes it feels people are willing to ditch all this for very little gain.
Let's face it, how many apps out there can't really run a beefy relational DB with potentially an absurd amount of RAM?
The bug in question: https://github.com/boltdb/bolt/pull/452
Serious question: BoltDB something like a Key Value store equivalent (congruent?) to SQLite?
The redis go client had probably some design issue.
One of the likely causes for issues is file descriptor limits or TCP tuning that you suffer when stress testing two dependent networking apps on the same server. BoltDB, being embedded, just works out of the box.
On another note, really need to find time/excuse to play with go... It's been on my radar for quite a while now... but the use of JS all around has been very nice for my productivity.
Not exactly surprising. Redis is single threaded so if you will need to spawn multiple instances if you want to satisfy all cores. Which is what I assume is happening here.
EDIT: Anyway for my programming philosophy, what the OP did, regardless of performances, totally makes sense. I think easy to use and deploy is a key value in software ;-)
You should be able to pipeline the recording, but likely not the get operations.
What I take from his post is that his software was nicely designed to be able to easily change the backend which will allow him to easily switch back to Redis in the future when he will get corruption from BoltDB because of the intensive read/write which is not the use case.
If any data page is partially written then it doesn't matter because the new meta page hasn't been written to point to it. If the new meta page is partially written then it is detected through a checksum and the previous meta page is used (thereby rolling back the transaction). That's how Bolt works as well.
That being said, all code has bugs (yes, even LMDB). Bolt has a large amount of test coverage as well as randomized black box testing that it uses to help minimize those bugs.
Are you guys still single process? That killed a few pretty important lmdb use cases. I've always been curious why you did that.