JunoDB: PayPal’s Key-Value Store Goes Open-Source
medium.com
medium.com
It would be nice to see some benchmarks or just a mention of any kind of number. TiKV is a CNCF donated project with roughly the same architecture and has been deployed in larger clusters than 200 nodes.
I'd be very interested in learning why JunoDB isn't used as a SoR (or why PayPal doesn't consider it suitable for SoR).
When to use it? You must be sure you have a case where the additional guarantees relationship databases provide are not necessary, and you either need the simplicity of usage or deployment. Or in rare cases, speed, but whereas most people seem to act as if this the main reason, I consider it a relatively poor reason. A relational DB with its guarantees turned off (i.e., no relations, no transactions, tables with just a key and a value) perform fairly closely to a key/value store for a wide range of scaling needs. There is a top end where this matters, but fewer programmers need this than think they need this. Still, it is a valid concern, and if you don't need relational guarantees it may let you scale down the instance size.
Most of the cases I see where key/value stores are a good choice relate more to the simplicity than performance. I love me my Postgres, but nothing compares to the simplicity of just tossing up a Memcache somewhere and solving my problem, if that's what my problem calls for. No schema, no migrations, an API so simple my local programming language may well simply integrate it with my native associative array syntax... if you are careful to use it only where you aren't going to need relational functionality the bang for the buck can be very nice.
1) The system that is reading the data from Redis is a front-end system. I have concerns with security. As a front-end system, it can be accessed from the Internet and potentially hacked. As such, I make sure this system does not have access to the DB.
2) Preserve DB connections, size (affects costs of storing backups and time to hydrate), CPU, and memory. I would really like to keep my DB as trim and spry as possible. I am using traditional relational DB and not distributed. So in an effort to keep my complexity in DB management down, I offload some work to Redis. I could have created a second relational DB that has low security requirements and no relational object mapping, but I didnt think of that till now (I will explore this further in the next few days). In my case the data that is stored in Redis is accessed a lot and consists of data from multiple tables, so it was a good pick for moving to Redis. Performance wise, I suspect relational DB could perform jus as fast, but again I want to offload traffic off of the DB to not have to grow it or go distributed. If I was a better DB admin, I could probably created views or other relational DB features. But I felt it was easy to just let my API backend code (which has access to data that is also not in DB or might need to be formatted or calculated) construct the final data object then store a copy in Redis, whenever the data is updated.
I guess I'm like you - I really don't understand why someone who had another choice would opt in to a KV store (barring something like memcache, or something really high performance using an embedded KV store).
That was at least a chief use case spotlighted in the original Dynamo paper by Amazon that what the precursor to AWS’ DynamoDB paper.
Not to say that couldn’t be done with Postgres but of course they were dealing with insane scale on Amazon Day.
Using a key-value store for shopping carts can work for awhile, especially for the use-case you describe, but fails when system functionality grows beyond retrieving only by a cart ID.
And when using a persistent store which does not provide ACID capabilities, the system will ultimately have to enforce at least atomicity and consistency via server logic.
Either you manage your system so it never ever uses any other key than a cart ID (services have been running for decades keeping the same unique ID, that's not some unreasonable thing).
Or you migrate your data to match the completely wild new requirement, and taking costly steps to deal with a funfamental business change would be seen as reasonable in most orgs.
not that you're saying they don't, but some people might interpret your comment that way.
It depends on how DynamoDB is used[0]:
Transactional operations provide atomicity, consistency,
isolation, and durability (ACID) guarantees only within
the region where the write is made originally.
Transactions are not supported across regions in global
tables.
Granted, this likely handles most use-cases and the restriction enforced makes complete sense.0 - https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
The ideal use-case is when there is one, and only one, property used for retrieval which is guaranteed unique by the system. Less ideal, but often very performant, is when retrieval always uses one property which may not be unique.
Once retrieval requires anything other than a single predefined property, querying key-value persistent stores degrade into linear searches.
In essence everything is keyval. A sparse array is an ordered keyval store with integer keys (also technically everything is ordered, too, but some orders are stable, and useful, while others aren't). A dense array is an adjacency-optimized version of a sparse array where the key is implicit based on computable offset within a larger dense array, your address space. RAM address space is also variations on that theme. Raw disk storage. And file systems. Everything is. Maybe I spend too much time messing with storage, but I can't see it any other way at this point.
It'd be pretty silly to use a relational database for something that trivially shards across servers, often doesn't have any consistency requirements, and only needs 3 columns (request, response, expiry time).
I don't think it's even possible to use postgres as a traditional cache. Eg this trigger looks extremely slow and I would be unsurprised if a human holding ctrl+shift+r could single-handedly DoS a cache like this: https://stackoverflow.com/questions/26046816/is-there-a-way-...
> I don't think it's even possible to use postgres as a traditional cache.
The link I posted shows that you can use Postgres as a traditional cache. Specifically, the SO link showing that Postgres cannot do caching (expiring old data) is explicitly addressed in the link I posted.
One would have to got out of their way to miss the point. And miss it entirely.
You can put a fucking text files on the disk as cache and `find -atime X -delete` to clear it, doesn't mean you should use it.
> the SO link showing that Postgres cannot do caching (expiring old data) is explicitly addressed
Nope, UNLOGGED tables don't expire old data. They just fill up forever unless you use a trigger like the one in the answer I linked.
I also don't see anyone claiming they can saturate their network interface with postgres the same way they can with any KV store.
I recently tested it, and I could trivially saturate a 10GBit interface with postgres.
Total conjecture on my part.
If you want to have tier based storage, where you trade off latency for increased data size, meaning disk space is your limit and not your RAM.
edit: Redis supports fsync at every query in AOF mode
https://redis.io/docs/management/persistence/#ok-so-what-sho...
Personally, whenever I see a system like this I try to look for why they didn't go with an existing solution. And in 99% of cases either an existing solution never existed or they had a unique requirement that necessitated building something from scratch.
PostgreSQL is great for single instances but is poor when it comes to high availability and horizontal scalability.
Be curious what use cases where engineers are writing their own database but it is single instance.
There is plenty of absolute code abominations powering "world's largest applications"
Engineer competence is also vaguely related to quality of the infrastructure, yes, you need smart people to make big complex things, but you also need smart people to manage ungodly legacy enterprise spaghetti
Redis in comparison is a different thing if I'm not wrong.
Also, Redis was not originally distributed.
The question is definitely interesting. And to be fair, Juno was originally in memory.
That distinction is massively important when you're talking about distributed systems as key-value has far less edge cases to consider. Also its architecture is quite different as it doesn't have the concept of proxies.
Basically in the realm of databases the two are nothing alike.
I had a friend that worked for ISIS: Innovative Solutions In Space. They had a lot of problems with all sorts of financial institutions once that other ISIS started becoming better known. They've since rebranded to ISISPACE for that reason.
> JunoDB storage server instances accept operation requests from proxy and store data in memory or persistent storage using RocksDB. Each storage server instance is responsible for a set of shards, ensuring smooth and efficient data storage and management.
> JunoDB is PayPal's home-grown secure, consistent and highly available key-value store providing low, single digit millisecond, latency at any scale.
what do they mean by 'consistent' here?
While it has persistence options, they’re for durability and backup, not to increase the storage available.
JunoDB appears to store data primarily on disk, and limits storage by disk size, but then caches in memory as necessary perhaps. Quite different in behaviour and trade offs to Redis.