ActorDB: A distributed SQL database with the scalability of a KV store
actordb.com
actordb.com
a) warrant this kind of complexity b) be able to give up ad-hoc queries
So I guess if I ever need to build a highly scalable, fault tolerant, distributed todo list.. I'll look into it. Not sure how I will aggregate data to run the dashboards that my MBAs will be up my butt about though. Guess we can cover that detail later?
b) ad-hoc queries tend to be the first thing out the window once you hit a certain size (query load or dataset size). ActorDB arguably gives you the largest amount of query flexibility out of the new breed of databases, while still being strongly consistent and persistent (not memory only).
If so, this might be the first one of these new databases that actually helps with my use case. (Though I'll wait to see what jepsen makes of it first).
It's a great idea, but I think the "mission statement" needs to be restated.
In this case it's that JOINs are still expensive. Well, this is actually a pretty big deal, because JOINs are typically the reason I settle for either a NoSQL or a SQL store.
Again, I'm sure this has other merits that may distinguish itself from the rest of everyone and their grandma building a database these days.
When I saw that in the README, it said to me that I can have a distributed SQL database as long as I did not need JOIN or REFERENCES which are commonly used features of a relational database. Am I incorrect?
This is the sacrifice ActorDB makes to be distributed.
1. A NoSQL database but with SQL as the API and some support for relations. This is quite nice because I have often found NoSQL APIs to be lacking in many aspects.
2. A set of fully relational databases that need to be managed and/or load balanced. This is a use case I have. If the system administration is significantly easier than the tools natively provided by Postgres, MySQL, etc, then I can see this being a valid use case. The use of SQLite does not concern me since it has legitimate downsides [1][2].
For instance if you were to create your own news.ycombinator.com you would require 3 types of actors: frontpage, thread, user
frontpage:
- CREATE TABLE threads (storyid INTEGER PRIMARY KEY, author INTEGER, title TEXT, url TEXT)
thread:
- CREATE TABLE comments (id INTEGER AUTO INCREMENT, parent INTEGER, user INTEGER, txt TEXT)
user:
- CREATE TABLE info (id INTEGER PRIMARY KEY, email TEXT, password TEXT)
- CREATE TABLE posts (id INTEGER PRIMARY KEY, threadid INTEGER)
There would be many thread and user actors, but only 1 frontpage actor.
>Use case: reliable distributed counters
I am not understanding the example provided for this use case. In my implementation of this use case I use Redis back by an RDBMS. Redis provides a cache and a distributed counter while the RDBMS engine handles all the OLTP and OLAP.I currently have a job that takes snapshots of the real time Redis counter at given intervals and inserts them into my analytics table.
How would the ActorDB way simplify or improve this model?
- no single point of failure (redis and likely the rdbms as well)
- no need to maintain two very different systems
- scalable. As in plug in a new cluster and it will proportionally increase capacity.
We are working on more detailed use case examples. As well as a pretty big bugfix release due out tomorrow or the day after.
>no single point of failure (redis and likely the rdbms as well)
Your README mentions replication in the Operational characteristics, but it does not cover partition tolerance?- Is there one master per cluster?
- What happens when the network becomes segmented?
It is also unclear of how this affects multiple clusters.
- How is a global consensus of the data reached?
- What happens when the network becomes segmented between clusters?
>no need to maintain two very different systems
SQLite isn't a full featured RDBMS. It mainly lacks support for stored procedures and concurrency. It is a very nice solution for small, self contained, horizontally scalable usage scenarios. However, once the use case involves generating custom globally unique identifiers (like a 7 character alphanumeric string) it generally falls apart. >scalable. As in plug in a new cluster and it will proportionally increase capacity.
Redis and RDBMS have well known use cases for replication and partition management. Redis itself is single threaded, so it is horizontally scalable by your budget (and likewise supports sharding and replication albeit the network segmentation can lead to issues). More full featured RDBMS like Postgres, MariaDB, Oracle, and SQL Server support concurrency and granular locking with similar scaling strategies.I do not see the value proposition in consolidating the infrastructure in this case.
All queries (reads and writes) go through the master and are not committed if master does not reach a majority of configured servers in the cluster.
Individual actors are not meant to be scalable. They are meant to be fast enough. For instance if you were to create your own news.ycombinator.com an actor would be this entire conversation tree. Sqlite would be sufficient. You don't need concurrency on a per thread level.
ActorDB provides a global uniqueid generation feature (independent of sqlite).
Using postgres or mariadb would be very problematic or even impossible, because they are not designed for this kind of use. ActorDB needs to be able to move actors to new clusters when they are added, while still execute queries on them at the same time.