5 predictions on the future of databases (from a guy who knows databases)
gigaom.com
gigaom.com
That's certainly one way to interpret Facebook's dilemma (I know a few DBAs, and this statement certainly isn't false). Another way is that nothing else scales as well as MySQL, or nothing else offers the same tools to manage them at scale.
Relational data stores aren't going anywhere; there are still quite a few use cases they fulfill that NoSQL, NewSQL, and Graph DBs do not. This just triggers the "use the right tool for the job" reflex in me.
Can someone explain this to me ? In terms of acid principles wouldn't you want your transactions to be durable.
>Durability means that once a transaction has been committed, it will remain so, even in the event of power loss, crashes, or errors. In a relational database, for instance, once a group of SQL statements execute, the results need to be stored permanently (even if the database crashes immediately thereafter). To defend against power loss, transactions (or their effects) must be recorded in a non-volatile memory.
First, it's worth mentioning that you use in-memory databases when you have performance requirements that far exceed physical capabilities of hard disks or even SSDs. The workload will have a high degree of contention, high concurrency, etc.
Second, you're probably working with non-human-generated data. If you're not a bank that needs to guarantee a debit and credit went through (and that's slow human-generated data anyway), then you're looking at in-memory technology because you can guarantee every read will hit memory.
For writes, you basically have to use more machines to guarantee durability.
Any in-memory database worth its salt will write a transaction log to disk, so in the event of a power failure, it will read back from disk into memory.
You can increase the probability of 0 data loss, even in high contention and concurrent workloads, if you run the dataset across a set of machines, which both multiplies the number of disks capable of writing a sequential log as well as storing an extra copy in memory for high availability.
For the 99.999% of people who will never need to worry about Google or Facebook's "scale", stored procedures are absolutely the only sane way to build large systems.
I never understood this argument. Maybe you can elaborate a bit? Almost all our apps use sp for data lifting work (most of the so called "business logic"), and we never got scaling problems. Well, we don't have that much users or data either, but how many users/data do you have that your stored procs could not keep up with? What numbers are we talking about?
What makes the db slow is throwing bad sql at it, not Stored Procs.
> What makes the db slow is throwing bad sql at it, not Stored Procs
Absolute statements are absolutely wrong. :)
Particularly with DBs like PostgreSQL where you can (and people do) run Python scripts as stored procs. Great way to slow down your DB.
That is not much info :)
Just for reference: a simple query by pk index on my veryveryvery low end pc¹ on a table with 1mil rows takes about 0.04 ms (so says pgsql analyze). That is 25000 queries per second. What type of app are we talking here where the db needs to scale? Something like reddit? Stackoverflow? Just curious. Never saw such a beast where a DB server needs that much power. Where I work we have 200 apps on the same DB server, and it's idling most of the time.
¹i5 CPU 2.40GHz with 2GB of RAM
> run Python scripts as stored procs. Great way to slow down your DB.
There is more than one way to shoot yourself in the foot :)
Stored procs are used to allow rich transactions without using client-side transaction control. For all of the concerns about stored procs (of which some are valid), the fundamental issue is that external transaction control is performance's nemesis. It's not a coincidence that few No/NewSQL systems support external transaction control, and those that do don't publish many benchmarks using that feature.
Again, I think VoltDB is going places. I haven't been able to actually dig into it personally due to time constraints, and when I realized Command Logging was only available in the Enterprise Edition that effectively killed a few use cases for me.
One of the uses I hope to investigate is to use it as a nice materialized-view playground for bridging data on-demand from various databases, in coordination with Presto:
https://www.facebook.com/notes/facebook-engineering/presto-i...
In any case, it's a shame it's rarely discussed on HN.