Predictions on the future of databases
gigaom.com
gigaom.com
https://en.wikipedia.org/wiki/The_Third_Manifesto
The difficult problems since that time turned out to be scalability and parallelism, neither of which were addressed by The Third Manifesto. So instead of D, we got Map Reduce and NoSQL.
Who knows what the hard problems of the next 15 years will be.
http://www.dcs.warwick.ac.uk/~hugh/TTM/Tutorial%20D-2013-05-...
Which provides a description of Tutorial D, an implementation of the prescriptions.
Or "seem" -- I have really no idea how that happened and didn't notice it until too late to edit.
D would overcome the object-relational mismatch by making the relational side much, much more powerful. It's really about unifying the ideas of "tables of data are relations" and "relations are logical propositions", so that one could have an über-powerful general purpose language that just happens to also be a natural interface for databases.
D is essentially a set of low level prescriptions that the authors felt would be required in order to achieve such a language.
A RDBMS is a very slow key value store, so anything that can't do transactions or can't do anything else, is always going to be faster at its extremely limited set of abilities.
From an engineering perspective if you need 10 HP to run your 5 KW generator, a RDBMS is like installing a 10000 HP marine diesel and then not having any idea how to start or maintain it or even where to get the fuel. Obviously every objective performance metric would be better for a 10 HP lawnmower engine in that app, if all you'll ever plan to use is 10 HP and you have no idea how to use anything more advanced anyway.
There are also very loudly trumpeted anecdotal situations or contrived thought experiments where certain unusual technologies fit a unusual situation very well. This is, oddly enough, unusual.
I've read and experimented with the "7 DBs in 7 weeks" book and it IS very interesting but I can't find any business cases to actually use any of it, which is somewhat frustrating. And my experience is how you end up with people writing CRUD apps to store cooking recipes that none the less use NEO4J because they really, really, want to add a line to their resume that they used NEO4J, not because the app needed it.
"I don't know what I'm doing, but someone who knows what they're doing anecdotally solved a completely different problem using tool XYZ, so for lack of any better idea, lets copy them".
If you're familiar with cargo cult science there is an enormous miasma of cargo cult engineering fogging up the entire database arena not just nosql.
Another concept that needs to be in the discussion is the "no silver bullet" rule from programming applies to database design, like it or not. Can't just sprinkle magic nosql pixie dust on any old random problem and expect it to work, any more than applying any random programing fad to any random problem will work.
The (old and new) tools are actually pretty interesting, although often poorly engineered (by the end user) and implemented. Its the persistent anti-patterns and non-engineering design technique that I'm properly arrogant and condescending toward.
Hammers are a cool new invention and have some great unusual new applications, but they don't install deck screws any better than the legacy screwdriver. Laborers on the job randomly mixing screws nails screwdrivers and hammers on the job and then internet discussions about how hammers and/or screwdrivers suck is nearly physically painful to watch.
Maybe, but it's awkward. Graph or tree structures are still painful to store in relational databases, and a lot of problems turn out to involve those shapes.
> We have a working, declarative query language that Just Works(TM), for which we have written very good optimizers.
Maybe, but the tooling is still terrible. Where are the libraries? Where are the IDEs? Where's the integration that makes it easy to call procedures in my application language from SQL and vice versa? General-purpose programming languages have got better and better in the last 25 years, while SQL has stayed static.
> So, to sum it up: Exactly why should we abandon this?
Look at the ORM problem. Why do people continue to use these horrific bloated, leaky tools? Because it's really nice to be able to store and query in the native language of your application. We need storage tech that's better at supporting this.
On a large scale I think the future will be interesting, will "real databases" remain in IT land or will it move out in the world?
Don't laugh, I'm just barely old enough that when I started out, printing, file management, physical media management, static network address configuration (not just IP in the old days), and backups were done solely by IT professionals in the machine room, and now for better or worse every noob does it (or tries to do it) themselves. If in the future someone comes up with software to automate away DBAs...
Or it might go the other way and the corporate business database of the future will be super turbo hyper i-Excel.
It doesn't meet every single need you might ever have, but it sure meets a whole lot more then any other database I've seen. http://www.datomic.com
I've asked about this on their Google Group: https://groups.google.com/forum/#!topic/datomic/zkMV50VKw2s
As of right now, there are still questions left unanswered.
I'm not saying they made a bad choice, I'm just impressed at how they've managed to make MySQL fit their needs.
Impressive indeed.
ie. no joins - the part of the database that makes it relational.
What's the point of joining in the application level? If you're going to join, why not do it in the database? That should be both faster and more convenient (unless your schemas aren't relational at all, in which case I don't see why you'd use a relational database)
Another issue is connections between DB servers. Generally you want your DB server to be as fast as possible, so having it handle the connections to other servers slows everything down. If you offload the DB server connections to your web server, you can easily scale by just adding more web servers and having each DB server handle only its own data.
The best solution would be some type of 'mysql proxy' that could run on every web server that would transparently handle the joins between all the different mysql data servers. I think I saw a project attempting to do this awhile back, but didn't really keep track.
> that many deem inferior.
I guess most of those "many" never actually run anything close to the scale of the Facebook (or Wikipedia, which also runs MySQL).
If you bother to pay attention what FB DB engineers say you will know that they have their arms shoulder deep into the innards of the database itself, I/O stack, all that jazz.Most of us can't afford that, and want something that's pretty good out of the box. I've used Mysql a number of times, but it always seems to have more gotchas for what I need to do with it than Mysql.
Lucence base for search. RMDB for money/transactions. And other noSQL type for just data that isn't critical, don't care much about relation, and just need it to be fast.
Of course, perhaps PostgreSQL and other RMDB can have noSQL attributes then noSQL would have some good competition.
PostgreSQL if they can make it easier to clustered, like cassandra or elasticsearch. And get a better fuzzy string and other text search capability in there then now we're talking. But I guess it's just me dreaming.
It's actually the smaller SQL databases (MySQL, PostgreSQL) that are being used for non mission critical uses.
I think it was a failure because of the complex law background and interfacing with all kinds of legacy systems it has to draw data from.
And from reports coming out it was the Oracle Identity system that was causing many of the headaches.
Their users don't want innovation. They want a stable platform with moderate and infrequent changes.
Also, it would be nice of you to step in and contribute to an open source database engine if your expertise is such that you can dismiss what FB did with them as lacking innovation. I'd wager they'll be all ears open for your explanation how to do multimaster replication and distributed joins efficiently while remaining ACID compliant.
The first layer is business need, without one there is no point. You need to start here. Facebook is empty here, so no need for innovation and no concrete ideas about what to innovate about, means no point discussing arch or later tech.
The next layer is architecture. Do you need joins? Well then use them. Do you need transactions, or not? Basically a big key value store?
The final step / layer is tech. You can do a boring key value store on mysql just fine. There might be a slightly faster or cheaper software choice. Or maybe not. Or maybe some nosql thing fits the architecture.
This is a huge problem I see with "the whole market is going to be disrupted totally and only I know how" type articles. So I herd cats around a quasi-geographic database at work (not directly geographic, but related). The only way to change its scale or business needs is to modify the geography of the United States or dramatic demographic changes that would even theoretically take generations to breed the eventual customers. So layer one can't change for that job.
The arch doesn't seem like it can be improved very much. So layer two can't change for that job.
Therefore because layer one and layer two are static, layer three will utterly change completely. Because he says so. LOL
http://www.h2database.com/html/main.html
I can't think of a better reason NOT to use it. Can someone who knows a lot about db please comment on this.