[1] https://en.wikipedia.org/wiki/Object-relational_impedance_mi...
I think a more correct analogy would be that table are like classes, columns are the properties, and rows are instances. And so defining foreign keys is like setting a pointer to a parent instance.
There is not direct analogy for methods, but you can use function/trigger to do the same job.
PostgreSQL is actually an object-oriented RDMBS, it's not because you are meant to manipulate these objects through SQL that they are less powerful. And SQL is actually Turing Complete with PostgreSQL.
It's clearly not convenient for general programming, but as soon as data manipulation is involved you benefit from a lot of built-in optimization.
This is why object/relational mapping (ORM) has been called the Vietnam of Computer Science[0].
[0]: http://blogs.tedneward.com/post/the-vietnam-of-computer-scie...
And yes it breaks down because it is similar but different. Row "instances" can perfectly exists outside a table and even be defined outside any table via composites types or even on the fly.
Just saying a RDMS can be a powerful object-oriented environment using solely SQL and no ORM at all.
tables are like functions, columns are the arguments, rows are the invocations
I can't agree with that statement, though I'm sympathetic to why you would say it. A foreign key is something that essentially doesn't translate from the world of Tables/Rows to the world of Classes/Objects.
In the relational data model the foreign key is a convenient place to include an index for joining on. What does join even mean between two classes? The closest you'll typically get is nested objects, but that's not quite right. A row produced from table joins isn't a nested data structure. It's still flat. The join operation in relational algebra has no direct equivalent in OOP.
That's all an ORM is, a good default gateway to the tables that you don't have to write by hand.
Around this time, NoSQL databases started popping up, and a lot of my colleagues moved on to them since it fit so well with the trend toward denormalizing everything. Some loved them and dove in full-force.
Personally, I kept working with normalized data, used a caching layer to handle the denormalized versions of the data and learned more about scaling with Master / Slave configurations, and honestly felt very much like I was being left behind.
In order to see what the fuss was about, I tried a small personal project with MongoDB (this is pre-redis, I think), and honestly enjoyed the simplicity of it. And then my toy project got a sudden popularity bump from 20 users to 40k in a day and my project just died. I spent two weeks trying to keep up and then kinda got it working, but couldn't keep the site alive for more than a day. And since it was just a toy project, I just gave up completely.
I've only used non-relational stores for portions of projects that explicitly warrant them, since.
Another aspect of "too complex" is sometimes the true data structure of the problem is correct and also too complex. Some programmers bite off small chunks and chew on them, then push back that the entire data structure of the problem should become the small successfully chewed up chunk. A model that encompasses all of the concept of "number" is very complicated, so a programmer starts off writing a small simple integer library and pushes back that the definition of number should become the simple integer... then reality impacts and as the system evolves and demands are made, the concept of "number" needs floats, rats, complex, base conversions, maybe worse, and the original very complicated design is the only successful way to implement the business requirement of "number". Whoops. If you baked into the cake at the start how to handle rats with zero denominators life would be a lot easier and safer than bolting it on later or praying the app level code handles it, for example.
Loved the vampire/garlic analogy!
Fyi... the "SQL vs noSQL" means at least 2 different ideas and some of your replies are highlighting one aspect but not the other.
1) "SQL vs noSQL" can mean "SQL syntax (e.g. joins) vs object syntax to save/retrieve documents (e.g. Javascript JSON/BSON)" . This is probably the main driver of MongoDB adoption. (They don't care about Mongodb's scaling aspect; they just like the easier syntax.) The counterpoint to this idea is that "people are using terrible db engines like MongoDB because they are unwilling to learn SQL syntax."
2) "SQL vs noSQL" can mean "OLTP RDBMS engine (e.g. MySQL/Postgre/Oracle) vs distributed db engine (e.g. Hbase/Cassandra/etc)". The counterpoint to this idea is that "people are deploying to distributed db engines when their use case actually fits in a 1-node RDBMS engine."
From your wording, it looks like you're talking about #1. To that point, many programmers don't like the compexity of 10-way joins of a dozen normalized tables to reconstruct a customer data entry screen. With a document-oriented noSQL db, you just retrieve the entire denormalized "document" with no joins. However, there are tradeoffs to the "easier" noSQL engine such as performance.
Historically databases got used a lot where the short term cost of normalization was low and the long term cost of denormalized data was extremely high, like accounting or bank records or medical records.
The world has some data store applications in completely different environments, where the cost ratios are wildly different and the bandwidth load is extremely high. Like feeding the whole logging output of a webserver cluster into a DB for "data mining" or whatever. So minimum total cost in those weird new applications involves doing the opposite of what usually is remains correct in the older, still profitable applications.
Inevitably resume stuffing being what it is, neophillia, you end up with people trying to pound square pegs into round holes to boost their resume, look cool, gain experience, or for the sheer joy of trying something new. There's also the rush of transgressive behavior, short term thinking, etc. So naturally you get the extreme over-reaction of people arguing that corporate financial records should be nosql or your bank balance or credit record should be based on nosql, etc.
1. Bad experiences with cumbersome ORMs and refusal to learn how databases work: basically fighting with performance or normalization problems and concluding the problem was the concept rather than using it poorly (“JOINs/subselects are hard so I'm doing 10k queries. SQL isn't web scale!”). A related component of this was optimizing for only part of the problem – e.g. someone's writing a web form and they found it appealing to slap arbitrary key:value pairs into a NoSQL store because they hadn't gotten around to writing the validation, reporting, etc. code which actually needed structure.
2. Hype, hype, hype: people would look at papers coming out of large places like Google, Yahoo, etc. with impressive numbers and think they needed the same infrastructure to impress the other cool kids, ignoring the fact that those papers mentioned traffic / data volumes many orders of magnitude higher and the huge companies could hire more engineers to make up for the extra time needed to hit that kind of scale. A lot of that thought was influenced by the pre-SSD era so the assumptions about when you exceed the limits of a single server really don't hold up to serious consideration now, or often even back then.
Another problem is trying to store dynamic schemas in SQL, e.g. data for a CMS with user-defined entities. You either have to expose a version of SQL to the client, which is dangerous and/or a huge hassle (see above), or you have to implement SQL-in-SQL with rows masquerading as columns. Neither is ideal.
For me it was a much different problem though that sold me: SQL has no concept of revision control or conflict resolution, and all update tracking must be done in-band, with manually incrementing versions and timestamps. Doing master-master style synchronization between two SQL databases (e.g. server and client) is a pain. Doing it with e.g. CouchDB was a breeze, because of its git-like revision tracking. Being able to shunt JSON in and out was a huge benefit, as simple arrays and hashes do show up everywhere. Being able to suck in data from production into a developer database using its built-in replication features was great too.
If SQL serves your purposes well today, that's great, but there are plenty of reasons to want to move beyond it. I'd like whatever Post-SQL is to handle nested data types, particularly if it comes with functional-programming-style algebraic closure of the resulting constructs. SQL-in-SQL should just be SQL.
In the meantime, I will just design data-first and acknowledge that how I _pull out_ the data is the main constraint anyway, SQL or not. For arbitrary queries, there's dedicated indexing and searching solutions like ElasticSearch that can actually deliver on doing that efficiently, without lots of careful babysitting.
To speculate, noSQL key-value stores became very popular because they allow you to model imperfectly defined situations, and to update that model quicker than in traditional relational-database type situations.
Consider the difference between writing some classes in Java, or bunging everything into a dict in Python. The Java solution could be more formal and well documented, but a pain in the ass if the model changes significantly. The Python solution is a bit more janky and probably sluggish, but can adapt quicker.
I think there are more people in things like research, self-teaching, iterative game dev, and exploratory startups that value the adaptability than there are in things like old school business consultancy for whom the underlying model might be static and well described.
Many sites and apps found themselves in a situation where they would gladly trade strong consistency, for more performance and eventual consistency.
Further, schema changes for relational databases with billions of rows have to be very carefully orchestrated to avoid downtime. Databases with more flexible schemas allow schema changes more easily (though obviously the code most accommodate the fluid schemas).
Various NoSQL solutions were developed to fill these needs.
Sure, they were overused for a bit (though that fad has basically passed), but they definitely arose to fill a real need.
I don't believe this is the best way to decide on a technology.
Of course, the reality of the situation is all that matters, and most database engines at that time did not offer such features. So, I don't necessarily fault people for trying to find solutions that simply worked for their needs.
"noSQL" should be read as "not-only sql" and I think sql in this acronym should be interpreted as "relational databases". I think the relational model is convenient for a lot of purposes but at some point (maybe ~10 years ago) some companies where trying to use it for everything even when those things didn't actually fit in that model. So, "noSQL" for me actually means "alternatives" for different kinds of data time-series, graph, key-value, documents, etc. with redis, neo4j, graphite, mongodb, etc.
What I would like to see in the future is some ecosystem based on small pluggable modules that you can use to build the embedded store for a micro-services. Instead of using a full-feature database system (relational or not) you assemble your store like Legos with only the stuff you need. Something like this https://github.com/level/levelup/wiki/Modules
everything you work on in school is not like this. you spend a lot of time studying in-memory data structures and doing tiny projects (ranging from 100 lines of code homework all the way up to 5000 lines of code midterm project) that implement in-memory data structures and algorithms to work with them.
so you get some inexperienced developer who only has their schooling to inform them and they're all stuck thinking in terms of arrays, lists, hashmaps, and trees. learning a new data model has a lot of cognitive overhead so the allure of a data persistence tool that gives you almost the same data structures you already know is very strong. "Why would I learn about relational data modeling when everything is just a hashmap anyway?" thus MongoDB as extremely popular with 23 year olds and extremely scorned by 33 year olds. 10 years of job experience drives the point home.
Actually, pre-RDBMS those were the dominant systems. The concepts invented from RDBMS were the new kid on the block that was hot and cool and more powerful that supplanted the original "NoSQL" system and had remained dominant until we hit "Internet scale" and people wanted to 1) scale to huge data volumes requiring distributed databases and processing cheaply and 2) wanted to return to finding alternatives to the dominating SQL platforms.
SQL supports linear ordered or even branching versioning fairly easily (though only recently have many SQL-based DBs had decent tools for temporal versioning), and SQL-based object-relational databases (Postgres and Oracle, for example) have supported both document-oriented views over classical relationally-structured data and document-oriented data storage since well before the NoSQL craze.