I think one thing that's not mentioned enough in the SQL vs NoSQL debate is the benefit of powerful storage types. For example, when storing IP addresses in Postgres, you can use the inet datatype and easily query results if they fall within a given cidr range. Example:
SELECT * FROM audits WHERE ip_address << '10.0.0.0/20'
gives you any matching address between 10.0.0.1 and 10.0.15.254Blindly picking a SQL DB (mysql/postgres etc.) is quite expensive from the get-go (a production ready mysql/postgres would cost ~30$/m).
Mongo costs ~$10/m (MongoDB Atlas), Google's Datastore is Pay as you Go (so your initial cost is close to $0 till you get paying customers), AWS's DynamoDB is similarly priced as well.
Sure, sadly all those noSQL solutions get really expensive as your usage goes up to normal non-webscale proportions, but at that point you have the $ to invest in a SQL solution.
The above was mentioned with bootstrapped startups/services in mind. Not your usual million funded valley companies.
Since it's free below 10K rows and only $9 for 10M. And, the dataclips feature always comes in handy.
For backing up postgres, all you have to do is setup a cron job that backs up the postgres' data directory to S3/Google Drive/Dropbox every hour/day.
If you want proper replication and failover then you can probably use 2 digital ocean droplets each for $5/month and another $5 VM for the application server itself.
If anyone remembers the most prominent and convincing articles on this topic, I would love to share that with my team
They didn't chose NoSQL, they were forced to. I'm fairly convinced they started with relational stores. If a company or product grows to a point where relational data doesn't work, that's a problem you want to have.
The mistake is either thinking you need to design for facebook scale from the beginning OR thinknig that you can cut time in a startup by not having to bother with those pesky schemas that just slow you down.
Yes, I agree that an key-value store that lets you do embarrassingly parallel reads and writes is useful for scaling, but has't it been said enough that You Are Not Google[1]?
[1] https://blog.bradfieldcs.com/you-are-not-google-84912cf44afb
Why not just directly `COPY` your log files to something like Redshift?
The "important" events for data science are still stored on postgres (i.e. logins for checking if a login was malicious)