Databases for data scientists
josiahparry.com
josiahparry.com
Databases weren't big in science because they used to be row oriented and most questions in science are column oriented. Most workloads in science are of the map, filter, reduce variety. If you used a regular rdms for those it would have extremely poor performance and scalability.
Computers have caught up with most business sized datasets. Today you can fit every human begin alive in a database that can fit on a single Postgres database without too much trouble and even run queries against them. This was not the case in 2012 when the term was invented.
Nontheless, every person working with data and programming should know database concepts.
You should be keeping track of those things, but I see little benefit and lots of downsides in using an SQL database to do so.
It’s a tool that consistently gets underestimated at the same time it’s also consistently improving.
I hate XLSX with a passion. The main issue is that as a file format it's too heavy, slow, clunky, and does weird data type conversions, but that's not the only problem.
The main problem of excel is "mission creep". It just does too many things in a half-assed way. So when you get a bunch of excel files to process for your data analysis task, you can't be sure someone didn't put a bunch of charts, merged cells, used formulas, used strange currency or unit formatting gimmicks, structured the data in a weird way to make it visually appealing to management, and so on. It's all so tiresome.