CSV is a great format for humans to comprehend.
For big data systems, CSV is a arguably worse format
Row wise storage doesn't provide compression benefits like columnar storage which can significantly reduce storage needs.
Data could contain the same delimiters ( comma, newline ) as used by the parsers and introduce errors in computations. binary formats( parquet/ORC) can eliminate this issue while providing other benefits as well.
A data quality framework like deequ should help catch errors in data before you introduce the same to downstream applications in ELT.(not infalliable )
For ELT processes, always filter first before other processing.
Plan your partitions according to expected querying patterns. If reports are run country wise , then country is a good partition key or if date wise then date is a good parition key.
Be aware of data skews , you might introduce skew during partitions ( ex : partitioning by country and a few countries have large number of records compared to others.
Random thoughts
I have a rule that no-one is allowed to use the words big,small,fast or slow. You must quantify.
I've met too many people who think that 100MB is big data or that 1 Gbps is a fast internet connection.
- How big is the average data unit?
- How are you going to analyse and process this data? (What kinds of questions will you ask it?)