https://github.com/capitalone/DataProfiler
We’re working to automate much of that as well in a Python library.
The end goal is to:
1. point at any dataset and load it with one command
2. Calculate statistics and identify entities with one command
3. Generate robust reports with one command.
Regarding the data wrangling... Believe it or not, even automatically detecting a delimited file with a header is hard work. Imagine a header can be on the 3rd row and has a title and author ship one rows 1 and 2 respectively. Further, the delimiter might be the “@“ symbol!
The linked library wrote handles that scenario. But that’s just CSVs, there’s also Json, parquet, Avro, etc etc...
This is an extraordinarily deep and complex field.