I had to ingest and explore enough dirty CSV/TSV data to use CLI tools only to get a first glimpse what's there.
Whenever facing anything non trivial I go for CSV to parquet, and preferably write intermediate parquet data sets. Then DuckDB / SQL queries for slicing and dicing.
I am yet to encounter some readable to a newcomers awk/uniq/sed combos for intersecting bunch of CSVs with 30+ columns.
On the other hand parquet becomes a turtle if one tries to squeeze i.e. 12k numerical columns into it.