Linking a few trillion records doesn't seem that difficult. It should be doable with a good data warehouse and a reasonable entity linking model. I suspect that we'll find more than a few instances of fraudulent behavior once the data is linked.
My father was nearly pushed into ~2 Million dollars worth of brain surgery that was unnecessary. Not only was the procedure unnecessary, the price for it was >5X what a top-3 hospital would have charged. I only became privy to this once I pushed him to come to Mass General Hospital (MGH) for a second opinion. The surgeon we saw at MGH also believed the suggested procedure to be dangerous.
I wonder if it's possible to cross-reference mortality/complication rates with prices...
with open(filepath) as fd:
first_line = fd.readline()
cols = []
for col in first_line.strip().split(','):
col2 = f'''"{col.strip('"')}" text'''
cols.append(col2)
cols2 = ','.join(cols)
print(f"create table {table_name} ({cols2});")
print(f"\copy {table_name} from '{filepath}' csv header;")
this variant will ingest whatever trash is in your CSV fields as-is (cast & cleanup later)run the output in a psql instance connected to your db
(important note: \copy is a psql client command and it is critical to use \copy instead of COPY in many cases where the server process may not have the permission to read your CSV file. with \copy you can read any file the user that launched psql client has permission to read. to make things more confusing it is indeed possible to stream stdin through psql but you use the regular COPY for that instead of \copy)
If you split on commas, your code will fail for quoted fields with commas in them.
Never heard of ndjson, can’t see one publishing this data in a format that isn’t nearly as common as something like csv (or regular json which some of the data is published in).
Not sure why NDJSON is considered simpler, as json objects can be arbitrarily nested. Breaking into records is easier, but parsing is harder.
And following the CSV spec is much easier than following the JSON spec. And there's only like three edge cases.
Unless you involve MS Excel or, worse, MS Excel on macOS. OTOH, pitfalls are: UTF-8 BOM, comma vs semicolon, single vs double quote, multiline cell content and escaping, escaping in general...
JSON and NDJSON can also be much larger than CSV files if the wrong structure is used.
(disclaimer: I'm one of the zsv authors)