"Oh, they're going to force us to publish our prices are they? Well we'll publish so much data it'll take a herculean effort to make it readable to anyone that doesn't work in data engineering"
"Oh, they're going to force us to publish our prices are they? Well we'll publish so much data it'll take a herculean effort to make it readable to anyone that doesn't work in data engineering"
Linking a few trillion records doesn't seem that difficult. It should be doable with a good data warehouse and a reasonable entity linking model. I suspect that we'll find more than a few instances of fraudulent behavior once the data is linked.
My father was nearly pushed into ~2 Million dollars worth of brain surgery that was unnecessary. Not only was the procedure unnecessary, the price for it was >5X what a top-3 hospital would have charged. I only became privy to this once I pushed him to come to Mass General Hospital (MGH) for a second opinion. The surgeon we saw at MGH also believed the suggested procedure to be dangerous.
I wonder if it's possible to cross-reference mortality/complication rates with prices...
Never heard of ndjson, can’t see one publishing this data in a format that isn’t nearly as common as something like csv (or regular json which some of the data is published in).
Not sure why NDJSON is considered simpler, as json objects can be arbitrarily nested. Breaking into records is easier, but parsing is harder.
And following the CSV spec is much easier than following the JSON spec. And there's only like three edge cases.
Unless you involve MS Excel or, worse, MS Excel on macOS. OTOH, pitfalls are: UTF-8 BOM, comma vs semicolon, single vs double quote, multiline cell content and escaping, escaping in general...
JSON and NDJSON can also be much larger than CSV files if the wrong structure is used.
(disclaimer: I'm one of the zsv authors)
with open(filepath) as fd:
first_line = fd.readline()
cols = []
for col in first_line.strip().split(','):
col2 = f'''"{col.strip('"')}" text'''
cols.append(col2)
cols2 = ','.join(cols)
print(f"create table {table_name} ({cols2});")
print(f"\copy {table_name} from '{filepath}' csv header;")
this variant will ingest whatever trash is in your CSV fields as-is (cast & cleanup later)run the output in a psql instance connected to your db
(important note: \copy is a psql client command and it is critical to use \copy instead of COPY in many cases where the server process may not have the permission to read your CSV file. with \copy you can read any file the user that launched psql client has permission to read. to make things more confusing it is indeed possible to stream stdin through psql but you use the regular COPY for that instead of \copy)
If you split on commas, your code will fail for quoted fields with commas in them.
Basically a document dump - https://en.m.wikipedia.org/wiki/Document_dump
It might be a little exciting to be an underwriter right now :D
The problem is that the health care costs situation results in many deaths and very severe economic consequences for much of the country.
Until lying, cheating, and scheming, and screwing over the public have consequences like prison time, you can expect executives to do everything possible to avoid complying with the spirit of laws like this.
There probably was an effort to create a more useful and sanely worded law that would provide a uniform format for rules that could reduce dataset sizes by a factor of 100, but was killed by the healthcare industry because it would require some implementation costs on their end and make the data files actually useful.