It seems like their database[0] has a column for the `cdm_url` of all of these hospitals. The challenge is like being able to read all these HTML, PDF, XLXS, CSV, etc pages of very different formats and turn them into usable data
Just my guess
[0] https://www.dolthub.com/repositories/onefact/paylesshealth/d...