edit: As I reflect, I'm amused to recall that this was early enough in my path that I didn't know about DB indexes, so I was very proud that I figured out how to basically roll my own indexes by pre-sorting the columns by lat and lon. I don't remember whether my solution actually prevented a full-table scan, but it felt like a major breakthrough at the time.
I'd be very curious to read more about the data cleaning phase when you get there. Specifically, how hard it is to combine this data and construct good schemas.
It's entirely possible that two surgeons with offices next to each other could be getting reimbursed at wildly different rates for their most common procedures for their most common procedures by the same provider.
If you're that provider, you ABSOLUTELY want to know what the surgeon next door is getting paid the next time your group is negotiating with the insurance provider.