The HIPAA and other regulations aren't any more annoying than any other modern programming practices these days. As for liability in the case of a breach, that's what E&O insurance is for. It does raise the barrier of entry a little, but it's not by any means prohibitive.
I think the bigger issue is what format this data is in. Most medical and medical records data is in a variety of proprietary, non-open and difficult to integrate technologies. One look at HL7 is usually enough to send a programmer back into the loving embrace of anything else.
Working with this data is time intensive and expensive and I'm guessing most heathcare companies don't see it as worth the cost.
It's not worth the cost 99% of the time. The only reason we're still making real breakthroughs is because of the research institutes doing basic research on public funding. Even then, they're swinging for the low hanging fruit.
> I couldn't have gotten out of there fast enough for the way saner land of tech companies.
You know you have a problem when tech companies seem saner in comparison.
In many cases scientists are required to submit data to journals when they publish. However, data is a funny term! Using genomics data, there are many levels of data that you'll encounter - everything from raw image files (many gene sequencers are actually automated digital cameras, taking pictures of florescent markers attached to the DNA strands) to intermediate sequence files to formatted pieces of selected data. Then, there's the whole toolchain used to go from sample to formatted final analysis? What's required to be submitted? How long should it be kept? Who checks all this to make sure it's not NES roms or Shakespeare instead of the right data?
Finally, there's the big questio: how can we be sure that the data captured in the intermediate or final steps of analysis actually originates from the raw data? Should scientists store the raw data (TBs and TB), intermediates (GBs and GBs), or final analysis (MBs). For how long?
To answer your question about meaningful long term research - I've personally seen grad students' careers effectively ruined due to shitty data storage hygiene.
As someone in the field I beg to differ :). And regardless of who pays in the event of a breach, the effects may be sufficient enough to shut down the company for future projects.
You can basically self-certify, but most serious companies will bring in an outside contractor on an ongoing basis to certify compliance. Staff needs to be trained, computers need to be managed, software changes have to be very thoroughly reviewed, updates become slow. It makes it pretty unattractive to enter into for a lot of devs.