Show HN: Visualization of Longevity and Mortality
infino.me
infino.me
In particular I think there need to be options added to normalize the y-axis of the graph beyond raw counts to percentage of deaths by age group to control for the fact that the number of deaths per bin is not the same.
It would also be nice to add a little more categorization to the causes of death, in a tree-like structure. For example, all vascular disease, with cerebrovascular disease, CVD, etc, as subtypes.
Also it would be nice to be able to ask questions like, "which states have the most (or least) fraction of deaths by, e.g., CVD"? Do some states have smaller or larger gender gaps in particular diseases?
I come from an academic background. Made this in my PhD: amass-db.org
But my core passion this days is the project that this viz is hosted at (www.infino.me). I think I can make a much bigger impact with more consumer-facing nonprofit apps than in publishing academic articles.
I would be interested in your rationale behind how the public would/could use this kind of data. I think the infino.me/health seems to be an example, pointing out major risk factors behind cancer and CVD. But don't you think this is already common knowledge?
Or is the primary goal to get people to voluntarily share their health and genotype data to get a big dataset for analysis, and maybe eventually provide a sort of personalized risk assessment? I wonder how the FDA views that sort of thing.
I did my PhD work in diabetes genetics. It runs in my family, and Im kinda pissed that its killing off a large fraction of us. The eventual goal is to make this into some kind of communal science effort where people can contribute open source analysis pipelines. I need the right kind of organizational structure however to keep the data safe and centralized but still allow open source research.
So, some kind of platform where algorithms can get in, results can get out, but raw data stays locked up. I want a world where this kind of research happens in the open, and not privately in biomedical corporations.
What pisses me off is how much can be done to prevent it and how little people are willing to do so, especially when their genes can tell them point blank that they really ought to be doing something and doing it right now before the very bad thing happens that will negatively impact them forever.
Much deeper, and more valuable, would be information on the optimal time of day to eat, exercise, etc, and more information on the specifics of the exercise needed (in terms of intensity or heart rate zones).
TLDR, exercise matters, so does genome, but what we really want is more specifics. This might inform better lifestyle adjustments.
Which is to say while I think you could build a billion dollar company out of that last 20% (and you can correct me on the percentages here, really, please do so), it would have to not only provide such information, but also present it in such a way that the recipient got up and did something with it. And that's the hard part, no?
But although "algorithms in/results out" sounds good in principle, I think it will be hard to implement in practice. You would have to make algorithms run without network access to prevent a bulk_send_data_to_ip() type of function from being written, but that would hamper complex programs requiring external data.
In general I think the only realistic way forward is to take the 1000 genomes approach of finding people who are willing to take the privacy risks of truly open-sourcing their data. But it sounds like an interesting idea and I hope I'm wrong and your approach turns out to be workable.
Its a continually evolving thing. I imagine it would be years before I get to that stage. Depends on if I find funding or university help.
Hit me up at info@infino.me if you'd like to chat more.
Made using the dc.js library. It's interactive! You can click on any of the plots to refilter the data.
Sorry mobile users! The best Ive been able to do is show a screenshot for many mobile browsers. Even so, the draw performance is terrible since there is so much data being loaded. While smaller dc visualizations can work on a mobile device, its just too dang slow to show 10's of thousands of records. A laptop however has no problem.
On the frontend, there really isn't anything new aside from whats already been featured in the many examples for dc.js: https://dc-js.github.io/dc.js/
What people might benefit from is the python scripts Ive used to processes the CDC data so that it is nicely compressed.
My core vision of this project is to better understand this overlap.
Thanks for reporting the bug.
I burn ~3200 calories a day and I walk >10 miles a day. This is the 100th percentile of your data and that floors me.
So I have to say that step 1 is getting people to get off their asses and move. It'd be nice if they cut down on red meat and sugar intake while they were at it, but small steps, no?
I do wish there were a way to anonymously calculate the very comparative statistics you're generating, but alas I don't see one. Am I missing something?
Really what we need is some kind of homeomorphic encryption process that would enable large scale Genome Wide Association Studies (GWAS) to be performed without actually divulging the underlying raw data. Until that happens though, scientists like myself have to contend with the privacy concerns of many.