HNHacker News
TopNewBestAskShowJobs

visualsearchsv

88 karma · joined January 28, 2016

submissionscomments
visualsearchsv··on Schizophrenia, Hubris and Science
""" We'll only start making real progress in curing these kinds of illnesses when we form a deep understanding of whatever is going on in that very large component that we already know is not purely genetic. """

You have zero clue regarding what you are talking about. Your comment is just meaningless combination of platitudes. Even if we find a target and develop a drug that treats 20% of the cases, it still counts as a real progress if not a revolution. Look at the other comment in this thread by Chris Patil.

As you say in your "But I'm no scientist or medical expert", yes that's the reason why you find an obviously flawed blogpost appealing. Since the post was written/designed to make you feel better about your pessimism/skepticism.

This trend of "Contrarianism for the sake of contrarianism" is a sophisticated version of click-baiting where bloggers/commenters displaying classic case of dunning kruger spread FUD about any new discovery/result.

Commenters on Reddit/HN quickly lap up these articles/comments since it resonates with their sense of technological pessimism, often disguised as skepticism. As a result its now fashionable to criticize every new discovery/progress and get guaranteed pageviews/upvotes. It helps if its in form of a scary slideshow or contains assertions or scenarios that are not under question.

visualsearchsv··on Ask HN: What should we fund at YC Research?
The second article is undoubtedly flawed in several aspects, that's why I put clifton et. al. first, which I think lays out the case for studying non-formal models.

Regarding

"""This also, however, ignores the fact that most statistical inference drawn from such queries will be nonsense even without differential privacy."""

This is not true. Just because the number returned by a count query on a very large dataset (~100 Million visits) is very small (~100) does not automatically means that the result is nonsense or can be disregarded as error. Doing that requires understanding the query and a hypothesis with good prior on expected outcome. E.g. intersection of two rare diseases. Where you would otherwise expect it to be very small, but there might be an underlying genetic reason / physiological process which might lead to higher prevalence. Or a group of hospitals using tainted batch of medicines leading to unexplained increased mortality.

Consider this paper where there were only 1000 cases (only 248 strokes) per 1.6 Million patients (even larger if you consider the entire 20 Million patients present in the data). However in spite of the small number the authors showed that the increase was statistically significant by comparing with same period a year later.

http://www.nejm.org/doi/full/10.1056/NEJMoa1311485

Again I am not denying what you wrote in the blogpost. But in medicine and the analyses for which such databases are used, the investigators have access to very good priors.

visualsearchsv··on Indian Women Seeking Jobs Confront Taboos and Threats
I wish that NYTimes would differentiate between Individual indian States. Individual states have significant autonomy and depending upon the government elected can range from moderate to leftist. Largest indian states have populations similar to large nations such as Mexico ~ Maharashtra. And have huge variation in metrics like literacy rates ranging from 93% (Kerala ~ Canada ) to 63% (Bihar ~ Philippines).

~ is used to show country with similar population.

https://en.wikipedia.org/wiki/Indian_states_ranking_by_liter...

http://www.economist.com/content/indian-summary

visualsearchsv··on Ask HN: What should we fund at YC Research?
This data is available through Agency for Healthcare Research & Quality HCUP project. Getting access to it is straightforward if you are affiliated with a teaching hospital and/or a university. There might even be some researcher at your institute who already has access to it. Getting access to it as a private entity (E.g. a startup) is more challenging and often requires a stricter review.
visualsearchsv··on Ask HN: What should we fund at YC Research?
We do something similar. We precompute/aggregate exhaustively by following certain aggregation strategies. The aggregated statistics are further processed to ensure privacy.

Differential Privacy cannot be directly applied since the underlying assumptions are too strong. An important consideration is that the error/noise added is independent of the answer. Which means that the system becomes unusable for almost all queries other than general trends.

By restricting the query structure, we no longer need large amount of noise. Privacy of hospitals and providers is also very important and cannot be encoded in the Differential Privacy framework. Again this is still a hotly debated issue. But even the most vocal supporters of differential privacy agree that it might not be directly applicable for healthcare domain.

Following are some of the paper that discuss this:

http://www.openu.ac.il/personal_sites/tamirtassa/download/co...

http://www.jetlaw.org/wp-content/uploads/2014/06/Bambauer_Fi...

visualsearchsv··on Ask HN: What should we fund at YC Research?
Hi zo1, its perfectly legal, in fact the entire program (HCUP) is carried out by the US federal government agency itself (Agency for Healthcare Quality and Research) for last two decades. We have only made the system available internally to a select group of doctors who sign agreements with the US government. Misuse of this data is punishable by a felony. I have personally met the director of the agency and they know about it, and have seen a demo.

Regarding your point second point, the patient ownership of the data is not very well understood legally. Since most of it is transactional in nature. E.g. consider the Sorrell vs IMS Health judgement by the Supreme Court few years ago. Following paper in Harvard Journal of Law and Technology gives a good overview of the issues.

http://jolt.law.harvard.edu/articles/pdf/v25/25HarvJLTech69....

visualsearchsv··on Ask HN: What should we fund at YC Research?
Thanks, Alex. I agree that such system should be global / non-profit & free. We have also studied privacy issues surround such system in detail. A large motivation is that regardless of the privacy preserving technique employed the system is a huge improvement over current practices which involve sending out entire data to each individual research group.

Today a physician suspecting novel association e.g. Adverse event particular to a co-morbid condition. Usually has to wait multiple weeks for Ordering data from US government, Developing SAS scripts and finally conducting the study. Wrth our stack the underlying question can be answered within seconds. With enough legal permissions we can modify it to deliver only the required data, for further statistical analysis. Sure such a system might not be completely public, but there is nothing that prevents us from giving access to say all board certified physicians globally.

visualsearchsv··on Ask HN: What should we fund at YC Research?
Large Scale Medical Data Mining research, similar to OpenAI. Specifically Computational Healthcare a Search and Aggregation Engine for Medical Records & Claims. We believe that this is a classic Software eating the world situation and the time is perfect for it.

Here is the link http://www.computationalhealthcare.com

We have access to almost 130 Million de-identified medical records from approximately 36 Million patients (~10% of US population) this includes all Inpatient, ED, Ambulatory Surgery records between 2006-2011 from California. To put simply if you lived in California and went to a hospital there is 95.9% chance that we have your data. This data has been available for quite some time but its use has been hindered due to lack of good software. The data has led to significant research, e.g. my collaborator (not me) published a paper showing risk of strokes following pregnancies in New England Journal of Medicine last year.

At Cornell Tech & Weill Cornell Medical College, we have developed a Search and Aggregation engine that will revolutionize how researchers and physicians use this data. Imagine your mother with Leukemia in Remission just got admitted for Pneumonia. With our software, the Physician will be quickly able to asses likelihood of this occurring and rule out any confounding adverse events. Or consider that there is a rare combination of diagnosis e.g. Graves Disease and Clotting disorder that is indicative of a unique genetic mutation likely to offer novel insight into disease process. With our software questions like these can be answered within second, Today & Right now.

The Data, Legal structure and fully functional prototype are available right now. We were counting on support from AHRQ, but sadly the agency has run into trouble due significant budget cuts.

visualsearchsv··on Show HN: Visual Search using features extracted from Tensor Flow inception model
Yes, being able to compute quickly is especially important in reducing query latency, much more so than during indexing. What stood out for me in the paper was that out of box performance of VGG (trained on imagenet alone) was as good as fine tuned alexnet.

I am interested in assessing if there are any tricks that could be used when querying from a mobile device. In such cases feature extraction can be performed on the device itself, with only feature vectors sent over the network. In case of pinterest, another special case is that a lot queries are performed on images already present in the system. The user simply readjusts the bounding box to highlight the object of interest. In this case they can simply pre-compute 4~20 crops per image. Online feature computation is much more expensive / complicated than offline.

visualsearchsv··on Show HN: Visual Search using features extracted from Tensor Flow inception model
According to their KDD paper, they showed modified / trained alexnet performing as good as VGG with significantly less computation time. Not sure what they actually use in practice.

http://www.kevinjing.com/visual_search_at_pinterest.pdf

visualsearchsv··on Show HN: Visual Search using features extracted from Tensor Flow inception model
Inspired by the Pinterest paper at KDD on implementing Visual Search, I have created this barebones but functional implementation of a Visual Search server using ~450,000 Female Fashion images crawled from an image aggregation website. The code uses Pool 3 layer features extracted from the Google's latest Inception model using Tensor Flow. The extracted vectors are then indexed using an Approximate Nearest Neighbor implementation from Nearpy. The AMI provided contains both images and pre-computed vectors.

In future I plan to add more images ~2 Million in the same domain. Test various combination of nearest neighbor indexes and multiple vectors per images using some form of a multibox style detector. I will also add a script to launch spot GPU instances via Cloudformation to economically index images using S3 and SQS. I am building a companion iOS swift app however, since Tensor Flow hasn't been ported to iOS yet, its still in development.

https://engineering.pinterest.com/blog/building-scalable-mac...