HNHacker News
TopNewBestAskShowJobs

cafebeen

462 karma · joined December 6, 2014

https://cabeen.io https://www.linkedin.com/in/cabeen/ ryan@saturnatech.com
submissionscomments
cafebeen··on An adult fruit fly brain has been mapped
Budgets are finite, and most science funding involves some decision making about how to allocate resources.
cafebeen··on One dead as London-Singapore flight hit by turbulence
It’s more like keeping a fire extinguisher in your kitchen. Some people may never need it, but grease fires happen often enough to plan for it.
cafebeen··on After 14 years in the industry, I still find programming difficult
Greg Lemond wrote about fitness, "it never gets easier, you just get faster", which perhaps applies as well to programming.
cafebeen··on GNU Octave
I think you're both right, but you might be thinking of different tasks and libraries. If we're talking about solving a specific type of PDE, then in that scenario he's probably right that MATLAB will work better out of the box, but if we're talking about logistic regression on data from a messy CSV table, then python will work great and have better usability than MATLAB (IMO). He might also be thinking of versioning and packaging headaches in python, which are quite a pain the first time exposed to them coming from a walled garden environment.
cafebeen··on Brain-Imaging Studies Hampered by Small Data Sets
Here's a link to the associated paper just out in Nature:

https://www.nature.com/articles/s41586-022-04492-9

Marek et al. Reproducible brain-wide association studies require thousands of individuals (2022)

"Magnetic resonance imaging (MRI) has transformed our understanding of the human brain through well-replicated mapping of abilities to specific structures (for example, lesion studies) and functions1,2,3 (for example, task functional MRI (fMRI)). Mental health research and care have yet to realize similar advances from MRI. A primary challenge has been replicating associations between inter-individual differences in brain structure or function and complex cognitive or mental health phenotypes (brain-wide association studies (BWAS)). Such BWAS have typically relied on sample sizes appropriate for classical brain mapping4 (the median neuroimaging study sample size is about 25), but potentially too small for capturing reproducible brain–behavioural phenotype associations5,6. Here we used three of the largest neuroimaging datasets currently available—with a total sample size of around 50,000 individuals—to quantify BWAS effect sizes and reproducibility as a function of sample size. BWAS associations were smaller than previously thought, resulting in statistically underpowered studies, inflated effect sizes and replication failures at typical sample sizes. As sample sizes grew into the thousands, replication rates began to improve and effect size inflation decreased. More robust BWAS effects were detected for functional MRI (versus structural), cognitive tests (versus mental health questionnaires) and multivariate methods (versus univariate). Smaller than expected brain–phenotype associations and variability across population subsamples can explain widespread BWAS replication failures. In contrast to non-BWAS approaches with larger effects (for example, lesions, interventions and within-person), BWAS reproducibility requires samples with thousands of individuals."

cafebeen··on Scientific progress despite irreproducibility: A seeming paradox
A good example from physics is this observation related by Feynman in 1974:

"[...] It's interesting to look at the history of measurements of the charge of an electron, after Millikan. If you plot them as a function of time, you find that one is a little bit bigger than Millikan's, and the next one's a little bit bigger than that, and the next one's a little bit bigger than that, until finally they settle down to a number which is higher.

Why didn't they discover the new number was higher right away? It's a thing that scientists are ashamed of—this history—because it's apparent that people did things like this: When they got a number that was too high above Millikan's, they thought something must be wrong—and they would look for and find a reason why something might be wrong. When they got a number close to Millikan's value they didn't look so hard. And so they eliminated the numbers that were too far off, and did other things like that..."

cafebeen··on Computational Power Found in the Arms of Neurons
Alas, the possibility of comments like this are one reason that some researchers hesitate to release their code.

Even if there are issues, it's tremendously valuable that they shared this, so others can reproduce and build on their work. That said, most scientists would like to have better quality code in science as well, but I think the incentive structures and other demands that make it challenging...

cafebeen··on Biology is the New Tech: Letter from a conference on CRISPR
Yes, there is surely evidence that intelligence is heritable, but it's worth giving the research on this a close reading. Current approaches can only explain 10% of the variance in intelligence (https://www.nature.com/articles/nrg.2017.104), plus the effects from the environment. My main issue with the article though was the assumption that a "best" set of genes exists without futher discussion (in addition to obvious ethical concerns)
cafebeen··on Biology is the New Tech: Letter from a conference on CRISPR
Agreed that intelligence is a good thing, but the question is about whether there is a set of genes that are “best” for producing it. At the moment it’s unclear which genes those are, and even if we found some genes related to intelligence, they may have other negative consequences that make the idea of a “best” set of genes questionable
cafebeen··on Biology is the New Tech: Letter from a conference on CRISPR
The bigger issue seems to be determining what "best" means. While there are a handful of diseases associated with specific genetic variants that could obviously be cured, the traits people would likely try to optimize (e.g. intelligence) seem to involve a complex combination of genes and environment, and an attempt to select for such genes could incur some potentially negative side-effects.
cafebeen··on Averages Can Be Misleading: Try a Percentile (2014)
I find raincloud plots to be the best of both worlds, and they have been my preference lately:

https://wellcomeopenresearch.org/articles/4-63/v1

They can visualize common statistics like box-plots, the distribution shape like violin plots, as well as the raw data!

cafebeen··on New iPad Air and iPad Mini
Perhaps a person's interests, income, or lifestyle are salient variables as well? I would guess that each person's circle of friends has some of these in common in addition to age (likely true for HN discussions as well). For a less biased picture, here are some statistics about tablet use by age:

https://www.statista.com/statistics/805143/us-tablete-users-...

and for a baseline, iPhone use by age:

https://www.statista.com/statistics/203063/iphone-users-in-t...

It's striking how uniform the distribution of tablet use was across age! There was a bump after age 25, though that was mirrored in the iPhone usage. This is really salient at the low and high end: age 0-18 was 20% of tablet usage and only 6% of iPhone usage, and age 55+ was 20% of tablet usage and only 14% of iPhone usage. Given all of this, it seems like tablets are quite broadly adopted devices

cafebeen··on Programming Books You Wish You Read Earlier
This is essentially a list of textbooks for undergraduate coursework in computer science. I'd guess is the important parts are the amazon affiliate links...
cafebeen··on Neural Networks, Manifolds, and Topology (2014)
In a trivial sense, it seems that a Euclidean topology would do the trick, no? I'd be curious to hear counter-examples, of course. In my mind, I suppose what's missing from the "manifold hypothesis", as stated above, is that the manifolds should be more useful than raw data, for example, does their structure respect our notion of object categories or are they sufficiently low-dimensional for visualization?
cafebeen··on Neural Networks, Manifolds, and Topology (2014)
To play devil's advocate... given that manifolds are so general, wouldn't it be more surprising if natural data didn't form some lower dimensional manifold? That is, a manifold only requires some consistent notion of neighborhoods and a way to locally parameterize objects, which is a fairly low bar to pass, and perhaps trivial if you knock off one pixel in the case of images. Maybe the surprising point (which is not said explicitly) is that the manifold is very low dimensional compared to sensor data, but I suppose it's hard to formulate a hypothesis about the magnitude of the reduction besides it being "lower".
cafebeen··on How to Eliminate the Dreaded “Blind Spot”
These seem like good guidelines, although I would add that drivers should also be aware of other cars that may merge from the side, so you cannot depend entirely on the mirrors. This happens quite frequently in heavy traffic when folks are eager to take any space available
cafebeen··on A Visual Exploration of Gaussian Processes
One major application is in geospatial statistics for a variety of fields that need to perform regression of irregularly samples across space. Although it is typically called Kriging, it is mathematically equivalent to Gaussian processes from my understanding:

https://en.wikipedia.org/wiki/Kriging

cafebeen··on The dubious distinction, and literary legacy, of Leo Szilard
While there are some amusing anecdotes in the article, the characterization of Szilard as a failure is baffling.
cafebeen··on Modern human brain organization emerged only recently
Note, the original paper is about evolutionary patterns of overall brain shape, which is far from indicating how brain organization evolved. The important parts of brain organization are encoded by patterns of structural and functional connectivity, which are internal to the brain and not well depicted by the overall shape. The paper is great, but the write-up just jumps to some distant conclusions.
cafebeen··on MacBook Pro? No
I find the current MBP makes a lot more sense if you mentally substitute “deluxe” for “pro”.
cafebeen··on Paying top employees the highest salaries in the market
i think you mean, there will always be at least one person who is unhappy.
cafebeen··on Apple has acquired Workflow, an automation tool for iPad and iPhone
Reading about Workflow makes me wish HyperCard were still around. If all of iOS was essentially a HyperCard stack, this the kind of automation provided by Workflow would possibly be more natural and powerful. I hope this is a sign that Apple is moving away from the hard line between users and programmers, but maybe that's too optimistic.
cafebeen··on Programs that have saved me 100+ hours
I use make for scientific data analysis, and it has saved me a tremendous amount of time. It's definitely worth the investment to learn the language, in my opinion. I know many folks have tried writing their own version to "fix" the arcane parts, and maybe they're better in some use cases, but make is thoroughly documented and tested and quite general purpose.
cafebeen··on Buck – A build system developed and used by Facebook
Tools like Make are also useful for simple data analysis workflows, and I'm curious to hear any thoughts from any Buck users as to whether it would be useful in those cases too.
cafebeen··on The Competitive Landscape for Machine Intelligence
I think there are many companies currently or on their way to using AI to run a significant portion of their business, e.g. financial and advertising sectors. The hard part is demonstrating intention in that scenario, as many learning algorithms are stochastic. Even if there is some group to be held responsible and they give you the training data, the model will slightly vary when run repeatedly. It seems likely that people could use that fact for things like money laundering and insider trading, for example, using biases in the model that are hard to detect.
cafebeen··on The Competitive Landscape for Machine Intelligence
I think this is a major challenge too, although it depends. Some models are easy to reason about, e.g. Bayesian graphical models, while black-box approaches like deep neural networks are not.

One especially problematic issue is: if a model is too complex for humans to reason about, then a business could encode any kind of illicit behavior in the form of model parameters they like. Even if someone could prove that the model is biased one way or another, there is complete plausible deniability for the business, i.e. "I didn't make that choice, the learning algorithm did". We're in for some very interesting legal battles related to this, I think.

cafebeen··on A bot crawled thousands of studies looking for simple math errors
It's not coddling as much as professional courtesy. Mistakes are often brought up at conferences and in peer review, but most scientists within a specific research area try to be on good terms with one another and don't see value in publicly shaming colleagues for their mistakes.
cafebeen··on A bot crawled thousands of studies looking for simple math errors
Well, scientists are people too, so they have feelings and are not perfect. I think there are polite ways provide criticism and corrections, hopefully without humiliating people with good intentions. If it's a simple mistake, they can email the author and suggest an erratum. If it's more serious, then it's common to write a response in the journal. I think more public forms of criticism are perceived as bullying because a general audience can't typically judge the magnitude of the error, since they're not scientists working in the field, so it can be unjustly damaging to someone's reputation.
cafebeen··on Software faults raise questions about the validity of brain studies
Despite the title, software was not at fault here. Rather, the paper found a higher than expected false positive rate due to a choice in statistical modeling. The conclusion was that nonparametric tests can avoid the issue.

So, none of the results are truly invalid in the way of a hypothetical software bug, since we know what modeling assumptions were used in the previous studies. Personally, I don't see a major problem, since most fMRI papers are exploratory, and we should be reproducing the major findings anyway. We should certainly start using nonparametric tests from here on out though!

cafebeen··on Google Unveils Neural Network to Determine the Location of Almost Any Image
I'm just going to leave this here:

http://graphics.cs.cmu.edu/projects/im2gps/

← PreviousPage 2 of 7Next →