2,428 karma · joined March 14, 2012
Look at the way we regulate harmful chemicals. After studies find that chemicals found in consumer products are harmful, manufacturers spend several years redesigning their products to avoid them. After manufacturers have voluntarily redesigned their products, the chemicals are banned. The manufacturers often replace the harmful chemicals with structurally similar chemicals that have not yet been shown to be harmful. When those similar chemicals are shown to be harmful, the cycle repeats. This is the history of phthalates, which are used to soften PVC plastic. The CPSC first considered banning one phthalate (DEHP) in children's toys the 1980s. Before that happened, all large manufacturers replaced it by another phthalate (DINP), which also turned out to be harmful. Both DEHP and DINP were finally banned by congressional action in 2008.
Technology is not a chemical, but anything that people will be spending significant time interacting with is likely to have some effect on their wellbeing, potentially small but potentially large, potentially positive but potentially negative. The type of harms these chatbots might cause are difficult to predict in advance. It could take a couple years to determine whether the current generation of chatbots are actually a net positive or negative for society. In that time, they may become too entrenched for us to do anything if they do turn out to be harmful. When Facebook first came out, I doubt many people foresaw the negative impact that it could have on teenagers. Today, many studies show that social media use has a negative impact on kids' wellbeing, but there is no going back to a pre-social media world.
Prospectively testing and regulating tech in the same way we regulate pharma seems insane. OTOH, we do want to make sure that world-changing tech actually changes the world for the better, and the only obvious way to do this is through some kind of regulation. Although regulation will undoubtedly slow growth, many of us would be happy to sacrifice some growth for improved wellbeing. At this point, everyone has seen the productivity-pay gap plot showing that most of the improvement in US productivity since 1979 has not translated into increases in real wages. Our current "growth at almost any cost" strategy benefits corporations much more than it benefits the average person.
In the Recht et al. study, the reason the new test accuracy is wildly outside of a binomial confidence interval around the original test set accuracy is that the distribution is different. The CI only applies to data drawn from the same distribution.
ML research still suffers from replication issues; such is the nature of the scientific incentive structure. However, these issues generally come in the form of poorly tuned baselines, buggy code, and claims with insufficient experimental/theoretical justification. Outside of some isolated cases, publication bias and cheating at hyperparameter tuning do not seem to be major factors.
----
[1] Statistically speaking, to compare two models on the same dataset, one does not care about the accuracy numbers but instead about the number of examples model A gets right that model B does not and vice versa; see McNemar's test.
The article proposes and falsifies a different hypothesis (that aphantasic subjects actually have a mind's eye but have the delusion that they don't) but the scientific argument is equally valid against this hypothesis. There is a difference in implicit behavior (in this case, priming during binocular rivalry), so the difference between aphantasic and non-aphantasic individuals seems to go deeper than metacognition.
The key claim in the article, that gradient descent could not discover physics from equations seems, like it is a statement about neural networks, not gradient descent. Given sufficient training data, a neural network can probably learn to model physics. I sympathize with the concern that it's very difficult to translate a neural network's knowledge into human concepts, but I see no reason to believe that optimizing the same system with an evolutionary algorithm would make this problem any easier. You could e.g. try to do program induction (which was supposed to be the future of AI many decades ago) instead of modeling the data directly, but choosing to perform program induction does not preclude the use of a neural network. Neural networks trained by gradient descent can generate ASTs (e.g. http://nlp.cs.berkeley.edu/pubs/Rabinovich-Stern-Klein_2017_...).
[Edited to remove reference to universal approximation; as comments point out, even if a neural network can approximate a function, it isn't guaranteed to be able to learn it. But I am reasonably confident that a neural network can learn Newton's second law.]
The authors used logistic regression to try to determine whether a subject will remember a word or not, which the classifier did better than chance, but still did pretty badly, with an AUC of 0.61. Then, when the classifier said the probability of remembering the stimulus is less than 0.5, they sent some current through some electrodes. The set of electrodes to stimulate and the current were selected in consultation with a neurologist and fixed at the start of the session. They found that stimulation in the lateral temporal cortex was associated with a significant (but just barely) increase in recall compared to no stimulation or stimulation outside of lateral temporal cortex. (But it's unclear whether this decision to look at effects in LTC vs. outside of LTC was made a priori. If it was not, and many comparisons conducted before arriving on this story, then the effect may not be statistically significant after adjusting for the comparisons.)
Beyond the question of whether the outcome was selected post hoc, the main problem with the study is that, unless I have missed it, there is no control to demonstrate that selecting the trials on which to stimulate using the classifier is better than stimulating on every trial. This control seems necessary to demonstrate that the linear classifier (which is apparently now "artificial intelligence") is in any way useful. Otherwise, this paper has little scientific value, short of possibly providing another data point regarding the effect of stimulation upon memory.
Link to paper: https://www.nature.com/articles/s41467-017-02753-0#Sec19
There's a difference between ideas for which there is not enough compute right now and ideas that are computationally intractable according to our knowledge of the physical universe.
With that said, I believe that CNNs are merely one approach to understanding images that, given enough data, appears to work quite well. It is quite possible that, by encoding a stronger prior regarding the world into the network architecture, you can accomplish the same goals more accurately with less data. The appeal of the capsules work is that the approach is substantially different from the CNNs that have been tweaked to recognize images over the last 5 years, but still appears to achieve good (and sometimes superior) performance on difficult tasks.
log((65/35)/(5/95))/(pi/sqrt(3)) = 1.9646484930140864
The justification for this procedure is that the judges are assumed to make decisions by dichotomizing a continuous variable with logistic distributed error (this one statistical justification for logistic regression; see https://en.wikipedia.org/wiki/Logistic_regression#Latent_var...). The mean difference in the continuous variable is given by the log odds ratio times some constant, and the standard deviation of the continuous variable is pi/sqrt(3) times the same constant. Because the logistic distribution resembles a normal distribution (see https://en.wikipedia.org/wiki/Logistic_distribution and figure in paper), the standardized mean difference given by this method will approximately equal the standardized mean difference of the latent continuous variable.