HNHacker News
TopNewBestAskShowJobs

simonster

2,428 karma · joined March 14, 2012

email: simon@simonster.com github: https://github.com/simonster
submissionscomments
simonster··on AI is making it easier to create more noise, when all I want is good search
It seems inevitable that LLMs will eventually work for search. They just don't work yet.
simonster··on Bloom's 2 sigma problem
The bad versions should be developed, but they should not be marketed as a replacement for traditional education.
simonster··on For Chat-Based AI, We Are All Once Again Tech Companies’ Guinea Pigs
In modern society, we prioritize the ease of bringing new products to market over regulation and testing. The only industry we prospectively regulate for safety is the pharmaceutical industry. Otherwise, regulation is retrospective and slow.

Look at the way we regulate harmful chemicals. After studies find that chemicals found in consumer products are harmful, manufacturers spend several years redesigning their products to avoid them. After manufacturers have voluntarily redesigned their products, the chemicals are banned. The manufacturers often replace the harmful chemicals with structurally similar chemicals that have not yet been shown to be harmful. When those similar chemicals are shown to be harmful, the cycle repeats. This is the history of phthalates, which are used to soften PVC plastic. The CPSC first considered banning one phthalate (DEHP) in children's toys the 1980s. Before that happened, all large manufacturers replaced it by another phthalate (DINP), which also turned out to be harmful. Both DEHP and DINP were finally banned by congressional action in 2008.

Technology is not a chemical, but anything that people will be spending significant time interacting with is likely to have some effect on their wellbeing, potentially small but potentially large, potentially positive but potentially negative. The type of harms these chatbots might cause are difficult to predict in advance. It could take a couple years to determine whether the current generation of chatbots are actually a net positive or negative for society. In that time, they may become too entrenched for us to do anything if they do turn out to be harmful. When Facebook first came out, I doubt many people foresaw the negative impact that it could have on teenagers. Today, many studies show that social media use has a negative impact on kids' wellbeing, but there is no going back to a pre-social media world.

Prospectively testing and regulating tech in the same way we regulate pharma seems insane. OTOH, we do want to make sure that world-changing tech actually changes the world for the better, and the only obvious way to do this is through some kind of regulation. Although regulation will undoubtedly slow growth, many of us would be happy to sacrifice some growth for improved wellbeing. At this point, everyone has seen the productivity-pay gap plot showing that most of the improvement in US productivity since 1979 has not translated into increases in real wages. Our current "growth at almost any cost" strategy benefits corporations much more than it benefits the average person.

simonster··on Unbundling Tools for Thought
A brief look at the first few paragraphs of Vannevar Bush's Wikipedia article (https://en.wikipedia.org/wiki/Vannevar_Bush) would clearly establish that the answer to your question is yes.
simonster··on Over twenty percent of cable TV bills are bogus fees, study says
The DOT Full Fare Advertising Rule requires that airfare prices include all taxes and fees. Perhaps the cable industry needs a similar rule.
simonster··on A visual proof that neural nets can approximate any function
The universal approximation theorem guarantees that a finite-width neural network that approximates the function to within some epsilon exists. But, regardless of the approximation method, there is no way to certify that a given approximation method is sufficient for an arbitrary continuous function given only a finite number of samples (i.e., without oracle knowledge of the underlying function), which is the typical situation where neural networks are applied. I can construct a continuous function that has an arbitrary (but non-infinite) number of peaks in an arbitrary interval. Thus, any method that approximates the function within some epsilon for all possible inputs within that arbitrary interval must encode an arbitrary amount of information. I can also ensure that whatever the number of samples is, it's not enough to properly approximate the function.
simonster··on A visual proof that neural nets can approximate any function
By this logic, (naive) matrix multiplication is not O(n^3) because it is a function of the precision required. The size of the neural network required to approximate a given function to within some epsilon does not change with the dataset size.
simonster··on A visual proof that neural nets can approximate any function
Yes, neural nets are successful is in large part because they are asymptotically more efficient than other models. Training time is O(n) with O(1) memory, and prediction time is O(1) with O(1) memory. Compare to e.g. kernel methods, which have nicer theory behind them, but kernel least squares is O(n^3) with O(n^2) memory to fit and O(n^2) with O(n^2) memory to predict. The coefficients are larger for neural nets, but if your data are big enough, the asymptotics win out.
simonster··on Yann LeCun, Geoffrey Hinton and Yoshua Bengio win Turing Award
Schmidhuber's take on the Wright brothers:

https://www.nature.com/articles/421689c

https://www.nature.com/articles/d41586-019-00491-5

simonster··on Adding layers to the middle of trained network without invalidating the weights
There's at least one existing paper about this idea (https://arxiv.org/abs/1511.05641). Also, it is possible to initialize a convolutional layer so that it passes through its input, but initializing the weights properly requires a little more work than calling tf.keras.initializers
simonster··on Google hit with €1.5B fine from EU over advertising
Yes. They are a significant enough expense that they are a separate line item on SEC filings (search for "European Commission fines" in https://www.sec.gov/Archives/edgar/data/1652044/000165204419...).
simonster··on Myths in Machine Learning Research
I don't think machine learning suffers from the same kind of p-value-driven replication crisis as other fields. It is true that people don't generally perform proper statistics to compare machine learning models [1], but ML research has two things going for it that other scientific fields do not. First, comparing machine learning models on the same test set corresponds to a within-subjects analysis, generally with tens of thousands of subjects, so the noise level is low. Second, because ML researchers don't generally perform hypothesis tests, they care solely about effect size and not about significance. If my model gets all the same examples right as the previous state-of-the-art, plus 10 more, then my model is statistically significantly better, but on a test set of 10,000 examples this corresponds to a 0.1% accuracy improvement, which is generally not big enough to publish. By not doing hypothesis tests, ML researchers actually tend to be more conservative than their p-value-driven counterparts in other fields.

In the Recht et al. study, the reason the new test accuracy is wildly outside of a binomial confidence interval around the original test set accuracy is that the distribution is different. The CI only applies to data drawn from the same distribution.

ML research still suffers from replication issues; such is the nature of the scientific incentive structure. However, these issues generally come in the form of poorly tuned baselines, buggy code, and claims with insufficient experimental/theoretical justification. Outside of some isolated cases, publication bias and cheating at hyperparameter tuning do not seem to be major factors.

----

[1] Statistically speaking, to compare two models on the same dataset, one does not care about the accuracy numbers but instead about the number of examples model A gets right that model B does not and vice versa; see McNemar's test.

simonster··on Myths in Machine Learning Research
#3 is actually wrong. The results of Recht et al. do not show that people are performing validation on the test set. If this were true, one would expect a poor correlation between accuracy on the original CIFAR-10 test set and the new test set, whereas the authors observe an extremely high correlation. The results actually indicate that attempting to follow the same dataset collection procedures as the creators of CIFAR-10 results in a dataset that is slightly harder than the original dataset (at least for models trained on the original dataset). The follow-up paper (http://people.csail.mit.edu/ludwigs/papers/imagenet.pdf) makes this point explicitly. The fact that the relative ordering of models is preserved on the new dataset suggests that the creators of the models didn't cheat, or at least didn't cheat enough to invalidate CIFAR-10 test set performance as an evaluation metric.
simonster··on The blind mind: No sensory visual imagery in aphantasia
> What if no-one had a mind's eye and Aphantasia is simply the lack of a delusion of a mind's eye.

The article proposes and falsifies a different hypothesis (that aphantasic subjects actually have a mind's eye but have the delusion that they don't) but the scientific argument is equally valid against this hypothesis. There is a difference in implicit behavior (in this case, priming during binocular rivalry), so the difference between aphantasic and non-aphantasic individuals seems to go deeper than metacognition.

simonster··on The blind mind: No sensory visual imagery in aphantasia
These sound like hypnagogic hallucinations. I think they're pretty common, although the precise experience varies by individual.
simonster··on Neural scene representation and rendering
It is possible that the problems are related—-it may be that, to achieve human-like generalization, neural nets need to learn in a human-like environment, instead of from a folder full of images. But time will tell.
simonster··on Neural scene representation and rendering
I don’t think this is a novel idea, but it is still a great topic for a PhD. While the results in this paper look impressive, my suspicion is that the system doesn’t generalize particularly well. (I suspect this from experience with similar, albeit simpler, ideas, as well as from looking at the datasets.) If you can make a system that generalizes to new environments and objects, or a system that works with real-world natural image/video data, that would be a tremendous accomplishment.
simonster··on The limitations of gradient descent as a principle of brain function
Yep, the minima are in the same locations. However, if the problem has multiple minima, then the parametrization can affect which minimum gradient descent actually reaches.
simonster··on The limitations of gradient descent as a principle of brain function
tl;dr ordinary gradient descent is sensitive to the parametrization of the problem but natural gradient is not. This is an important fact (and one that is fairly well-known within the ML community), but it is not totally clear to me why it should be particularly relevant to neuroscience.
simonster··on Empiricism and the limits of gradient descent
It's a fair point that the Universal Approximation Theorem does not guarantee that the weights can be learned. OTOH, the physical laws that the article states a neural network cannot discover are computable functions.
simonster··on Empiricism and the limits of gradient descent
There are a couple of factual errors here. First, the difference between backprop and evolution is smaller than the author indicates. The error signal used in modern backprop training is stochastic because it is computed on a minibatch (which is why it's called stochastic gradient descent). This stochasticity seems important to achieving good results. And the most popular evolutionary algorithm in the deep learning world is Evolution Strategies, which effectively approximates a gradient. Ordinary genetic algorithms are not gradient-based and have recently shown promise in limited domains, but can't compete with gradient-based algorithms for supervised learning.

The key claim in the article, that gradient descent could not discover physics from equations seems, like it is a statement about neural networks, not gradient descent. Given sufficient training data, a neural network can probably learn to model physics. I sympathize with the concern that it's very difficult to translate a neural network's knowledge into human concepts, but I see no reason to believe that optimizing the same system with an evolutionary algorithm would make this problem any easier. You could e.g. try to do program induction (which was supposed to be the future of AI many decades ago) instead of modeling the data directly, but choosing to perform program induction does not preclude the use of a neural network. Neural networks trained by gradient descent can generate ASTs (e.g. http://nlp.cs.berkeley.edu/pubs/Rabinovich-Stern-Klein_2017_...).

[Edited to remove reference to universal approximation; as comments point out, even if a neural network can approximate a function, it isn't guaranteed to be able to learn it. But I am reasonably confident that a neural network can learn Newton's second law.]

simonster··on Intel Delays Mass Production of 10nm CPUs to 2019
Given that TSMC just started high volume production of 7nm chips (https://www.anandtech.com/show/12677/tsmc-kicks-off-volume-p...), this sounds bad for Intel.
simonster··on Machine learning algorithms used to decode and enhance human memory
It's not totally clear to me whether there is a real story in ML/AI or neuroscience.

The authors used logistic regression to try to determine whether a subject will remember a word or not, which the classifier did better than chance, but still did pretty badly, with an AUC of 0.61. Then, when the classifier said the probability of remembering the stimulus is less than 0.5, they sent some current through some electrodes. The set of electrodes to stimulate and the current were selected in consultation with a neurologist and fixed at the start of the session. They found that stimulation in the lateral temporal cortex was associated with a significant (but just barely) increase in recall compared to no stimulation or stimulation outside of lateral temporal cortex. (But it's unclear whether this decision to look at effects in LTC vs. outside of LTC was made a priori. If it was not, and many comparisons conducted before arriving on this story, then the effect may not be statistically significant after adjusting for the comparisons.)

Beyond the question of whether the outcome was selected post hoc, the main problem with the study is that, unless I have missed it, there is no control to demonstrate that selecting the trials on which to stimulate using the classifier is better than stimulating on every trial. This control seems necessary to demonstrate that the linear classifier (which is apparently now "artificial intelligence") is in any way useful. Otherwise, this paper has little scientific value, short of possibly providing another data point regarding the effect of stimulation upon memory.

Link to paper: https://www.nature.com/articles/s41467-017-02753-0#Sec19

simonster··on Taking a Picture of a Supernova While Setting Up a New Camera
I'm somewhat amused that, although this article suggests that Buso's work was vital to the Nature paper, as does the first sentence of the paper itself, he is author 7 out of 21: https://www.nature.com/articles/nature25151
simonster··on Greedy, Brittle, Opaque, and Shallow: The Downsides to Deep Learning
Let's assume that evaluating a program takes one microsecond and one atom, and that we can parallelize the search across every atom in the observable universe. If the AGI program is 500 bits long, it will take about 10^57 years to find by brute-force search.

There's a difference between ideas for which there is not enough compute right now and ideas that are computationally intractable according to our knowledge of the physical universe.

simonster··on Nematode Neural System Uploaded to a Computer and Trained to Balance a Pole
It's a more accurate neuron model of some pretty weird neurons. In nearly all organisms, the vast majority of neurons fire all-or-nothing action potentials ("spikes"). C. elegans neurons do not.
simonster··on Understanding Hinton’s Capsule Networks, Part I: Intuition
Beyond the initial stages of the network, current SOTA CNNs use strided convolution in addition to (Inception, NASNet) or instead of (ResNet, DenseNet) max pooling. But my impression is that this has more to do with computational efficiency than anything else. Even with max pooling, you can maintain spatial information if you construct the preceding filters properly. But what's important in the example in the post is not the absolute locations of the parts of the face, but the spatial relationships among them, and this is actually something CNNs appear to be reasonably good at handling. CNNs achieve superhuman performance in identifying faces from natural images, so I doubt that a CNN would have trouble telling apart the faces shown in the article.

With that said, I believe that CNNs are merely one approach to understanding images that, given enough data, appears to work quite well. It is quite possible that, by encoding a stronger prior regarding the world into the network architecture, you can accomplish the same goals more accurately with less data. The appeal of the capsules work is that the approach is substantially different from the CNNs that have been tweaked to recognize images over the last 5 years, but still appears to achieve good (and sometimes superior) performance on difficult tasks.

simonster··on Why Does the Neocortex Have Columns? A Theory of Learning Structure of the World [pdf]
As a neuroscientist who just started doing ML research, I would call this paper cargo cult programming. If you cobble together a hodgepodge of ideas from neuroscience and build a network to accomplish some trivial task with no baseline to compare it to, I find it really difficult to take anything away from that. Ignoring the cortical column aspect, I'm not particularly convinced that Hawkins's model is a better approximation of biology than a typical deep neural network, just different (and likely far less capable, if you were to apply it to a challenging task). Why not start with a network that we know works and make a biologically-inspired change, and then see if that improves performance on a well-studied problem? If it does, then you have 1) an improvement on the previous network and 2) weak evidence that your idea of how the brain works may be right, if we assume that the brain is a highly optimized information processing device.
simonster··on Why I left Academia: Part I
If you spent 7 years of your life on a project only to haven it stolen out from underneath you by an unscrupulous "mentor", you're saying wouldn't get emotional about it? As far as I can tell, her story reveals exceptional resilience: She did, after all, get her Ph.D., despite having a normal emotional response to the revelation that people she trusted for a substantial proportion of her adult life were conspiring to screw her over.
simonster··on Impossibly Hungry Judges
For anyone who's curious, the procedure used is described in http://www.aliquote.org/pub/odds_meta.pdf. The standardized mean difference is given by:

  log((65/35)/(5/95))/(pi/sqrt(3)) = 1.9646484930140864
The justification for this procedure is that the judges are assumed to make decisions by dichotomizing a continuous variable with logistic distributed error (this one statistical justification for logistic regression; see https://en.wikipedia.org/wiki/Logistic_regression#Latent_var...). The mean difference in the continuous variable is given by the log odds ratio times some constant, and the standard deviation of the continuous variable is pi/sqrt(3) times the same constant. Because the logistic distribution resembles a normal distribution (see https://en.wikipedia.org/wiki/Logistic_distribution and figure in paper), the standardized mean difference given by this method will approximately equal the standardized mean difference of the latent continuous variable.
← PreviousPage 2 of 15Next →