Failing 15% of the time is the best way to learn, say scientists
independent.co.uk
independent.co.uk
The Independent article is unfounded speculation about this applying to the way humans learn, without any discussion of whether the model is actually applicable. (Most things humans are trying to learn aren't binary classification tasks.)
I don't know how to separate out the difference between learning something through immersion and seeing the broader world view that nestles those ideas within them.
Is there a mistake in equation 3? shouldn't it read:
ER = integral (-inf, 0, p(h | ABS( Delta ), sigma ) dh ?
for positive Delta I agree, but with negative delta either the integral should run from 0 to +inf, or the absolute value of Delta should be taken?
Throwing up cards you can answer in 1 second is not wasting anyone’s time. It is more likely to be encouraging to people to remind people of what they know.
In my understanding, that's closer to optimal. Effortful retrieval is much more effective at strengthening future retrieval.
Wanikani does an amazing job at SRS. I wish Anki followed its model.
These are "leeches". They burn productive study and review time.
I have a theory that these items have lower adjacency to past experience or knowledge and it's difficult to form mnemonics or other connections. Or they're less novel and don't cause our brain to take interest. That's where all of my leeches lie -- in the realm of things I don't particularly care about.
A good leech management algorithm will back-burner unproductive items so you can focus on the rest of the concept population. There are different types of leeches too -- things you don't get during introduction, or things that you can commit to short term memory but won't stick for long. A good algorithm will identify all of them and block them.
Leech management is ok, but I think active positive reinforcement is also necessary. Remind the user of how much they know as well.
My hypothesis is that we model the system we're studying and simulate many 'attempts' for every real world attempt. I.e. we grow a low-fidelity, but much faster, model of the system in our brain that we can use to make medium-low confidence predictions about the real system many times for each time we test against the real system.
So when you say you fail 95% of the time, I'm saying each of those failures actually have 200 mini-successes embedded that you can still use to train your mental model.
Once burned, twice shy. And often that results in irrational aversion to huge classes of behaviors just because they appeared in the larger context of the failure of an endeavor as a whole, which I'd say is not a good way to learn from failures.
Sometimes people go into a situation confused, fail, and don't know how to interpret why they failed or what parts caused the failure. I think it is worth distinguishing that type of failure from the type you're talking about. Why? Often when you tell someone you don't know how to do something or you think you'll fail at something, they have your type of failure in mind and they encourage you to just try again.
It seems necessary to include a model of self to reach the kind of predictive ability necessary to learn from few examples.
If you failed to succeed in the expected amount of time/effort, it's a failure. Maybe you add some tolerance of going past the estimate before classifying as a failure, but it's still rooted in the expectation.
For example, if it took you ten years to pass kindergarten, you've failed.
If you got 100% right, you already know everything that was being tested.
If you got 50% right, you don't know if you are guessing or if you should be picking up on any features.
So you would expect that the rate would not be close. 50.1% would be similar to 50% for most intents and purposes. Similarly 99.99%.
So you might expect that the optimal learning rate would be close to 75%/25% in general. This would apply to humans too because it is a statement of the information you need to solve the problem, not a statement about the algo.
This paper finds it to be 85%/15% for a particular algorithm. Perhaps humans learn similarly, perhaps not. However, you might expect the optimal examples to be somewhere in the 65-85% range for any particular algorithm.
That's approximately the vale of the area under one tail of a normal distribution, from one standard deviation above the mean to infinity.
I'm not statistically mature enough to say whether it's just coincidence. For one thing, oodles of natural phenomena in no way follow the normal distribution.
I found other articles about this: https://eshapard.github.io/anki/target-an-80-90-percent-succ... https://vladsperspective.wordpress.com/2017/03/14/optimize-y...
Specifically in performance marketing spend, 15% of the budget is very often allocated to "new initiatives & new partners", with the thought process that it'll either allow to find a previously un-identified improvement, or it'll allow to learn what to avoid in the future on the 85% of spend.
How is this different from information theory?
That's a different sense from learning as discovery, or at least, learning as search. In searching a graph of possible hypotheses, yes, it is a better rule to look for opportunities to halve the search space.
No, AFAIK it literally never existed. That was, as I recall, a popular misinterpretation of what was itself an unwarranted generalization made by Malcolm Gladwell based on a paper with much more limited scope and conclusions.
> This one sounds very similar in trying to quantify a very chaotic and qualitative process.
The actual direct conclusion—that this error rate is optimal for a variety of machine-learning processes—does not seem to ha r the problem you describe. The suggestion in the paper that this extends to “biologically plausible” neural networks that may model animal learning also does not seem problematic in the way you describe. The news article’s claim that this is a finding of a sweet spot for human learning is, while it is a possibility suggested by the paper, simply unwarranted as a conclusion.
It's certainly plausible that a quantifiable sweet spot of this type exists for some kinds of human learning at the optimization of effectiveness in a curriculum that can be dynamically scaled to individual learners could effectively be guided by it, but there is not a strong reason without actually testing in concrete human learning scenarios to believe that the particular number here is a guide to that.
> As though getting 100% on my calculus quizzes indicated that I wasn't learning.
It doesn't say that. It says you would probably be learning faster if the test would be more difficult.
Same category error with the second.
If you get 100% on every pretest, you're probably not learning anything new from being told the answers.
The subject is so sensitive that I often get banned from forums for saying what I know to be true.
If you don't fuck it up, people will think your food is at least 'good'. Making it 'great' is the hard part.
Well, you get the idea.
OTOH, I graded for a Physics prof where 50/100 on his test was about average.
However, you do want to be a "B" student when learning, just not in scoring.
Eventually (at a prestigious university) I found I could no longer "coast", and studying required real work. It wasn't easy to come to grips with that reality, and I wish I'd learned earlier.