Experiment: Paying for performance works (for teachers in india)
marginalrevolution.com
marginalrevolution.com
Maybe tests can be improved - but it is currently the only reliable way in which the performance of students can be measured.
A teacher motivated by incentives receives their motivation externally. Motivation in such a form is never sustainable I don't believe. Over time, I would expect the motivation teachers receive from incentives would decrease over time.
Though I must admit, I don't really buy your thesis that incentives only work short term. If they did, you'd expect entrepreneurs to be unmotivated beyond the short term, for example.
http://notebook.lausd.net/pls/ptl/docs/PAGE/CA_LAUSD/FLDR_OR...
So I suppose that a more careful cost-benefit analysis needs to be made.
So. You are proposing to throw teachers out after six years of teaching. Then, assuming every teacher stays the full six years, 1/3 of all teachers will be in that low-experience category that teaches so much less effectively. At present, according to the study you cited, the figure is more like 1/5. It seems that your proposal will result in 65% more pupils being taught by teachers who do not do a good job.
This is not the only reason why your suggestion seems to me unlikely to be a good one, but it seems a pretty compelling reason.
If we get rid of teachers after 6 years, what would that teacher do? How would this work exactly?
As for "what will teachers do" after 6 years, they can either continue to perform well or find a new job (just like a trader or salesman). The purpose of school is education, not jobs for liberal arts grads.
"they can either continue to perform well or find a new job" How to measure?
Yes, paying teachers more for higher test scores is a very good way of increasing test scores - this was found in california schools years ago. It had an opposite effect on students actual learning though.
That is assuming the unlikely proposition that preparing for the test doesn't directly prepare you for the conceptual portion of the test. I say unlikely because test makers have been trying to make such tests (eg the SAT) for decades, and test preparation companies (eg Kaplan) have been demonstrating that you can prepare for them after all.
You can't escape the circularity. If you're measuring performance with a test, then you can't really distinguish preparation for the test from actual performance. In this study they had much stronger results in their second year, than their first. Which strongly suggests that teachers did a better job of preparing for the test after teachers saw the tests that would be used.
Incidentally, bringing up the SAT is a red herring. The SAT was originally meant to measure g/intelligence/"aptitude". It did this by measuring a body of knowledge which is not explicitly taught in school, giving questions like wheel : car -> pick one of [leg: horse, egg : chicken, computer : TV]. Obviously, you can improve performance by explicitly teaching that knowledge (e.g., memorizing analogies).
A subject test does not suffer from this problem, since it is designed to measure subject knowledge rather than aptitude. You can improve performance on a multiplication test by memorizing multiplication tables, but so what? Multiplication performance is what you are trying to measure, it's not a proxy for something else.
Now whether they're incenting the right things (test scores) is debatable, but why is this newsworthy?
At the moment, it's effectively impossible to fire teachers after two or three years, no matter how poorly they do: see, for example, http://www.newyorker.com/reporting/2009/08/31/090831fa_fact_... .
I was mostly taking issue with the headline's (and article's) apparent surprise that offering money for higher X caused X to increase. Not sure why an experiment was needed to establish this.
Also, your rubber room article is horrifying, but I don't really see the relevance... except that the reward mechanism has been completely divorced from performance of any kind.
Sure there may be some motivational issues, but they would have to be systemic for this sort of ploy to work.
"At the end of two years of the program, students in incentive schools performed significantly better than those in comparison schools by 0.28 and 0.16 standard deviations (SD) in math and language tests respectively...."
That's 0.28 for math and 0.16 for language. I'd say that's pretty significant - especially for a large sample size. In this case presumably because of the sample size it's more than being about "power/tools/money".
Thus the results are not statistically significant aka with in normal random error etc. So when did 0.3 of not significant become significant.
From the wikipedia article on statistical significance (http://en.wikipedia.org/wiki/Statistical_significance): "Given a sufficiently large sample, extremely small and non-notable differences can be found to be statistically significant, and statistical significance says nothing about the practical significance of a difference."
This isn't even "extremely small". From the report: "The mean treatment effect of 0.22 SD is equal to 9 percentile points at the median of a normal distribution. We find a minimum average treatment effect of 0.1 SD at every percentile of baseline test scores, suggesting broad-based gains in test scores as a result of the incentive program."