And even if the model was good at measuring teaching performance on average, that could just mean that out of every hundred teachers evaluated, two got higher alphas than they deserved and two got lower alphas.
And even if the model was good at measuring teaching performance on average, that could just mean that out of every hundred teachers evaluated, two got higher alphas than they deserved and two got lower alphas.
On the other hand, lets compare it to our current system of ignoring teacher quality. Assuming 25% of teachers are significantly above average and 25% are below (anyone have data on this?), that's a 50% error rate.
Let me just address your main statistical misconception, however: You can’t assume that after all the measurable factors in your statistical model are taken into account, the only remaining input to performance is the skill of the teacher.
This is not an assumption. The assumption is that after all the measurable factors are taken into account, the remainder are unbiased (i.e., have mean 0) or at least have bias smaller than ignoring all data.
Can you hypothesize an external factor which would, in a single measurement period (either a semester or year), significantly reduce the scores of 20-30 students of a single teacher while not affecting the scores of other demographically similar students? And do you really believe this occurs so often that it would make measurement worse than ignoring all data?
Yeah. Classes, like any small social group of people, have their own dynamics. While the teacher should be able to exercise some control and restraint, they'll obviously be less successful as teachers if that's where their energy is going. Whereas a well-behaved class, they can focus all their energy on teaching. If you adjust for the socioeconomic factors of each student individually, you can separate them out, but you can't separate out the overall dynamics of the group from the performance of the teacher.
This is precisely what VAM does. Each student gets a predicted grade and they are combined to get an overall class grade.
If things like well-behavedness are correlated with socioeconomic status, they will be incorporated into the predictor. Thus, a low SES class will have a lower predicted grade, and their poor behavior will be part of the reason why.
The only component of student behavior the predictor will not account for is whatever component of behavioral factors which are uncorrelated with SES, but are correlated with the distribution of students between classes. I.e., there will need to be a systematic reason why rich white kids who are more poorly behaved than the average rich white kid go to classroom 101 but not classroom 201.
(If they are uncorrelated with the distribution between classes, they will merely contribute to sampling error. See my response to jbooth illustrating how the law of large numbers makes errors like this a small problem. Feel free to plug in your own numbers.)
The root cause is that they share a classroom with other students who are misbehaving. To oversimplify wildly, take 4 demographically matched classes:
1) A class with some disruptive students , with a bad teacher : outcome - bad
2) A class with entirely generally well-behaved students, with a bad teacher : outcome - not so good
3) A class with some disruptive students, with a good teacher: outcome - not so good
4) A class with entirely generally well-behaved students, with a good teacher : outcome - good
How do you distinguish between classes 2 and 3?
1) Average over the set of classes taught by any given teacher, maybe drop best/worst (you know, stats 101 stuff). Simple assumptions: do teacher evaluations once per year, 4 classes per semester (so data from 8 classes to evaluate teacher). Assume average scores of 80, and disruptive students reduce scores by 30, bad teachers reduce scores only by 10 pts.
Assume disruptive students uncommon, unlucky teacher gets 1 class with disruptive students.
Case 2) scores [80,80,80,50,80,80,80,80], mean=76. Mean drop top/bottom = 80.
Case 3) scores [70, 70, 70, 70, 70, 70, 70, 70], mean=70. Mean drop top/bottom = 70.
Assume disruptive students are common, unlucky good teacher gets 4 classes with disruptive students, lucky bad teacher gets 3.
Case 2) Scores [80, 80, 80, 80, 50, 50, 50, 50], mean 65, MDTB=65.
Case 3) Scores [ 70, 70, 70, 70, 70, 40, 40, 40], mean 58.76, MDTB=60.
Assuming gaps of 5-10pts are statistically insignificant for one year, you occasionally are unable to distinguish between good and bad teachers in a single year. So every 2-3 years, you to a multi-year review. Now your sample size is up to 24 (from 8), and most likely both the good teacher and bad teacher have both had a lucky and unlucky class or two.
2) Include disciplinary problems in the predictor. Thus, the error is reduced from a class with some disruptive students (happens occasionally) to a class with some students who became disruptive for the first time ever (happens 1/12 as often).
And we haven't even got into the troublesome part of reducing a student's behaviour to a single number. Or the unfortunate way that useful correlates of underlying behaviour stop being useful correlates once you reward people for meeting them. (I'm sure there's a name for this phenomenon, but I have forgotten it)
I support the idea of collecting data. I obviously want to analyse it as rigorously as possible. But there are just too many complex interactions going on when you have 30 people in a room trying to learn for our statistics to produce reliable numbers. At least, that's my intuition. It occurs to me that if you actually collected the data, you could do an ANOVA, and have a reasonable stab at attributing the variation in outcomes to various factors, such as individual students, teachers, subjects, the class they're in, interactions between any combination of the above, and "other". You'll need a lot of data, mind. My guess is <10% of the variance would come down to the teacher alone. But this is just a guess - no doubt people have done this before. They've probably done it multiple times, with different answers depending on the different measures they used, whether they adjusted for socioeconomic factors etc. There's probably review articles summarizing those.
Look, I’ve worked in the search-engine biz for over five years, and while I myself haven’t gone beyond stats 101, some of my co-workers have studied this stuff at the grad-school level, and they are constantly arguing with one another about how to measure the “quality” of a search engine, and then, given that measurement, what particular kind of statistical model to use in order to predict, given a query and a bunch of potential results, which result has the highest “quality”. (Note that in the particular subfield of search that we are working with, spam pages and black-hat SEO are not really concerns.) And we (like Google and everyone else in this industry) have the luxury of trolling through millions of clicks’ worth of log files and we can hire semi-skilled labor to train our statistical models.
And compared with measuring the “quality” of a teacher in a classroom, measuring the “quality” of a search engine is trivially easy.