And we haven't even got into the troublesome part of reducing a student's behaviour to a single number. Or the unfortunate way that useful correlates of underlying behaviour stop being useful correlates once you reward people for meeting them. (I'm sure there's a name for this phenomenon, but I have forgotten it)
I support the idea of collecting data. I obviously want to analyse it as rigorously as possible. But there are just too many complex interactions going on when you have 30 people in a room trying to learn for our statistics to produce reliable numbers. At least, that's my intuition. It occurs to me that if you actually collected the data, you could do an ANOVA, and have a reasonable stab at attributing the variation in outcomes to various factors, such as individual students, teachers, subjects, the class they're in, interactions between any combination of the above, and "other". You'll need a lot of data, mind. My guess is <10% of the variance would come down to the teacher alone. But this is just a guess - no doubt people have done this before. They've probably done it multiple times, with different answers depending on the different measures they used, whether they adjusted for socioeconomic factors etc. There's probably review articles summarizing those.