On the other hand, lets compare it to our current system of ignoring teacher quality. Assuming 25% of teachers are significantly above average and 25% are below (anyone have data on this?), that's a 50% error rate.
Let me just address your main statistical misconception, however: You can’t assume that after all the measurable factors in your statistical model are taken into account, the only remaining input to performance is the skill of the teacher.
This is not an assumption. The assumption is that after all the measurable factors are taken into account, the remainder are unbiased (i.e., have mean 0) or at least have bias smaller than ignoring all data.
Can you hypothesize an external factor which would, in a single measurement period (either a semester or year), significantly reduce the scores of 20-30 students of a single teacher while not affecting the scores of other demographically similar students? And do you really believe this occurs so often that it would make measurement worse than ignoring all data?