Here are two more or less obvious ways.
1) Average over the set of classes taught by any given teacher, maybe drop best/worst (you know, stats 101 stuff). Simple assumptions: do teacher evaluations once per year, 4 classes per semester (so data from 8 classes to evaluate teacher). Assume average scores of 80, and disruptive students reduce scores by 30, bad teachers reduce scores only by 10 pts.
Assume disruptive students uncommon, unlucky teacher gets 1 class with disruptive students.
Case 2) scores [80,80,80,50,80,80,80,80], mean=76. Mean drop top/bottom = 80.
Case 3) scores [70, 70, 70, 70, 70, 70, 70, 70], mean=70. Mean drop top/bottom = 70.
Assume disruptive students are common, unlucky good teacher gets 4 classes with disruptive students, lucky bad teacher gets 3.
Case 2) Scores [80, 80, 80, 80, 50, 50, 50, 50], mean 65, MDTB=65.
Case 3) Scores [ 70, 70, 70, 70, 70, 40, 40, 40], mean 58.76, MDTB=60.
Assuming gaps of 5-10pts are statistically insignificant for one year, you occasionally are unable to distinguish between good and bad teachers in a single year. So every 2-3 years, you to a multi-year review. Now your sample size is up to 24 (from 8), and most likely both the good teacher and bad teacher have both had a lucky and unlucky class or two.
2) Include disciplinary problems in the predictor. Thus, the error is reduced from a class with some disruptive students (happens occasionally) to a class with some students who became disruptive for the first time ever (happens 1/12 as often).