Why Your Boss Is Wrong About You
nytimes.com
nytimes.com
I was once a manager at a medium-sized consulting company. I had an employee we'll call "Roy." He was always involved in the critical parts of large, profitable projects. He generally got 3.5 to 4 out of 5 in his peer reviews. There was another guy, "Jim." Jim was mostly on smaller, simpler systems and maintenance work. He generally got similar scores in his peer reviews. For their first performance review, I gave Roy a 4.2 and Jim a 3.8. Roy got a 7% raise and Jim got a 5% raise. Later on, I was the lead dev on a project with Roy and Jim as my team. Once the project got rolling, getting Roy to actually produce code was like pulling teeth. I talked to the other leads and found out that Roy talked a good game in front of clients but required a baby-sitter to actually get anything done. He was always on the critical components not because he was a great developer, but because the lead or architect was already paying extra attention to the critical components and could more easily manage the babysitting. Jim, of course, got all his work done on time with minimal fuss and even stayed late to finish some of Roy's work.
I went back and looked at those peer reviews... Roy's lead dev had given him a review that averaged to 3.5. The soft skills were mostly 4's, the technical skills were mostly 3's and 4's. (He was an OK developer when you actually got him to work.) My only clue would have been a 2 in "Works Independently." Jim didn't have other devs on his projects so he had Project Managers giving him 3's and 4's for soft skills and 4's for technical skills.
A poorly designed performance review was actually worse than no performance review.
And none of this came through in the peer reviews, beyond the 2 in Works Independently? I think the review system here broke down in the peer area - no one who was working with him was communicating the vital info on this one, beyond just numeric scores.
Exactly. The review process wasn't actually measuring anything useful, it was just a list of traits that upper management thought were important.
This story also reflects how clueless I was as a manager, that I was only talking to my people to gather information.
Move him over in sales, he will properly do very well there since he knows the technical part very well but has human skills as well.
Flip side: if you are a small enough group that know each other very well, performance reviews can be a good opportunity to raise awkward subjects.
It's a tool - it can be used or misused.
Of course, the author completely overlooks any solutions besides his own pet method, suggesting "taxpayers can't ask for more than that".
They certainly can. For teachers (the situation he uses as a lead in and fade out), they can use Value Added Modeling. In fields with more subjective performance, 360 reviews (reviews by your peers, underlings, overlings and self, discarding any singleton viewpoints) are very effective.
But I guess it's more fun to push your own toy method (and hopefully get hired as a consultant) than propose serious solutions.
I mean, you've described your own experience as getting out of the hidebound checklist mediocrity of university life to the dynamic 1-man-shop life. Similarly I pick startups to work at instead of big lumbering corporations.
I was just explaining where the unions tend to be coming from when you sit across the table from them. I think the seniority rules are dumb, particularly last-in-first-out which is probably the worst thing about the current state of your typical municipal agreement. But I think you're mischaracterizing the opposition to standardized tests.. some % of teachers are lazy and don't want to teach the material but the strongest objections tend to come from some of the best teachers who resent having to "teach to the test" instead of using their judgment about the best way to teach/evaluate their students.
How are you judging the urban teacher against the suburban one? Is the kid who is getting abused at home ever going to match the one with two great parents? What percentage of your students eat breakfast? What percentage of your students have English as their first language?
What percentage of the individual class are troublemakers? What tools are teachers allowed to use to subdue them? Principals differ on this.
Are there bad teachers? Certainly! Do the metrics identify which teachers are the bad ones? Often, no.
From my first post: For teacher..., they can use Value Added Modeling.
Come up with a statistical predictor of performance, based on observable quantities such as income, ESL, race, free lunch, and previous year's performance.
Teacher alpha = actual performance - predicted performance.
All the factors you described are either highly correlated with easily available quantities (low income), or they are directly measurable. Thus, the predictor performance is likely to include them.
I'd be curious to see if you could even hypothesize a systematic error (as opposed to sampling error) with a large effect that is not highly correlated with with easily observable base quantities (income, demographics, prior performance).
But that statistical predictor of performance will always be horribly flawed. Applying a rigid set of a dozen statistics to human beings always misses the mark in a ton of cases. The question is whether it's better than the current system -- personally I'd want to see several iterations experimented with before I'd feel confident at all that it wasn't making things worse. "First, do no harm" and all that.
One huge problem a lot of good urban teachers have with testing is that the tests as written really don't "speak to" poor urban and minority students. The teachers complain that they spend more time teaching the students how to take the test and less time teaching them the material in a way that's relevant to their lives. I haven't ever taught, so I'm just repeating this, but I can see where they're coming from.
http://en.wikipedia.org/wiki/Law_of_large_numbers
The question is whether it's better than the current system...
Without objective measurements, how could we ever answer this question?
I mean, just to start with:
* Are our tests measuring the right things? * Are we measuring the right things about the kids? * Are we capturing all of the data for the things we do measure about the kids?
Those are huge unanswered questions to just be like "oh yeah just average the results for the income quintile, bang, done, we know who the good teachers are".
If a student's being sexually abused at home, and test scores dramatically decline year-over-year, is the teacher a worse teacher for it? Just one edge case. Throw in 4 more and you've got a significant fraction of the scores that the teacher will be evaluated on. These are students' futures and teachers' jobs we're talking about, remember.. I think general teacher resistance to having some sort of oversimplified system shoved on them from above is a feature -- that system needs to be good enough to buy them in before being implemented.
Suppose 1 in 100 children is sexually abused for the first time in any given semester [1], and a teacher has 4 classes of 30 students each. Then in 2.5% of classes, a teacher will have 4 sexually abused children (vs an average of 1.2). When this occurs, the percentage of children failing will increase by 2.3%.
So lets say on average, a teacher is expected to have a failure rate of 20%. 2.5% of the time (roughly once in their career) they get unlucky and have a failure rate of 22.3% instead of 20%. This will occur 2 semesters in a row about one time in 1500, i.e. it will happen to 1 in 50 teachers (assuming a 30 year teaching career).
Alternative numbers, to show that I'm not using cherry picked numbers: P(abuse)=0.001, P(4 abuses in class of 120) = 2.56e-6, slightly greater impact (failure rate goes up to 23.3% over 20%). If sexual abuse happens at a higher rate (e.g. 2%), then we get into the rather implausible situation that 48% of children are sexually abused, resulting in them failing school.
[1] I don't know the relative proportions, but this seems like a high number. Among other things, it would imply that at least 24% of children are sexually abused, resulting in them failing school. We are also assuming the child has a good upbringing and good grades until the sexual abuse, which I imagine is not the most common situation. The assumption that it is the first instance of sexual abuse is important, because if it occurred in the prior year, it would likely affect their prior year grades and hence their current predicted grades.
Again, I do think that some sort of unified performance measurement is important. But it's really easy to screw something like that up. Teachers spend between 5-35 hours a week with a student depending on schedule and we're summing up that interaction with information that fits on half of the back of a business card? That doesn't tingle your spider sense at all? "Warning, lots of information loss here"?
When you base performance on a metric, then you re-orient everything around the metric. It is very very very important to get that metric right and a right answer probably doesn't fit in an HN post. I mean, nobody thinks engineers should be measured by kloc, let's not oversimplify for teachers either.
And even if the model was good at measuring teaching performance on average, that could just mean that out of every hundred teachers evaluated, two got higher alphas than they deserved and two got lower alphas.
On the other hand, lets compare it to our current system of ignoring teacher quality. Assuming 25% of teachers are significantly above average and 25% are below (anyone have data on this?), that's a 50% error rate.
Let me just address your main statistical misconception, however: You can’t assume that after all the measurable factors in your statistical model are taken into account, the only remaining input to performance is the skill of the teacher.
This is not an assumption. The assumption is that after all the measurable factors are taken into account, the remainder are unbiased (i.e., have mean 0) or at least have bias smaller than ignoring all data.
Can you hypothesize an external factor which would, in a single measurement period (either a semester or year), significantly reduce the scores of 20-30 students of a single teacher while not affecting the scores of other demographically similar students? And do you really believe this occurs so often that it would make measurement worse than ignoring all data?
Yeah. Classes, like any small social group of people, have their own dynamics. While the teacher should be able to exercise some control and restraint, they'll obviously be less successful as teachers if that's where their energy is going. Whereas a well-behaved class, they can focus all their energy on teaching. If you adjust for the socioeconomic factors of each student individually, you can separate them out, but you can't separate out the overall dynamics of the group from the performance of the teacher.
This is precisely what VAM does. Each student gets a predicted grade and they are combined to get an overall class grade.
If things like well-behavedness are correlated with socioeconomic status, they will be incorporated into the predictor. Thus, a low SES class will have a lower predicted grade, and their poor behavior will be part of the reason why.
The only component of student behavior the predictor will not account for is whatever component of behavioral factors which are uncorrelated with SES, but are correlated with the distribution of students between classes. I.e., there will need to be a systematic reason why rich white kids who are more poorly behaved than the average rich white kid go to classroom 101 but not classroom 201.
(If they are uncorrelated with the distribution between classes, they will merely contribute to sampling error. See my response to jbooth illustrating how the law of large numbers makes errors like this a small problem. Feel free to plug in your own numbers.)
The root cause is that they share a classroom with other students who are misbehaving. To oversimplify wildly, take 4 demographically matched classes:
1) A class with some disruptive students , with a bad teacher : outcome - bad
2) A class with entirely generally well-behaved students, with a bad teacher : outcome - not so good
3) A class with some disruptive students, with a good teacher: outcome - not so good
4) A class with entirely generally well-behaved students, with a good teacher : outcome - good
How do you distinguish between classes 2 and 3?
1) Average over the set of classes taught by any given teacher, maybe drop best/worst (you know, stats 101 stuff). Simple assumptions: do teacher evaluations once per year, 4 classes per semester (so data from 8 classes to evaluate teacher). Assume average scores of 80, and disruptive students reduce scores by 30, bad teachers reduce scores only by 10 pts.
Assume disruptive students uncommon, unlucky teacher gets 1 class with disruptive students.
Case 2) scores [80,80,80,50,80,80,80,80], mean=76. Mean drop top/bottom = 80.
Case 3) scores [70, 70, 70, 70, 70, 70, 70, 70], mean=70. Mean drop top/bottom = 70.
Assume disruptive students are common, unlucky good teacher gets 4 classes with disruptive students, lucky bad teacher gets 3.
Case 2) Scores [80, 80, 80, 80, 50, 50, 50, 50], mean 65, MDTB=65.
Case 3) Scores [ 70, 70, 70, 70, 70, 40, 40, 40], mean 58.76, MDTB=60.
Assuming gaps of 5-10pts are statistically insignificant for one year, you occasionally are unable to distinguish between good and bad teachers in a single year. So every 2-3 years, you to a multi-year review. Now your sample size is up to 24 (from 8), and most likely both the good teacher and bad teacher have both had a lucky and unlucky class or two.
2) Include disciplinary problems in the predictor. Thus, the error is reduced from a class with some disruptive students (happens occasionally) to a class with some students who became disruptive for the first time ever (happens 1/12 as often).
And we haven't even got into the troublesome part of reducing a student's behaviour to a single number. Or the unfortunate way that useful correlates of underlying behaviour stop being useful correlates once you reward people for meeting them. (I'm sure there's a name for this phenomenon, but I have forgotten it)
I support the idea of collecting data. I obviously want to analyse it as rigorously as possible. But there are just too many complex interactions going on when you have 30 people in a room trying to learn for our statistics to produce reliable numbers. At least, that's my intuition. It occurs to me that if you actually collected the data, you could do an ANOVA, and have a reasonable stab at attributing the variation in outcomes to various factors, such as individual students, teachers, subjects, the class they're in, interactions between any combination of the above, and "other". You'll need a lot of data, mind. My guess is <10% of the variance would come down to the teacher alone. But this is just a guess - no doubt people have done this before. They've probably done it multiple times, with different answers depending on the different measures they used, whether they adjusted for socioeconomic factors etc. There's probably review articles summarizing those.
Look, I’ve worked in the search-engine biz for over five years, and while I myself haven’t gone beyond stats 101, some of my co-workers have studied this stuff at the grad-school level, and they are constantly arguing with one another about how to measure the “quality” of a search engine, and then, given that measurement, what particular kind of statistical model to use in order to predict, given a query and a bunch of potential results, which result has the highest “quality”. (Note that in the particular subfield of search that we are working with, spam pages and black-hat SEO are not really concerns.) And we (like Google and everyone else in this industry) have the luxury of trolling through millions of clicks’ worth of log files and we can hire semi-skilled labor to train our statistical models.
And compared with measuring the “quality” of a teacher in a classroom, measuring the “quality” of a search engine is trivially easy.
I'm sorry, but a subjective review by a principal definitely has problems (personal biases), but it is a HELL of a lot better than judging teachers solely on how many years experience they have and how many meaningless masters degrees they have (French literature).
If you care about such things as getting a good performance review you should know that it measures likability far more than productivity.
If one of the KPIs on the CEO's dashboard is hits to the website, you should drive that up, if it's acquisitions then drive that up.
Performance reviews and KPIs are highly instructional as to what you should be optimizing for. It's the company telling you what they value most. If you find yourself in disagreement with the things being measured it's probably an indication that the value systems you use and your company uses are out of alignment. It's usually much easier to find a company that shares your values than change your companies' values to suit your own.