What the Dunning-Kruger effect is and isn’t (2010)
talyarkoni.org
talyarkoni.org
Now if people were perfectly accurate those lines would be drawn on top of each other.
But because either people underestimate their differences or correctly guess their true ability which the test doesn't full capture. Let's make the perceived ability line flatter.
People tend to overestimate their own ability, so we should move the perceived ability line up which leads to an intersection on the right. Causing us to find that the best performers are the most accurate, explaining DKE.
But the harder we make a task the more we push that perceived ability line down. And if we make a task sufficiently hard then the worst performers are the most accurate judge of their own ability showing the opposite of DKE, due to the intersection being pushed left.
So really DKE is an artifact.
Quadrant II of that graph is Dunning-Kruger. Quadrant IV is imposter syndrome. Quadrants I and III aren’t talked about but I suspect most people, most of the time, end up in these quadrants if you ask them about their ability to repair helicopter engines or perform open-heart surgery.
> But the harder we make a task the more we push that perceived ability line down.
That's only true if people misperceive how hard the task is across the population. Or if the measurement is trivialized because the distribution of actual performance is ovwrwhelemedr by noise.
It seems pretty intuitive why people would use task difficulty as a heuristic for relative performance.
One thing that I think often gets skipped over in discussions on Imposter Syndrome is that imposters DO exist (that’s kind of why we use interviews when hiring people).
The source of miscalibration is from the estimation of the subject area and the assessment of how much of that has been learned. I suspect that the error in the estimation of the subject space is the much larger source of error.
Exactly, it's like people who think the company CEO is a useless idiot because they don't see him writing React code or whatever their limited definition of "adding value" is.
Another example I see on HN - whenever a hiring thread comes up, people chime in how stupid whiteboard interviews are, and surely the poster would be a super-star at a FAANG if only they weren't gate-kept by these dumb interviews. Isn't it more likely that these super-successful companies use these interviews because they end up hiring the people they want to hire, and it's you, the poster, lacking a clue as to what that looks like?
Basically the heuristic is, if something/someone is successful and it seems stupid to you, start by assuming it's you who likely has something to learn.
I’m not sure it’s that these companies only hire the people they want to hire. It’s probably that highly successful companies know their interview methods don’t work 100% of the time, but they’re willing to tolerate the false negatives of not hiring a few good people if it means all those they do hire meet some minimum standards. And their current system is the best they know how to do.
We arrogantly post on Hacker News: “Oh look, this interview method sucks! It will miss out on this great candidate!”. Ok then, make your own trillion dollar company with a better interview method that filters out the hordes of bullshit artists while simultaneously never missing out on the best people. Filtering out the bullshit artists is more important than hiring every single productive person. In fact, if you hire too many bullshit artists, all the good people will get fed up and leave anyway.
Basically the argument is that it is much better to miss out on a good person than hire a mediocre/bad person.
As someone in the second bucket, the process feels rather random (eg. passed the Facebook L6 interview but failed the Google L4 one somehow, passed the onsite at half a dozen other companies), which is mostly because it is. Maybe some interviewer really wanted to see a topological sort algorithm implemented in 20 minutes and there's not enough time in the interview to derive it from first principles, or maybe an interviewer felt grumpy that day, and then that's it, better luck next time. Combined with the policy of not telling candidates what to improve on, it's a pretty frustrating experience.
Yes, there are reasons for the company to behave that way, but it's still unpleasant to be on the receiving end of it.
What's most illuminating about those kinds of comments is they rarely even try to view thing from the hiring company's perspective. Having been in that boat, you'll find:
1. The cost of a mishire is huge. And, even worse than hiring someone who is downright awful is hiring someone who is just plain mediocre, or slow. I mean, I've hired diligent, hard-working, otherwise smart people who could only get stuff out the door at about a third the rate of other devs. Having a performance discussion with those folks can be extremely difficult, because oftentimes there is little they can realistically do to go faster.
2. Experience can have little correlation with proficiency. I've interviewed people who were "senior engineers" and "architects" at major corporation who couldn't code FizBuzz, moreover they could barely code a syntactically correct function in their chosen language.
3. Many companies want to diversify their employee base by going outside just friends/referrals and well known college programs, but to "take a chance" on other folks means they need to get a strong signal during the interview process.
I'm not saying there aren't problems in hiring, but given how there is a giant economic incentive to make hiring as accurate and efficient as possible it's at least valuable to attempt to understand why a company might interview the way it does.
What does an off-site sample assignment tell the company? That you can write code when given a spec. That's great, but that's insufficient for what they are hiring for.
What it doesn't tell them:
- Can you elicit clarity on a vaguely stated problem? These are conversations that top engineers have with their business counterparts and other technologists constantly.
- Can you iterate with another person on a solution. The whiteboarding exercise is a dialogue, and while contrived, problems are often solved in joint manner like that.
- How do you do under pressure? Something is going wrong and it's affecting millions of users. Can you carry your weight when the team is fire fighting?
Obviously the whiteboard isn't perfectly correlated with the above, but it's as good an indicator as I can think of. Saying "just let me write something off-line" means people are oblivious to these other key attributes. FAANG devs aren't earning 300K a year for just being able to write code on a spec.
Then, hiring doesn't have to be this grandiose, talk to 20 people over 2 days in 20 30-minute session, but actually the VP has a final call, sort of decisions.
That works aside for the fact that it's horrible for everyone:
- The employee: they presumably quit another job to take this one and now they are out on their ass because the employer didn't bother to test for proper fit.
- The team: they invest in training the new person, get to like them, only to have them fired. Meanwhile they are short a person as hiring effort has to be restarted so they lose a ton of time understaffed.
- The company: dealing with all this terrible churn, stopping and starting the hiring effort, is there still severance in this world? How does unemployment work?
This is just the craziest proposal I've ever heard. It's like saying "dating is too onerous, people should just marry randos and not hesitate to get divorced if it turns out the other person brushes their teeth in an annoying way."
Still, there are practicalities that can make "hire fast, fire fast" difficult in the real world. First, many countries outside the US have much more onerous requirements to fire someone. Even in the US, it's quite trivial for any employee to take legal action if they think they've been let go unfairly, so most HR departments will require many months of detailed documentation (e.g. email conversation, performance reviews, PIPs) before someone can be let go.
My point is that there's no reason to somehow induce your company to have more bad employees by watering down interviewing. Literally no upside.
You're saying I am saying we should randomly give jobs. I am not. I am saying it should be, at max, a 1 hour phone call. Background checks can verify past employment and job titles. Just talk to them about work matters. Ask a fizzbuzz if you must. No additional data will meaningfully tell you much more anyways.
If a few more hours of interviews can avoid a bad hire, I think almost all experienced managers would opt for the extra interviews.
Can it, though?
I'm sure it can increase the manager subjective feeling of confidence, but does it verifiably, objectively produce better results?
The last position I hired for had 700 applicants. Paper filtered down to 75. Interview screened down to 30. Detailed interviews with 5.
How would you propose this work? 30 people seemed like a possibly good fit after 20 minutes. Should we randomly hire one? Should we hire all and have them fight it out madmax thunder dome style?
I think if the position has low startup time and is assembly line work or something we could hire all 30 and fire the poor performers. Outside of tech support, I’m not aware of any positions like this.
I’ve worked with managers who thought programmers were like this and would hire 10 or 20 at a time to see who stuck. It was very frustrating as an existing employee because they all had to be trained, reviewed, etc. After a year of arguing, I (and every other senior, over about two years) left that company.
If the additional interviewing time doesn't produce real improvement in outcomes over random hiring, then yes, don't waste the interviewees’ time and time the form is paying you for doing it.
And I think there’s a deontological vs. utilitarian aspect as I think if applicants learned I was randomly picking 1 of 30 pretty goods it would discourage high quality applicants.
For me, I think the additional 100 hours interviewing helps and is worth it.
I’d be interested in hearing of hiring managers and orgs who just hire randomly from minimally qualified applicants. I hear from lots of applicants who claim they are minimally qualified that this is a successful strategy. But, naturally, they are a bit biased and not very useful for deciding how to hire people.
My preexisting mental model is that from the 700 applicants, 30 could have filled the position adequately, with perhaps a ~5% gain across them. Each additional step in the interview process yields increasingly diminishing returns but is still worth it because ... you can choose to be picky in this market. The risks associated with a mishire are also presumably greatly reduced with each diminishingly selective step. Or is it that there genuinely exists only 2-3 people in that pool who are a good fit for the role, and you have little choice but to engage in kissing a whole load of frogs to find your prince?
We only extended an offer to one person and maybe a second if the first declined. If both of those two declined then we wouldn’t have offered to the other 3 detail interviewees.
I’d certainly like a better way that uses less time. But it does seem that I need to kiss a lot of frogs.
Also, the goal isn’t to get an “adequate” person but to get a really good or great person. I think that a good person can do 5x more than an ok person, maybe even more. So for this position, I’d rather keep looking than just get a body.
Again, if I just needed assembly line workers or warm bodies that could bill in a consultant sweatshop, that’s a different story. But I don’t want to be in such a hiring position.
This is really surprising to me, and I'm not sure how this can be the case. I would imagine that this would be true of a high level position at the top of their game, able to define their own schedules & goals. For a typical worker with clearly defined short & long term objectives defined by the organization, how does one become 5X more productive? Clearly my heuristics about the tech industry are way out of step with reality.
However, I could see something like this happening in my own field: academia. As a bioinformatics postdoc, the bread & butter work most of my colleagues do is routine, so they can hit their goals fairly predictably. My project, in comparison, is mostly ad-hoc, struggling to parse a novel dataset right on the edge of what is technically possible. My productivity is in the toilet, and I can imagine a different researcher in my position being 5-10x more productive. This is not how I imagined the tech industry working, though.
I remember a McKinsey study from the 80s or 90s that talked about 10x productivity but I can’t find it so maybe I’m misremembering. I don’t think people who actually try to measure this succeed well. I run away from anyone trying to precisely measure programmer productivity because it usually means some point haired boss will try to optimize on lines of code or something stupid.
So for me, it’s a bit of a hunch but a strong hunch. I’ve worked in shops of 100 devs where a single dev wrote the entire authentication stack that teams couldn’t handle. And I have lots more stories like this.
Nowadays I do “strategy” and I’d say the multiple is more than 10x in that there’s usually some magic mix to a good strategy that can’t be accomplished with giant committees and tons of hours. When it comes to creative tasks, I think the productivity leap from “ok” to “good” and from “good” to “super” is really high. Not every job is like this, but I think these are the most fun and so try to go towards them and away from commodity work.
I have limited experience in academia but have seen authors quickly crank out really useful papers where teams have been working for months. I suspect that’s more about just having good alignment of capability and need instead of some magic or intrinsic productivity power.
Personally my productivity is pretty sucky so I’m in a situation where I don’t meet my criteria but am lucky enough to get to work on cool stuff.
Do human resources departments actually evaluate their hiring criteria and the tests they use in experimental designs, e.g. by hiring two sufficiently large groups A and B based on different interview methods and checking a year later how the groups performed relative to each other?
Bear in mind that there are plenty of people who believe that they're good at doing something just because they have been doing it for a long time. Without an independent method for confirming such claims they are practically useless.
FAANGS evolve their process over time too.
One clear example - at FAANGS even if the manager likes you, it doesn't matter. You have to be evaluated by a committee of people who're indifferent to your manager, aren't under the gun to fill a slot and are less likely to be swayed by a one-time positive conversation.
Do you think this process came out of nowhere? That's just one example of them settling on something that worked, I am sure having previously tried other less structured approaches and tracking the results.
First, candidates who complain about the whiteboard interviews uniformly complain about it "not being a good representation of how I can write code" which betrays a lack of understanding of the other skills besides writing code that these firms care about. People who understand more can connect the interview type to the other relevant attributes, so it's already clear who's more right.
Second, yes, companies do a shitload of triangulation on recruiting performance including looking many years out.
No it's more likely that the people staffing FAANG companies absolutely suck at interviewing and vetting.
That's obviously a part of it, but if that was all of it, you would expect accuracy of assessment of relative ability to consistently get better with relative ability.
But, in fact, the result is basically “everyone seems themselves as closer to the 70th percentile than they actually are”.
Everyone focusses on the low end of DKE to use it as a low-brow dismissal, but the fact that people at the top of the distribution tend to underestimate their relative ability is also interesting.
> imposters DO exist (that’s kind of why we use interviews when hiring people).
Yes, the presence of imposters in hiring and hiring-policy-setting positions is why we use assessment tools that don't work very well to evaluate candidates.
One of the most memorable examples of this for me was early on doing calculus in high school. There was a time where I was like "ok, yeah, I understand all the maths stuff now," then I hit that first peak and suddenly realized "oh wow, I know only a tiny, tiny fraction of what there is to know!" When I later learned about Dunning-Kruger it resonated because it described what I had experienced years earlier and had being thinking about since.
And then, about two years after that, it gradually sunk in that I'd been missing everything about what the music was really about. I hadn't been able to pick out what made good playing good, so I couldn't tell that my playing wasn't. I had been focusing on a bunch of superficial stuff and missing the heart of the thing.
And now 20+ years on, I am vastly better at it than I was when I thought I was great. But I know there are loads of people much better. And I'm constantly worried there's some other epiphany waiting to happen when I will realize I'm still missing something absolutely key...
2015 https://news.ycombinator.com/item?id=10145480
Discussed at the time: https://news.ycombinator.com/item?id=1498136
Related from yesterday: https://news.ycombinator.com/item?id=25546787
A case I witnessed was a VP of Engineering who realized (after many years at that position, performing well) that he is more an engineer than a manager. His results were good but he did not like the work he was doing because his true talents were under-utilized.
He wanted to move to a position where he would have a more "technical" role, on the official "technical ladder" but it was very difficult for him to make his point. He liked the company very much though (in Europe).
I did not know what to really tell him at the time, I wish I had some good arguments for him to help his case.
One of the key things is that he did not want to change the company - he loved it, he loved what they did; he just took the wrong turn at some point.
For example an expert JavaScript developer can quickly produce an MVP with an original code solution. A less competent (confident) developer might also be able to produce a similar MVP but only in certain contexts, such as requiring a particular framework. The less competent developer may refer to themselves as an expert developer in React, or whatever other framework, but in doing so redefines the problem to fit a more narrow context not immediately aligned to the product.
But it is useful to think about this for our personal growth. There are many psychological traps and challenges that can slow down progress or even halt it.
Related:
https://daedtech.com/how-developers-stop-learning-rise-of-th...
Every single time I see the Duning Kruger effect quoted it's used to dunk on some person perceived as incompetent (usually a highly-placed executive) as a way to say "No, look, all people in that position think they're geniuses but they're actually morons!"
But... well, even if there weren't a thousand different possible confounders (regression to the mean, of plain confusion about the level of competence of the other students taking the test), the DK effect isn't that strong.
People act like it's an uncanny valley of stupidity and hubris, when as presented it's more like perceived competence is only roughly correlated with actual competence, with minor bumps. And yet people keep invoking it as if it were the explanation behind all the incompetence in the world.
It has been quite a while now ~9 years, but I had the chance to work at a pretty prominent tech company for about 8 months when it was roughly 13 people in engineering. It was my first job experience in "tech" proper; and I think, I was most fortunate, to have had the opportunity. Immediately, and without question, I knew for sure I was absolutely the most junior person there. It's not because anyone was unkind, quite the opposite; everyone else was so unbelievably experienced it was obvious. I had, up until that point, not encountered such levels of professional experience, or expertise in a field, outside of university.
It gave me perspective for my (own lack of) ability I likely could not have gotten otherwise. It was humbling in a wonderful way because it meant I had, and have, a great deal to learn. I was happy to be resolutely at the bottom.
While undeniably good for my person, and who I want to be in life, it has also come to cause me a lot of pain.
Having worked pretty tirelessly since that original job, in the pursuit of learning and mastery of my craft. And if I am honest, to my chagrin, found I now underestimate my abilities. I only realized this because I generally give people the "benefit of the doubt" and it kept/keeps (still trying) biting me. It's very much insanity: doing the same thing again, knowing it won't work, expecting a different result... "'cuz someone else must be smarter than me." What I should learn to say instead of "smarter than me" is "less experienced and very confident."
Erik Dietrich wrote an article about the Dunning-Kruger where he coined the terms "Advanced Beginners" or "Expert Beginners." Like Dietrich, I tend to observe these folks, have the capacity, to cause the most harm for people in the competent/proficient areas (like I find myself). I believe they mean well, in most cases, but disregard anything not directly from an expert. The worst case being there is no expert available to help check them, and they will unrelentingly default to themselves. A close second, when they are the gatekeeper to the expert, and are unlikely or unable to explain someone else's ideas.
It's easy to see why it happens though. Unable to scope problems/tasks correctly; the problem's resolution becomes rooted in "beliefs," and generally speakings, people view other's beliefs as equivalent. With equivalent choices, we tend to then make decisions based on our preferences.
... but I digress, I've ranted and rambled enough.
[1] - Link to the referenced article: https://daedtech.com/how-developers-stop-learning-rise-of-th...
Such as?
And what about the 2nd and and 3rd quartiles?
And even if the phenomenon is caused by the mathematics of floor/ceiling effect (low skill people bias their assessment upward because they know they can't have a negative percentile), it's still real effect in misestimating their percentile.
Baseball players at bar can be modelled by random numbers too. That doesn't mean variance in performance over time is purely random.