I'd like to do a small study personally, that should take me about 10-20 hours, and I'd like people's feedback here about what they think about the methodology.
- First, I'd put together ten generic keyword-laden resumes using machine learning from a large dataset - which means it should somewhat reflect an "average" programmer, without representing anyone in particular. It will literally therefore be the creation of a non-existent person.
- Next I would edit it by hand so that it doesn't seem grossly machine-generated. At this point I'm still blinded, and I'd have ten programmer resumes.
- Next I would generate a separate list of common first names and surnames. Since most people are white, likely most names are somewhat white but this shouldn't be that relevant, also given that people have the right to change their name legally. The importance of this is that with such generic names, it is almost certain that there are a large number of programmers with that name, and so googling by the employer won't point to anyone in particular. At this point I have ten resumes and ten names.
- Finally, I would find stock art of a generic black man and generic white man. I would try to pick neutral images that if I were such a person I could actually have taken of myself, and may actually use on a resume. I would only just find two, since the goal of the experiment would be to do A/B testing on just the effect of this image. This is what is being studied.
- For the experiment itself, I would create email addresses for the ten names, and write cover letters that match the ten resumes.
- From each email I would email twenty-thirty companies based on a keyword search from the associated resumes for positions that seem a match. I'd be careful to pick all different companies.
- I'd carefully rewrite the cover letters (in my own voice) about how excited I'd be to work there and what a great match the position seems to me, as they can see from the attached resume.
- Finally, and this is the tricky part, when actually attaching the resume, in a blinded way I will use a script to insert either the black or the white man's picture. I must not see it in order for this to be a doubly blinded experiment. This is actually easy to do technically: .docx files are just .zip files, you can rename them, change the image in the zip file, and zip it back. I don't need to record anywhere which of the two the script chose (randomly), because as soon as I send it it will be in the sent folder.
The same name and resume must go to a mix of companies, some being sent the picture with a black man, some being sent the picture with the white man. Since I'd be careful not to email the same companies, it doesn't matter if the same name/resume is sent to some companies as a black man and some companies as a white man.
What my prediction is, is that the picture will have a statistically significant effect on emails back. (I don't want to bother setting up ten real phone numbers, what a pain, so the phone number will have to be omitted from the resumes.)
Now, this method isn't perfect. For example, real programmers frequently have github profiles or a large online presence. Still, by sticking to the most common of common names, perhaps there should be enough of a profile to be worth an email back, especially given an enthusiastic cover letter, even if they can't find this programmer in particular.
Finally, after the experiment the recipients could be informed that they were part of a study. But perhaps this is not so important. After all, if they don't email back they likely have forgotten about the resume, and even if they do, if the applicant does not answer their email then the applicant must simply be busy with other offers.
Of course, for authenticity purposes, it would be better not to share this methodology here, where some people might read it and be tipped off.
But given that I am not certain that I am doing things right, and haven't run an experiment in the social sciences before, while at the same time I've heard many people here report on experimental methodology, I thought I would run this methodology past the HN crowd. What do you think? Is it scientifically well-constructed? Is it possible for it to show an A/B effect?
Short of literally reusing people's real resumes with a fake image, I don't know of any way to improve this proposed methodology.
-> Shall I run the experiment?
-> Is there anything I can do to improve it?
-> Would the results be meaningful in either case? (If it does show a statistically significant deviation in email-back percentages, and if it doesn't.)
Thanks for any thoughts.