Empirical comparison of programming language productivity (2000)
cis.udel.edu
cis.udel.edu
Why didn't they just measure them all?
In any other study of workers, I'm pretty sure even the author would feel self-reported data would invalidate the results (especially when all of the "best" scores were self-reported).
e.g. The ten chosen Honda dealerships reported that they can fix a Honda in 10-23 minutes compared to our measurements that Ford dealerships can fix a Ford in about an hour and a half.
Maybe we could get a version where the program was to implement a decent sized project requiring several developers, to emphasise the code quality angle? That would take a lot of bored programmers to do, though.