While I certainly don't want to give this sort of "benchmarking" any credibility at all:
allocating significantly different compute resources for different samples would be at least as laughable (eg. giving Go a dozen cores [per it's defaults] but constraining Python to one worker [per it's defaults]).