My research area has actually started doing that. Sure it's a CS discipline, but yeah we have a second component of our conferences where we automate our benchmarks, and other people run those benchmarks and validate the results.
It's absurd that reproducing results be any more of a challenge than simply re-running a freely-available program.
Software is unique in that it can be generally be duplicated and executed trivially.
There's no way to make it trivial to reproduce a test on the strength properties of a new ceramic. There is a way to do this for software, and it's rather silly that it isn't standard scientific practice to do so.
I realise I'm taking a strong line here, but I've never seen a good argument against it.