The way you work around this (apart from doing what you can to make the system as predictable as possible) is to accept that the data is noisy, and then working around it by doing multiple trials and following up with a statistical analysis on the results.
You can use the desired confidence to inform the warm-up, number of trials, benchmark duration, and so on.
If you end up with a multimodal distribution it can be worth tracking percentiles.