As you mentioned, micro-benchmarks are riddled with bias, and confusing results.
So putting micro-benchmarks aside and looking at ES5 / ES6 induced performance bottlenecks in actual apps, they are clearly present. (note, this is of-course limited to features actually implemented)
Unfortunately, a macro focused endeavor (only gauging full app performance) isn't as nicely actionable as the micro.
So in practice, in an attempt to produce high value but actionable feedback a hybrid approach, utilizing both micro/macro investigation yields the best results.
In addition the micro-benchmark vs optimizing compiler trap can be mitigated by investigating the intermediate and final outputs of the optimizing compiler.
Anyways, /rant.
There exist an unfortunate amount of JS performance traps, I wish where taken more seriously. Although more work, it would be quite valuable for someone to perform a root cause analysis to investigate potential bottlenecks brought to light by this post.