There’s a lot of ambiguity behind that statement (fetch, decode, queue, pipeline, etc.) I’m just saying what I read in the article, the peak bandwidth calculation was making the assumption of a peak throughput of one instruction per cycle, according to the author.
The best answer might be to profile it yourself and see what happens on your processor, it’s only a few lines of code and looks incredibly easy to do.