Of course you are right that this is not the whole picture, it is possible to be bottlenecked on neither memory bandwidth nor CPU throughput.
With respect to memory, two additional considerations are read amplification and latency.
Read amplification: Read amplification basically means that you are reading data but not actually using it. Suppose we have a table stored in row format with four columns C1, C2, C3, C4 which store 8 byte integers. Row format means that the values for the columns are interleaved, so the memory layout would look like this, where cn refers to some value for column Cn:
0x0000 c1 c2 c3 c4 c1 c2 c3 c4
0x0040 c1 c2 c3 c4 c1 c2 c3 c4
...
One property of conventional memory is that we can only read memory in blocks of 64bytes (the size of a cache line).
So if we have a loop that sums of all of the values in C1, we will actually also load all of the values for columns C2, C3, C4.
This means we have a read amplification of factor 4, and our usable bandwidth will be a quarter of the theoretical maximum.Latency: Suppose we also don't have any read amplification and make use of all values for each loaded cache line. Then after we finish processing a cache line, and load the next cache line, we might have to wait a long time (100s of cycles) before the new cache line actually arrives! So then we would actually be stalled because of memory latency during some parts of program execution.
One of the big advantages of using columnar layout (which stores all values for a column sequentially) is that it all but eliminates read amplification. Similarly, by accessing data sequentially we make it possible for the CPU to predict what data will be accessed in the future, and have it loaded into the cache before we even request it.
Of course there's a number of reasons why things might not work out quite so perfectly in practice, and those may why performance in this benchmark starts to degrade before the (usable) memory bandwidth reaches the theoretical maximum memory bandwidth. But as you can see it still possible to get very close, so it's safe to assume that queries which read much less data than the maximum bandwidth are constrained on CPU.