Picking a higher queue size can increase the peak throughput. The queue will smoothen out peaks and valleys in the workload, and the other end of the queue will get rarer into situations where there is no work. If you know the duration of the "valleys" in your workload you can size your queue exactly to get over them.
However a too big queue size can easily lead to higher latencies. In extreme situations the work might be already outdated at the point of time it's processed by the receiver. And backpressure on the producer is less given.
Queue sizes > 1 can also mask concurrency issues (like deadlocks), which will then only show up rarely in production when the queue is fully exhausted. I guess that's one of the main reasons why they picked the 0/1 rule.
Your comment (and others) have convinced me to do some more empirical testing and see how necessary buffered channels are for my goal.
The last sentence in the recommendation emphasizes this: use them with scrutiny.