1t model instances(opus, gpt,etc) are not running on a single GPU. The catch is how the cards communicate and how the model is broken up. There's a bit that goes into it but the answer is yes the more gpus the bigger the model you can run.
GPU interconnect speeds are a big bottleneck today for GPU's in AI applications. Data can't move between them fast enough.