Hard to say tbh. Could be infinite, could be zero.
In classification batch is going to help a ton (including stability) and you may not even get top results without a high batch, so infinite in that case. More realistically, consider that batch also helps with GPU I/O. I/O slows things down quite a lot, so if you're moving stuff in and out of the accelerator a lot then you can easily be I/O bound.
For generation the conversation gets more complicated. For GANs the smaller card could do better if you can fit a proper batch size there (probably not tbh but often closer to 16 or 24 so that may be better here). Here distributed tends to help more. For VAEs, Diffusion, Autoregressive, etc are more stable and can also hugely benefit from large batches.
So the truth is an unsatisfying "it depends." I know that may not be the answer you're looking for, but it is the real one. It takes a lot of work and tricks to scale (as anyone working in non-ML HPC will tell you very similar stories about scaling). But also note that the clock speeds aren't as big of a gap as you note though the memory gaps are.