It can be run on an A40 or A6000, as well as the largest A100s. But other than that, no.
How much VRAM does it use during inference?
~40 GB with standard optimization. I suspect you can shrink it down more with some work, but it would require significant innovation to cram it into the next largest common chip size (24 GB, unless I’m misremembering)