> Is there something that prevents AI software from augmenting VRAM with system RAM?
As I understand, it adds complication, the points where it is practical without just getting into an horrible churn of shuffling are limited based on model architecture, and (as a consequence) implementations may not always been transferrable between models outside of a narrow family.
On the other hand, there are some frameworks with general support which eases this.
On the third hand, shuffling stuff back and forth between system RAM and VRAM has a performance cost in the best case.
So, yes, and no, and yes.
Most (that I've seen) consumer-focused frontends for local models (text generation, image generation, audio generation) do it, to a greater or lesser extent.
The base model implementations that are getting thrown out by researchers don't usually focus on that, and its generally implemented by third parties after the fact, sometimes with commonalities that can be applied to other models that are closely-enough related.