6GB of VRAM! Luxury! I managed to reconstruct the lego scene using TensoRF[0] (not strictly NeRF, but a similar approach with similar results) last night on my nvidia-equipped T480 (2gb of VRAM)[1], so it's possible.
>But, there are really no viable models that run on consumer hardware like llama or stable diffusion.
This isn't a strictly accurate framing, there is no pre-trained "model" that you "run" inference on like with Llama or Stable Diffusion. You are training the model, from scratch, on each new scene. The viability of this on a given GPU depends on the combined size of the input + output, i.e. the resolution and number of input images and the resolution and compactness of the resulting data structure. There's nothing in principle preventing you from training tiny low-res nerfs from tiny low-res images, except that all the researchers in this space are working with standard datasets of a standard size on big beefy machines and their code is full of magic numbers. Also, many of the improvements on the original NeRF achieve their speedup through much hungrier data structures (voxels, multiresolution, etc). TensoRF appears to have a very compact scene representation (like the original NeRF) and very fast training (like instant-ngp) so it seems to be a sweet spot for low-end hardware - at any rate, it's the first thing I managed to get working on this laptop. The main downside seems to be that inference (generating new images) is quite slow, at about 14 seconds.
[0] https://github.com/apchenstu/TensoRF/ - it's apparently included in nerfstudio as well
[1] batch_size = 512 in configs/lego.txt (with your 6gb you'd get away with 2048) and compute_extra_metrics=False wherever it appears in train.py