The main gotcha for local models is insane hardware requirements.
Even for $10K you get mediocre performance.
Even for $10K you get mediocre performance.
Comes down to how much of the ambiguity we expect out of the model.
Need 64GB vram, 900+ GB/s speed, and a lot of system ram (128GB). Seems feasible. Hmm.
Maybe older GPUs work.