In principle you could have bidirectional PCIe x16 pipelining at it would move the roofline a little with fast DDR5, I think llama.cpp has a flag for it.
Or go rent a B200 on vast.ai for 4 bucks an hour or thereabouts, a single heavy Opus session for a couple hours is like a week of any model on vast or RunPods.
NVIDIA publishes something called NGC containers that generally work out of the box. I started running Qwen3.6-NVFP4-MTP locally and then I'll put something heavy on Baseten if I'm lazy or Vast if I want a good deal.
Opus (and maybe now 5.6) are still the strongest for like, the really delicate shit, kernel modules or something, but that's on pace to cross over this year, and the overtraining and misalignment are getting so bad when they phase 4.6 out I'm pulling my plan. I don't pay to get gaslit about Constitutional AI.
It's time to have an exit strategy.