ParentFull threadjustaboutanyone·Running llama.cpp rather than vLLM, it's happy enough to run the FP8 variant with 200k+ context using about 90GB vramView on HN