I'm more interested in running distributed inference for purpose built small language models than these coding LLMs.
Say a distributed inference for image processing, SDR, local weather monitoring etc. These will run on mediocre specs and produce dependable output.
Nicely done OP.