> For my local setup, I’m currently [..] and LM Studio as the inference server, although it would likely be faster if I just used llama.cpp directly
Is there any truth to this claim? LM Studio uses llama.cpp to run the models. I guess the overhead of LM Studio should be minimal.
After all LM Studio is a really easy way to host models, are there really major drawbacks?