2,304 karma · joined June 4, 2009
Edit: This is wrong.
> genuinely
We can tell.
Not eating vast quantities of sugar is good advice, but it's not the direct cause of T2 diabetes. The best evidence suggests that T2 diabetes is caused by the accumulation of fat in the liver and pancreas. See the twin cycle hypothesis. To prevent diabetes, one needs to maintain a weight low enough where the body isn't storing fat in the liver and pancreas (everyone has their own individual threshold for this). If you're pre-diabetic, lose enough weight and most people will regain insulin sensitivity.
vLLM has the best performance if you can fit your entire model into VRAM. llama.cpp is the usual go-to if you're partially loading into RAM. LM Studio is a sensible front-end to llama.cpp. ik_llama.cpp has more advanced CPU quantization strategies than llama.cpp. If you're running super large models mostly from RAM, ktransformers can sometimes be the highest performer, if it works for your model.
Don't use ollama. The entire project is just a series of stupid decisions like this.
But also, control and consistency. A local model cannot be changed out under your feet like an API model can be.
That's not an emerging practice, it's a tested strategy that is these days only used as a last resort by those desperate to fit a model in memory. Some models do better than others, but generally the model quality suffers greatly under those conditions.
There's no guarantee that it works with standard clocked memory either.