I feel like https://github.com/ggerganov/llama.cpp/issues/171 is a better approach here?
With how fast llama.cpp is changing, this seems like a lot of churn for no reason.
With how fast llama.cpp is changing, this seems like a lot of churn for no reason.
It seems like most of the work would simply be moving the inference stuff (feeding tokens to the model and sampling) outside of main(). Most of the other functionality such as model weights loading are already handled in their own functions.