Curious on how (if?) changes to the inference engine can fix the issue with infinitely long reasoning loops.
It’s my layman understanding that would have to be fixed in the model weights itself?
It’s my layman understanding that would have to be fixed in the model weights itself?
It can also be a bug in the model weights because the model is just failing to generate the appropriate "I'm done thinking" indicator.
You can see this described in this PR https://github.com/ggml-org/llama.cpp/pull/19635
Apparently Step 3.5 Flash uses an odd format for its tags so llama.cpp just doesn't handle it correctly.
It is a bug in the model weights and reproducible in their official chat UI. More details here: https://github.com/ggml-org/llama.cpp/pull/19283#issuecommen...