If you watch the progress of the reasoning in llama-server while it's doing the thinking, you can track its progress. Sometimes the dead ends it goes down or things that it considers and then disregards are themselves something useful to re-prompt it with later, and send it 'rolling downhill', to use the metaphor of another commenter here, in another direction towards the same effort.
Putting 3.6 35B-A3B into a state that lets it spend a lot of time in its reasoning mode before outputting an answer is probably not something that a web based SaaS LLM would tolerate, because it would frustrate many of the non technical end users who want a LLM to spit out an answer now.
It’s in the name :)
Everything that an LLM outputs is just a statistical language-based (no real grounding) prediction. Luckily with a model based on a large training set most common questions may elicit coherent responses from the training data, but you don't need to veer too far off into "questions less asked" territory to get responses based on training data mashups that amount to best guesses that are wrong, aka hallucinations. The unfortunate part of this is that as a user you may only catch this when asking a question about something you are already fairly knowledgeable about, then you give some pushback to the model and it cheerfully acknowledges "you're right - I made that up".