The same model will give the same result, and more processing power will simply enable you to get the inference done faster.
On the other hand, more resources may enable (or be required for) a different, better model.
The same model will give the same result, and more processing power will simply enable you to get the inference done faster.
On the other hand, more resources may enable (or be required for) a different, better model.
Is it wrong to think of this as misleading? Don't the results for exactly the same request differ because there are multiple output strings with the same computed weights?
Or do you include "multiple ways to phrase the same" in "same results" and I'm being a noob?
So if you want it to spend more "time" in a useful manner without changing the architecture, you have to get it to write down the temporary information in the tokens, as "think step by step" does or alternatively iterative prompts "write a draft for the rough structure" "now rewrite it better with more detail".
It also feels like a multiplication of required processing power but I have no clue yet how one could use the previous generation of weights of and the tokens themselves to improve, elaborate on, widen the range of predicted potential results.