It's all just about the awareness of contexts. Want to improve it? Simply add a term to the prompt to unlock more considerations. Assuming we've not reached the edge of the context window, every new word "unlocks" new vectors with more context the language models adds to the considerations.
The similarity with how the human brain (seems to) works is so remarkable, it doesn't even make sense not to use it as an analogue for how to better use language models.
When the results (same way of manipulating an LLM as manipulating a human brain ... using the right words) can be achieved the same way, why believe there's a difference?
This is stuff one can learn over time by using/researching 3B models. While most people seem to shun them, some of them are extremely powerfull, like the "old" orca mini 3B. I am still using that one! All they really need is better prompts and that approach works perfectly fine.
The biggest hurdle I've found is the usually small context window of such small models, but there's ways of cheating around that without sacrificing too much of the quality using small rope extension, summarizing text, adding context words or leaving out letters of words in the prompt, virtually increasing the size of the context window.
If you want to improve the results of your language model, you should become a mentalist/con-man/magician/social engineer. It sounds weird, but it works!