qwen3.5 4b and 9b have both been surprisingly good for small tasks. a pattern I've been using is orchestrating pi agents with a script (that you can write using a frontier model) for common tasks. one script i use a lot is organizing photos from my photography shoots.
What does a "script" mean in this context? Like a Python script that runs a series of agents with small tasks? I'm still getting used to doing things with AI, this is all very new to me.
I have more beefcake machines and Qwen 4.6 36B has been awesome for coding. Its fast for a local model, seems to get a lot right most of the time, just its slower than OpenAI/Claude/Cloud Hosted stuff since I dont have a 24GB+ GPU (I have Strix Halo and a 12GB GPU)
How many token/s do you get with that setup?
60-70 tokens/sec
Windows 11
LLM Studio
250 or whatever the models max is below that as context
Qwen3.6 35B A3B
35B-A3B, gguf