So where did you load up Qwen and how did you supply the pdf or photo files? I don't know how to use these models, but want to learn
If the model supports "vision" or "sound", that tool makes it relatively painless to take your input file + text and feed it to the model.
[0]: https://lmstudio.ai/
LM Studio isn't as "set it and forget it" as Ollama is, and it does have a bit of a learning curve. But if you're doing any kind of AI development and you don't want to mess around with writing llama-cpp scripts all the time, it really can't be beat (for now).