How does that work in practice? Do you connect the LLM up directly to the app somehow, or do you ask it what to do and then do what it says and see what happens?
How much RAM do you need to make those work acceptably well? I assume Llama 3.3 means the 70B model, so you need > 70GB. (So, I guess, a MacBook with 128GB?) In which case I guess you're also using 8 bits for the Qwen model?