I work on proprietary software, so no, that’s not really a fair question to ask.
However, if you’re willing to share how you’re prompting the AI and some examples of where you think the results are poor, we might be able to help identify what’s causing the difference. I’d be genuinely interested in understanding why we’re getting such different results.
Rather than assuming one of us must be wrong, it would be more useful to compare approaches and see what we can learn from each other.