It’s fair to ask people bragging about their amazing AI skills to show their code or GTFO.
Hope it becomes established.
However, if you’re willing to share how you’re prompting the AI and some examples of where you think the results are poor, we might be able to help identify what’s causing the difference. I’d be genuinely interested in understanding why we’re getting such different results.
Rather than assuming one of us must be wrong, it would be more useful to compare approaches and see what we can learn from each other.