Did you compile accuracy, F1 numbers, or anything like that? Do you have quantitative comparisons of results you got w/ different models?
Are you measuring accuracy with data wrangling prompts? Would love to learn more about that.
I'm skeptical of any claim that "A works better than B" without some numbers to back it up.