Just ran it on one of our internal PDF (3 pages, medium difficulty) to json benchmarks:
gemini-flash-2.0: 60 ish% accuracy 6,250 pages per dollar
gemini-2.5-flash-preview (no thinking): 80 ish% accuracy 1,700 pages per dollar
gemini-2.5-flash-preview (with thinking): 80 ish% accuracy (not sure what's going on here) 350 pages per dollar
gemini-flash-2.5: 90 ish% accuracy 150 pages per dollar
I do wish they separated the thinking variant from the regular one - it's incredibly confusing when a model parameter dramatically impacts pricing.