If the result is statistically significant, it just barely makes it. 84.8% isn't that much higher than 80.8% and they had only 250 prompts, if I'm reading this right.
The ~5% improvement reported here might just be an artefact of the data collection or random variation, rather than a consistent repeatable change.
> The ~5% improvement reported here might just be an artefact of the data collection or random variation, rather than a consistent repeatable change.
You're questioning method or data representativeness, not significance. 250 samples is just about enough to for a 5% difference in NHST (stddev is around .4, so 1.64 sigma is .4/15.8*1.64=0.04 for single sided testing).