The method feels problematic. The first trials are basically learning that one can associate celebrities with animals. Repeating that same type of task to the same group feels like latent learnings behind the first trial could account for a causal amount of the improvement.
Interesting overall still.
[edit] Full paper here. Issues persist, still cool. https://www.nature.com/articles/s41593-023-01324-5.pdf