Interestingly, this problem has been studied previously - an OpenAI paper called it the "noisy TV problem", since an agent could in principle get maximum novelty reward for just staring at a TV screen showing random noise.
The solution they proposed (https://openai.com/research/reinforcement-learning-with-pred...) called Random Network Distillation, uses a pair of networks, one learned and the other randomly initialized. The goal of the learned network is to predict the output of the randomly initialized one on the newly observed frame. If the output is easily predicted, it means the frame is similar to what's already been seen before, even if it has different TV noise or water patterns. They got some pretty good results on Atari and other games, so it'd be interesting to see if it works well on Pokemon.