As a hobbyist, I’ve wondered if the need for umpteen epochs just leads many nets to memorize datasets, especially when the performance jumps a lot from one epoch to another without much change during batches. It’s kind of disconcerting for those of us who don’t have millions of source images to train with.