And you can certainly say “I’m going to enforce x rules on ingest and if it doesn’t match it doesn’t pass”. Depending on your system you may end up with a lot of on call or self-enforced down time or data loss. You might white list a schema and then later find out you’ve been missing out on 3 months of high value data because your upstream provider added a new columns. Alternatively, you can let things float through, monitor changes, but don’t let it plug up your system. This is also a risky game to play, with its own set of downsides.
Striving for comprehensive testing is still super important though. I find probably 60-80% of my time is spent tooling test frameworks that let us address the various edge cases. We use pyspark, and the amount of test-oriented tooling around that is, at least to me, surprisingly underdeveloped.