The problem is that many tests aren't about asking the specific questions; those questions are mere proxies for the presence of foundational knowledge. They're testing whether a student is building up a rich mental model of the domain.
So we're not really testing students to see if they can do X times Y arithmetic problems, and we're not anticipating that future work involves being a calculator. Instead we're using these tests as proxies for whether a student has a good model of multiplication.
ML tech has the possibility of making someone look like they have the foundational knowledge. We can ask people to design "better" tests, but that also makes things harder to cope for everyone else! Must we all take increasingly difficult and fatter tests when a cheap proxy would've sufficed?
Also, perhaps some might say that foundational knowledge is not foundational because ML tech can one day do everything. Yes, this is a possibility. That one day there's no point in learning foundational math because ML can handily beat people just like it does in Go and Chess. But in such a world we won't be talking about students cheating in school, we'll be talking about massive economic revolution.