> With AI systems, almost all bad behaviour originates from the data that’s used to train them
Careful with this - even with perfect data (and training), models will still get stuff wrong.
Careful with this - even with perfect data (and training), models will still get stuff wrong.
How do you define "perfect" data and training? I'd argue that if you trained a small NN to play tic-tac-toe perfectly, it'd quickly memorise all the possible scenarios, and since the world state is small, you could exhaustively prove that it's correct for every possible input. So at the very least, there's a counter example showing that with perfect data and training, models will not get stuff wrong.
But you're right - if dataset is exhaustive and finite, and model is large enough to preserve it perfectly - such overfitted model would work just fine, even if it's unlikely to be a particularly efficient way to build it.