I often find myself going back to Eurisko as a fascinating example of the road not taken (or perhaps, the road taken to its logical end and abandoned) in machine learning. You had this system that instead of learning opaque weights in a network, learned heuristics that were fairly explainable. In some ways that's a much more useful companion than a model that's strictly superior in its decisions. I find this regularly in chess - you can prep with Stockfish and you know you're probably memorising the best line, and you can follow the sidelines to find very concretely _why_ that's the case, but it still feels like an inefficient way of learning. It always feels like a machine that yielded up fewer, more abstract principles would be helping me more. I recently thought about this for solving Rubik's cubes. I don't really want the fastest way to solve it, ideally I'd have a way that simultaneously optimises the number of heuristics (if you have such and such a colour here and here, do this, if this face is all white, do that etc) and the number of total actions. That seems like a tractable problem for a computer to brute force for you even without smarts.
I face this in my work constantly: you can come up with the most amazing model based on the most cutting edge architecture, and you can trust its predictions above any others. But then you've got to explain to a football coach why _they_ should trust it, and once that trust is built, distil it into few enough rules that they can train a team to reflect its expertise.
Sorry for the ramble!