>
People hear or read the words Machine Learning and assign a HAL like technology to it.It’s definitely hard to explain and idea like SVM, it’s applications, and how it works/what it does without a background in some linear algebra.
I agree with your first point—there is a general lack of understanding about ML/AI, even among knowledgeable laypeople. But on your second point, I think this illustrates the tendency for technical folks (e.g. ML engineers) to overemphasize the importance of specific algorithms (e.g. SVMs, CNNs) and implementations or frameworks (e.g. TensorFlow). These are the things that are important to us, so we try to convey that when connecting with others. Even when I intentionally intending to simplify things to share enthusiasm and understanding with a non-technical audience, it is easy to slip into unintentionally alienating statements like, “machine learning is just matrix multiplication and gradient descent done on a GPU”. It might be a subconscious way of justifying our hard-won knowledge and its value.
But, I think it’s possible to have demystifying conversations that help people build a genuine understanding of ML/AI. And it’s also possible to do so in a way that instills a sense of fascination and respect for the field and the work, skill, and resources it involves to do well.
The themes of probability, statistics, and linear algebra can be honored and elucidated by discussing their core relevance:
Probability—What does a statement like “there is a 70% probability that this image is of a Fuji apple” mean? How does that differ from a statement about an event in the future like “there is a 70% probability that it will rain in London tomorrow?”. How do those probabilities change depending on factors that we can measure (conditional probability—the heart of statistics and ML)? What is an expected value and how does it relate to the ideas of risk and optimal decisions?
Statistics—What is a statistical model and what does it mean to “build” one? How much data do we need to collect for this building process, and in what format and subject to what assumptions and methodology? What are the inputs and outputs of a model that are relevant to my problem? What is “ground truth”, and how do we get enough examples of it with enough confidence? How do I know my model will actually work in the real world (generalization), and how bad is it if the model is wrong (will my users suffer an injustice or die, or just eat the wrong flavor of apple)? What are sampling bias and statistical bias, and how do they relate (or not) to bias in AI systems? What is a distribution? An anomaly? What are clusters, and how do we define whether two things are similar or not?
Linear algebra—How do we store “unstructured” data like an image or document so that a machine can work with it? How can we use math and computers to (efficiently) transform the data from input to output? What does it mean for a machine to “learn”? What is a tensor and why is it flowing? Wait, you want how much money to spend on graphics cards?
I get variants of the above (non-trivial) questions from interested but largely uninitiated stakeholders in government and business quite often. Approaching these conversations in a “big picture” way that respects peoples’ intelligence and curiosity is tremendously more rewarding and productive than getting down into eyeglaze-inducing technical/architectural rabbit holes.