Neural networks need data to learn, even if it’s fake
quantamagazine.org
quantamagazine.org
We're using software that has encoded the rules we want to use to generate data. The ML is using the data to infer the rules. It seems like we need a better way to give our ML algorithms the rules that we already know rather than this processing-intensive process.
But perhaps that's the cost of using neural networks- they can learn any function, but they can only learn from data, examples.
Neural nets work by generalizing. This is important because GOFAI systems are brittle. We are never able to encode all the relevant rules for real world situations by hand.
Also, note that the synthetic data is generated from rules about how the world looks, not rules about what the model should do.
There may still be some direction for zapping the rules-of-the-world right into the circuitry of, say, a subnet of the model, without doing backprop. Something like a one-shot initialization of the network to approximate the given distribution.
The combinatory bit due to its branching nature lends itself poorly to leveraging SIMD hardware as statistical solutions do. And getting a right representation model is a substantial challenge. Everyone knew the importance of that ('knowledge engineering') by the onset of AI Winter but nobody had any good methods.
There are cases when we want a machine learning model to do the same thing as the process which generate it's data, like in the case of model's learning to replicate physics simulations, but even then the entire point is for the machine learning model to accomplish the same or a similar result but in a more computationally efficient way.
So in the end you still end up fitting some simplified vaguely similar model/function to your data and/or using sampling to keep things tractable.
And of course, as alluded in the article, sometimes we don't really know the "rules", or don't know how to articulate them in a way we can feasibly synthesize data according to those rules. Then people might reach for ML to figure out those rules, implicitly or explicitly. So it becomes a bit of an ouroboros situation. But still somewhere in that loop there's some domain knowledge injected somewhere by whoever is engineering the pipeline.
The only thing is — we start collecting data since birth and we also don’t have to pay for every little thing we put in front of our eyes.
If you’re training an AI, you have to pay for data. If you generate some of your training data, you pay a little less.
(Also I think we ask of AI more than we ask of humans. For example, it’s common to be face blind of ethnicities you’re not familiar with but we get pissed if an AI is face blind at all.)
No, I don't agree with this. My point is that we can learn from higher level information than raw data.
I can give you the formula of a complex mathematical function over 3 variables, and provide no data points.
You'll be able to tell me with perfect accuracy what values match the formula and which don't.
If you're just one guy who's face blind, no one really cares and you can't cause that much damage unless you're a border agent or something. Even then, it's not like you're affecting every single person passing through your country's border, just the ones who happened to be in the right (wrong?) place at the right time.
But if you deploy a face-blind AI as a security system that discriminates against minorities in thousands of establishments, suddenly it's like having thousands of clones of the same face-blind human and it can affect millions of people. If we're effectively making that many clones of a human, it better be good at its job
https://www.capitalone.com/tech/software-engineering/why-you...
Part of the reason was models generating synthetic data hone in on more key differentiable features
If there's nothing to be gained from training on pure noise, what is the most generic and generally useful sequence we could train on?
In his 2019 book "Human Compatible", Stuart Russell commented that 'Unfortunately NELL has confidence in only 3 percent of its beliefs and relies on human experts to clean out false or meaningless beliefs on a regular basis—such as its beliefs that “Nepal is a country also known as United States” and "value is an agricultural product that is usually cut into basis."'