Machine Unlearning Challenge
ai.googleblog.com
ai.googleblog.com
Dr. Mierzwiak: "Well, technically speaking, the operation is brain damage, but it's on a par with a night of heavy drinking. Nothing you'll miss."
- Eternal Sunshine of the Spotless Mind (2004)
- Futurama, Parasites Lost (2001)
- Simpsons, The Front (1992)
Edit: People seem to disagree, but is this not a tool for wiping ideas from your training data? Meaning they want to train on Wikipedia and remove concepts (per locale probably) easy as pie. This is Google outsourcing compliance with jurisdictions that require thought policing.
Even though they talk about removing biases from the model, I doubt this would make the model unlearn whole ideas/data patterns without severely damaging performance, which they clearly don't want.
Gradient descent obviously has none of that, it has not feelings or goals, and so there would be no preferential remembering or forgetting other than to do better next word prediction or rlhf or whatever. So it would be interesting to think about what a model should remember or forget in order to align it with our goals (because it doesn't have any of it's own).
Also, don't volunteer for Google. They have lots of money, they can pay for this stuff and if you know how to do it you have lots of options.
My totally amateur pet theory is that paranoia is threat pattern recognition gone bonkers.
Actually, most of what we do with brains is pattern recognition, and if there isn't enough good input they will make shit up.
Supposedly OCD is your brain’s cause-effect loop being too potent. As in, you have a random fear I.e. “stove is on, fire will burn down house,” you go to check the stove, and the act of checking (regardless of its being on) creates the feeling that you saved your house from burning down, so now you feel compelled to check every time. Paraphrased from the Huberman podcast.
He was also quite unassuming and very "normal". No trace of mental dysbalance.
While from a privacy perspective, not all 'data points' (i.e. memories) need be shared with the populace at large, to the individual they do constitute affective buildings blocks to deal with previous experiences. Similarly, the threshold for sharing (private) experiences depends on the communicative context one's in and how comfortable one feels sharing them.
Might similar mechanisms provide large models the ability to retain an internal recollection of 'traumatic' or 'problematic' data and become more attuned to the context its communicating in instead? Leaving blank spots or blocking out experience completely sets the stage for 'disturbing' a model's memory after the fact. While I don't want to draw a comparison between current NN architectures and organic ones, it is worth questioning our framing of currently emerging methods, as the incipient terminology can imbue how later practioners implement and think about these mechanisms.
This idea doesn't fully remove the influence of the target data (any previously saved gradient update from after a contaminated batch contains some information about the state of the network prior to update) but it may be a sufficient and efficient way to quickly reconstitute a network with far less influence from the problematic data.
Just an idea and I haven't tried it, so maybe it's bunk, but there you are!
On the face of it, I would expect the gradients to take about as much space as the weights. So you’d be checkpointing your network at every batch, in effect.