But if this technique scales up, then Anthropic has fixed that. They can figure out what different groups of neurons are actually doing, and use that to control the LLM's behavior. That could help with preventing accidentally misaligned AIs.
But if this technique scales up, then Anthropic has fixed that. They can figure out what different groups of neurons are actually doing, and use that to control the LLM's behavior. That could help with preventing accidentally misaligned AIs.
Hm. I wish they'd said more about that. Does that mean they found the same feature recognizers when training with the same training set? Or what? This tells us something, but what does it tell us?
Typically, you can take a pre-trained model and retrain it on your new dataset by only changing the weights of the last layer(s).
Some loss functions even measures the difference between the high-level features of two images, typically extracted from a pre-trained CNN (Perceptual Loss).
[1]Matt Zeiler did an amazing work on these findings 10 years ago (https://arxiv.org/abs/1311.2901).