1. Using GPT-4, generate a text explanation of a neuron's activations on sample input.
2. Using GPT-4 again, use the text explanation to simulate the neuron on some new text input.
3. Compare the result to the actual neuron's activations on the new text input.
They justify this by saying human contractors do equally poorly at coming up with text descriptions. However, the procedure is such a black box that it is difficult to make scientific conclusions from the results.
[1] https://openai.com/research/language-models-can-explain-neur...