Hmm, the article's summary table showed "only the top task layer(s)" for BERT fine-tuning. Is that correct?
The original paper [1] wrote:
"All of the parameters of BERT and W are fine-tuned jointly to maximize the log-probability of the correct label."