I have a smart RSS reader YOShInOn which uses BERT + a probability calibrated SVM as its main model, I want to make a general purpose model trainer for text classification that is able to do harder problems.
People who hold court on ML forums will tell you fine-tuned BERT is the way to go but BERT fine-tuning doesn't seem to be compatible with early stopping with anything like the training recipes I see in the literature. Compared to old days these networks soak up knowledge like a sponge, my hunch is that with N=10,000 samples or so you don't benefit from running more than one epoch because the the network doesn't have the capacity to learn from that many samples.
I find it depressing to find arXiv papers where people copy a training recipe from other papers for BERT and compare it 5-15 different text classification problems with maybe N=500 samples. My BERT experiments take about 30 minutes so it's no small thing to do parametric scans on them, particularly when the epoch count is one of the parameters. With "smart stopping" I'm not afraid of undertraining models so I could run trainings all night and believe I'm seeing representative performance as I vary parameters.
My plan is to couple ModernBERT to a LSTM or Bi-LSTM model as the literature seems to show that this frequently ties or beats fine-tuned BERT and my experience so far as I can build reliable trainers for LSTM whereas team fined tuned BERT is indifferent to the very idea of "reliable".
Another pet peeve is all the papers with N=500 samples where I regularly get N=10,000+ in systems that I use everyday and on a rainy weekend I can lay in bed with my iPad and switch to an Android tablet when the battery runs out and get N=5000 samples. [1] When I wrote my first text classification paper we found we needed N=10,000 to get really good models, sure the world knowledge in BERT helps models learn fast and that's great (a problem I worried about in 2005 and still worry about because I think the average person wants good results at N<10!) but I need calibrated usable accuracy and look at AUC-ROC as my metric, not "accuracy", F1 or anything like that.
Then there's the effort people waste with things that can't possibly work like Word2Vec, seems like people can read a lot of papers and not see it in front of them that Word2Vec is useless! I want to write a meta-analysis but instead I'm writing a diatribe and I'm not going to be happy until I repeat the paradigm with methods that are... repeatable, not for the science but for the engineering.
[1] with hallucinations as a side effect if it is a visual task but so what