I like the way this builds up from simple models to transformers. There's still a pretty big gap between bag-of-words and neural network models, though, and one step that helps bridge that gap is continuous bag-of-words models, where you create word embeddings and sum/average them together for all the words in a document. You can use precomputed embeddings to improve performance for small training datasets in a way that's analogous to fine-tuning a foundation model. This is more or less what libraries like fastText do.