But reality is usually much messier. The real training code will be littered with hundreds of ugly little tricks to make it work. A large part of it will be input preprocessing and data engineering, tricks to deal with exploding/vanishing gradients, monitoring, learning rate schedules and optimizer cycling, complexity for distributed training, regularization tricks, changing parts of the architecture for performance reasons (like attention), and so on.