Edit - for anyone unsure about what "BERT" is or its relevance, it's a transformer based natural language model just like GPT. However, where GPT is used to generate text, BERT is used to generate embeddings for input text that you can then use for predictive models (e.g. sentiment prediction), and that process is also demonstrated in the notebook.
*Edit 2 - The 17 hours are pretraining only, not including the time to train the tokenizer, or finetuning.