It's a pretty simple and direct mapping. A single character is 8 bits, and a single base64 digit is 6 bits. They perfectly align at 24 bits. So it simply has to learn how to map every 3 characters to 4 base64 digits. Otherwise, there's likely tons of base64-encoded text in the training data simply from scraping the web.
It's not perfect though. I tested it on a few sentences of text and it made a few mistakes. Due to the way that GPT tokenizes the input text, it can't really generalize the pattern, as mappings of text to tokens is somewhat random. It effectively has to learn how to map every unique combination of 3 characters to 4 base64 digits, of which there are up to 2^24=16,777,216 distinct mappings. Otherwise, the number of characters in each token varies, which can also lead to mistakes.
You can use this tool to see how GPT3 maps text to tokens and token IDs: https://platform.openai.com/tokenizer
As an example, the alphabet "abcdefghijklmnopqrstuvwxyz" maps to [39305, 4299, 456, 2926, 41582, 10295, 404, 80, 81, 301, 14795, 86, 5431, 89]. This is what I mean by it's fairly random.