It's important to remember that the "power" of gpt doesn't come from the model, but from the sheer scale of the dataset. It's trained on the
entire internet, in text form. You can 100% use a transformer architecture to train on binary data. But what data do you have hundreds of tebibytes of?
Language also follows very common and repeatable patterns. "Hello" is often followed by "How are you?", etc. Just like Zipf's Law dictates that some words are used exponentially more than others, there are linguistic and conceptual patterns that appear with predictable frequency. If your bits don't follow similar rules, the results might not be as clean.
I'm pretty sure you could code a transformer to work on binary or video data. Sounds like a great github project. But it's unlikely you'll have the scale of data to do anything close to ChatGPT.