You don't need Parallel corpora for all language pairs in a "predict the next token" LLM.
What I'm saying is that if an LLM is trained on English, French and Spanish and there is Eng to French data, you don't need Eng to Spa data to get Eng to Spa translations.