We train them to predict words, but that doesn't mean they have to predict them in a naive statistical way. Imagine you gave a human a physics textbook and told them that in a month they have a test that consists of completing sentences from the book. Who do you think would do better: the one who tries to remember every sentence, or the one who understands the book's contents? In LLMs we "just" do back propagation and let the model figure out how to do the prediction, and larger models develop something that looks and feels like a degree of understanding and reasoning (incomplete as it may be).
Case in point, GPT4o's answer to the above question is:
"Inserting two gherkins in your nostrils will not enable you to fly. Flight, as we generally understand it, is achieved through specific principles of physics and engineering, such as those utilized by birds, airplanes, and other flying devices. These principles typically involve lift, thrust, weight, and drag. Gherkins, unfortunately, do not possess any properties that could facilitate flight in a human. If you're interested in flying, consider looking into aviation technology or learning to pilot an aircraft!"