Think about how you learned anatomy. You probably looked at Netter drawings or Grey's long before you ever saw a CT or MRI. You probably knew the English word "laceration" before you saw a liver lac. You probably knew what a ground glass bathroom window looked like before the term was used to describe lung findings.
LLMs/LVMs ingest a huge amount of training data, more than humans can appreciate, and learn connections between that data. I can ask these models to render an elephant in outer space with a hematoma on its snout in the style of a CT scan. Surely, there is no such image in the training set, yet the model knows what I want from the enormous number of associations in its network.
Also, the word "finite" has a very specific definition in mathematics. It's a natural human fallacy to equate very large with infinite. And the variation in images is finite. Given a 16-bit, 512 x 512 x 100 slice CT scan, you're looking at 2^16 * 26214400 possible images. Very large, but still finite.
Of course, the reality is way, way smaller. As a human, you can't even look at the entire grayscale spectrum. We just say, < -500 Hounsfield units (HU), that's air, -200 < fat < 0, bone/metal > 100, etc. A gifted radiologist can maybe distinguish 100 different tissue types based on the HU. So, instead of 2^16 pixel values, you have...100. That's 100 * 26214400 = 262,440,000 possible CT scans. That's a realistic upper-limit on how many different CT scans there could possibly be. So, let's pre-draft 260 million reports and just pick the one that fits best at inference time. The amount you'd have to change would be miniscule.