Inter-subject variability is a huge problem in this kind of work.
As you say, they're operating at a much larger scale than individual cells (100k-1m cells in each voxel). Likewise, some early visual processing areas are broadly organized kind of like big, noisy bitmaps on the surface of the brain.
But for sophisticated machine learning-style analyses like these, the gross differences in representation and morphology (especially at higher processing levels in the brain) make it very hard to pool the data across multiple people. That's why they're preferring to use many sessions from a small number of participants rather than a single session from many participants (the standard approach).
[I worked on applying machine learning methods to fMRI for my PhD]