This means that you can split your image into tiles, process each tile individually, average the results, apply a final classification layer to the average and get exactly the same result. For reference, see the demonstration below.
You could of course do exactly the same thing with a vision transformer instead of a convolutional neural network.
That being said, architecture is wildly overemphasized in my opinion. Data is everything.
import torch, torchvision.models
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = torchvision.models.convnext_small()
model.to(device)
tile_size, image_size = 32, 224 # note that 32 divides 224 evenly
image = torch.randn((1, 3, image_size, image_size), device=device)
# Process image as usual
x_expected = model(image)
# Process image as tiles (using for-loops for educational purposes; should use .view and .permute instead for performance)
features = [
model.features(image[:, :, y:y + tile_size, x:x + tile_size])
for y in range(0, image_size, tile_size)
for x in range(0, image_size, tile_size)]
x = model.classifier(sum(features) / len(features))
print(f"Mean squared error: {(x - x_expected).pow(2).mean().item():.20f}")