Building powerful image classification models using very little data
blog.keras.io
blog.keras.io
So it kind of depends on what you are trying to extract from such large images. If it's macroscopic features then you should just downsample your image to something reasonable and feed it into some pretrained network like Alexnet or VGGnet. If you want microscopic feature detection, but don't care about macro features, then you should use a shallower convnet.
If you insist on having both but the resulting network can't be stored in memory and distributed system is not an option, you might want to look into using recurrent networks with visual attention. These sort of look at manageable chunks of the image and decide where else to look based on what they've seen so far.
[0]: http://arxiv.org/abs/1411.6836 [1]: https://arxiv.org/abs/1403.1840