As a result, higher pixel densities will cause light from a single point to be spread across multiple pixels, with different wavelengths going to different pixels. Software can probably reconstruct the image somewhat, but I'm sure you'll hit a limit where there is too much interference to get the image any sharper no matter how many pixels you squeeze in.
What you've described is chromatic aberration.
Diffraction refers to the spreading of waves around an obstacle -- such as the aperture blades in your camera. No lens need be involved -- you can get diffraction with a pinhole camera.
This results in an obvious merging of the blue/green strips. The more the diffraction the more they merge, the less sharp the images become.
Current APS-C sensors at 18 megapixels have pixels small enough that an aperture of around f/11 is the minimum size before this effect begins. As sensors get more and more dense, the required aperture size gets larger and larger. F/8 is considered a 'normal' aperture and once sensors hit that density then the payoff from increased resolution is significantly less and can impact image quality adversely.
This is why science missions use 2mp CCDs that they know the characteristics of, rather than some 40 megapixel phone sensor.
edit: I've just realised I've typed all this out with the wrong idea. You wanted a simple explanation of diffraction. Imagine a water tank with waves being generated from a source at one end. Put a wall with a narrow hole in the middle of the tank. If the wavelength of the waves is significantly less than the width of the hole, they will for the most part pass through unaltered. As the hole gets smaller, more significant changes to the waves occur. They begin to 'spread out' or 'diffract'. An intuitive explanation is difficult but this is a practical one that makes sense to people.
There's also technological barriers for creating large megapixel CCDs as opposed to CMOS. Large CCDs are usually created by "stitching" multiple sensors together, which increase their price as a square of the size.
You can shoot fullframe (35mm) camera at ISO 25600 and get decent results, but if you crank up signal amplification on iPhone to 25600 you'll get nothing but random color dots :-)
It would be foolish to suggest there was a single motivator for inclusion of a sensor. In reality it is a combination of our answers. Large sensors allows lower gain and better noise performance, enables smaller aperture use at acceptable sharpness, takes less resources to process, transmit etc.
You're not wrong in mentioning it though. Although FF at 25600 is going to be piss poor even if it's Nikon / Sony.
edit: corrected stupid misspelling
Nikon D600 at 25600: http://www.imaging-resource.com/PRODS/nikon-d600/nikon-d600G...
Not even close to "piss poor" in my book.
That being said, you don't go that high unless you're very desperate and the sensor itself can go up to 6400, the rest is software.
That was f/32. The rings around the street lamps are diffraction artifacts.
The higher your resolution compared to your aperture, the less photons you get per pixel.
Sensor size doesn't matter at all.
Diffraction cannot be solved as easily.
If your sensor counts a mean of 100 photons per pixel, then you'll see shot noise with standard deviation of 10 photons in each of those, for a signal-to-noise ratio of 10. If you quadruple your pixel size and now measure 400 photons per pixel, then your SNR goes up to 20.
This is why bigger sensors (more captured photons for same light level and exposure time) are fundamentally better at image capture.
I realize that is a very flawed analogy but it's the quickest I could come up with.
Here's more: http://en.wikipedia.org/wiki/Shot_noise#Optics