The diffuse cones send out a sampling of rays and attenuate the light from the light source, based on how many of those rays hit it, instead of some other object.
Instead of drawing light onto the scene and calculating how much passes through the viewport to the lens, ray-tracing cheats by working backwards, because photons traveling backward in time follow exactly the same rules as those traveling forward in time. Every photon that can travel backwards in time from the eye to hit a light source must have emanated from a light source with exactly the right direction and polarization to enter the eye. So the only photons calculated are the ones that contribute to the scene as viewed by the eye.
If I understand it correctly:
1) a point / pixel in the scene (as viewed by the eye) sends out a cone of rays, and the final color of this pixel is a combination of what those rays hit. This is the ray casting process, the reverse of light traveling.
2) the overall picture of the scene is the combination of pixels each calculated by the above ray casting process.
Am I right?
So for ray-tracing, you calculate along both paths and give 50% weight to each. Every time a ray hits a triangle in the scene, the material properties determine how the various components sum up to determine the color of the pixel.