As has been pointed out elsewhere in this thread, they have nice mathematical properties. But another important thing is that they typically work "well enough" for applications. Consider audio. Tones clearly have frequency, but they also have a position in time. Doing a sine-cosine decomposition of a whole song doesn't really make sense, since it has no way of saying that a tone on the piano is played at a given time.
So you would think that it would make sense to break the signal down into stuff with frequency and time. Some kind of wavelet probably. Maybe something that very accurately models what a human hears.
The thing is that chopping the audio stream up into windows, and decomposing those windows into sines and cosines, while a bit ad hoc, just works well enough.