There is the synthesis of sounds and there are a variety of models for that, such as FM synthesis, additive harmonic synthesis, physical modeling, wavetable synthesis, granular synthesis and more. Physical modeling is somewhat like ray tracing, modeling the materials and sound paths in math to imitate a real instrument (or any object) making sound.
There is modeling of acoustic environments, which is very much like the ray tracing of audio, dealing with reflections, damping, phase, refraction, dispersion and reverberation.
There is the physics of audio perception, how the ear works, how the brain perceives sound, psychoacoustics. What makes your ear "think" you are indoors or outdoors?
There are encoding standards, AAC, WAV, MP3, FLAC, which can also get into psychoacoustics.
There are algorithms, like DSP signal processing algorithms for audio which do reverb, EQ, echo, psychoacoustics, sound projection/direction, flanging, surround sound, and a lot more for shaping sounds.
There is the study of how to build optimal acoustic spaces, like concert halls and music studio performance rooms. More psychoacoustics, but some math and sensors also.
There is the art of trying to use all the above to accomplish something, a song, a soundtrack, a concert, a jingle.
So there are a lot of different aspects to audio, and I've probably left some out of the above list (voice, foley, sound effects, subsonic, supersonic,...). I'd suggest you first work on defining your problem better, what do you actually want to solve, and for what audience? There are many tools that do a lot of this already. If you start out with the most basic audio physics, how sound travels through a variety of materials, and what happens at transitions between materials you could probably spend many years working on just that (as a comprehensive system) before you even get to trying to generate sounds. I don't think there is a complete system that can be used to predict all aspects of sound propagation through materials (including air) because materials can be extremely complex, for example air has air currents, variations in humidity, air pressure is variable in time and space, and the sounds in air are also strongly shaped by reflections and the qualities of the surfaces it reflects from. Imagine the complexity of modeling all the trees, leaves, plants and rocks in a forest. You can "record" the effects of one specific space in a forest using impulse response recordings and apply it to sounds, but to model it from scratch and be able to recreate any forest location would probably be a lifetimes work, perhaps several lifetimes, unless you can find a clever way to do it (maybe generate artifical fractal landscapes based on parameters measured from real places, and then do acoustic ray tracing?)
Acoustic "ray tracing" is not like light ray tracing because sound is much lower frequency and diffracts, it bends around objects instead of bouncing off, and also bounces off. The phase of the interacting "rays" or waveforms as they reach the ear also matters a lot more with audio than with light which makes the computations more complex. In a light ray tracer light travels in straight lines from sources to virtual camera. In a sound ray tracer sound can bounce and refract and disperse from anywhere to reach the virtual ear. What is on the other side of a wall can matter to what a room sounds like.
Don't let my talking of complexity deter you though, look and maybe you'll find something nobody has thought of before. It might be a matter of plugging together some tools that already exist (e.g., a 3D definition of a room or outdoor space, sound source models, choosing a virtual ear location, and let the computer crunch for a few days.)