It's hard to find information on. The company used to be called "Arbitron" and they filed several patents on the technology, which is where I did most of my research.
There are 10 bands, each of which can carry 1 of 18 tones. 16 tones are to signal, so the stream is 4 bits wide, and the other two tones are used as a "STOP" and "SYNC" marker in the data stream.
All bands carry identical information, but each bit within a band is encoded with a different tone than any other band. So, it appears random or uncorrelated at first glance, but you can see the pattern pretty quickly if you do a long enough analysis. The same message gets identically repeated 12 to 13 times per minute and the message only changes once per minute. Here's an example isolation of the signal [1].
The coder monitors incoming audio and does a psychoacoustic pass to determine which level the tones should be injected at. If it can't find a good level for a tone, it uses the lowest possible level... which usually works out fine due to the 10 redundant signals in the watermark.
Encoders like MP3 see the injected tones as signal and will make the bits available to ensure they're encoded rather than correctly masking them out as non-audible material. The 10 bands help here as well, as even if the encoder fails to code the signal in one band, there's usually enough redundant bands that you can successfully decode the watermark.
[1]: https://www.sigidwiki.com/wiki/CBET