The Ghost in the MP3 (2014)
theghostinthemp3.com
theghostinthemp3.com
MP3's unwitting/unintended uncovering of the possibilities of weird filtering schemes points to the hugeness of unexplored timbre space. We are so familiar with our usual collection of instrumental timbres ... we have assumed that this galaxy is the whole universe. Timbre is waiting for its Hubble.
Jesus fucking Christ, man. I'm going to remember that phrase for the rest of my life. Good show, good show.
standing_ovation.gif
Some references you may enjoy:
* Music: a Mathematical Offering, Dave Benson. Explains the mathematical theory of how to compute the timbre of an instrument given its shape. Also how the timbre of the instruments determines harmony.
* Tuning, timbre, spectrum, scale, William Sethares. A deep exploration of the relationship between timbre and harmony. It starts by exhibiting a (synthetic) instrument whose octaves are dissonant, and the music that you can do with it. Then it doubles down on this idea to obtain a lot of fun.
* The Topos of Music, Guerino Mazzola. Really hardcore mathematical music theory, too scary for simple people like me.
EDIT: regarding telescopes for seeing timbre... have you ever seen an spectrogram? You can install a spectrogram app in your phone and see the spectra of sounds around you. There's a lot of things! My favorite is opening a rusty old door.
[0] https://archive.org/details/onsensationsofto00helmrich
Sethares sounds familiar. A key modern classic is 1985 Charles Dodge, Thomas A. Jerse: "Computer Music : synthesis, composition, and performance". A pricey book new (worth it!) it digs into everything from the basics to the advanced. You will not exhaust it!
Surprised but, for a limited set of users, this puppy can be accessed online. See the link for the list of topics. [1]
[1] https://archive.org/details/computermusicsyn0000dodg
An alternative to the computer is to take a little recorder with a good mike out into the wild and record, the rocks, the insects, the trees, the plants.... and water, the greatest improviser!
> Listening tests, primarily designed by and for western-european men, and using the music they liked, were used to refine the encoder.
> As previously stated, the MP3 codec was refined using listening tests designed by european audio engineers and featuring the music they chose. In a sense, each of these songs acts as a resonant filter for every file encoded in the MP3 format.
Through a simple accusation of racial bias, the author seeks to undermine the technical work that went into the development of the MP3 codec.
I'm always surprised when pointing out that something has design limits, and often limits imposed by racial bias in testing, is seen as such an evil thing. Acknowledging the limits is the first part of making use of something properly, wouldn't you say? Understanding that MP3 is likely (Although it would have to be confirmed via testing) to have poorer encoding on non-white artforms is important to choosing a media encoding, no?
I'm not sure how acknowledging the limits they created is "undermining the work that went into it", can you explain why you feel that is the case?
You Americans (I'm assuming) are really strange. Dividing the world into white and non-white art forms, white and non-white "ways of knowing" etc. is exactly how the Nazis and fascists thought here in Europe. You are going right back into the same divisive state of affairs.
> exactly how the Nazis and fascists thought here in Europe
Yes, we know. That's why we wound up with an orange man who didn't really care about race.
> You are going right back into the same divisive state of affairs.
Division allows the elites to control the masses more efficiently. It's like "divided we fall" might be true or something.
Prepare to see the division tear the United States in half. That's the goal, anyway.
That's amusing, because most of the people of colour I follow on the Fediverse give class-based intersectional analyses of political happenings. In most leftist circles, identity politics has been seamlessly merged with class politics in an effort to ensure that class politics serves the purpose of uplifting all lower class people, rather than just the white lower classes, and in an effort to eliminate discrimination and mistreatment of all kinds. This is the purpose of any form of left-wing ideology.
Indeed, Tony Cliff wrote on the topic of class and intersectionality back in the 1970s*, it isn't a new thing -- certainly not new to academic socialist circles. It is simply only recently that it has entered the majority consciousness. One can see that intersectional analysis does not conflict with class analysis, but compliments it and allows one better lenses by which to view the various conflicts and struggles in the modern world (or indeed, many of the struggles of the 20th century). See the unity he seeks by adding the lense of intersectionality, and see the error in your assessment that "Intersectional Analysis" is somehow adverse to and conflicts with "Class Analysis", I beg of you.
* - https://www.marxists.org/archive/cliff/works/1978/08/gays.ht...
Because of the cultural split, enforced by said systemic inequality, there is a cultural difference. It would be foolish and ignorant to ignore that difference and claim that it doesn't exist. Indeed, "Colour Blindness" as you are proposing, has been outed as racist ideology for over half a century (Jane Elliot explains this rather well). By ignoring the fact that people have different skin tone, and have been disadvantaged by it, you are choosing to ignore their very existence. If the entire human race were "homogenized" there would be much of value that is lost. Likewise, when we talk in terms of "white people" and "people of colour", we are simply recognizing and acknowledging the split that comes from that disadvantage that still exists -- in terms of the original argument, we are recognizing that this specific encoding of sound information is likely biased to encode white culture, white art, better. That is a difference that should be corrected, and to correct it we need to acknowledge it first.
In summary, I'm simply following the terminology given to me by people of colour, who often talk in these terms. The best I can do is follow their example and speak out in the manner that they have taught, or (better still) promote their voices on these matters.
Of course mp3 is limited by the people who made it, everything is. It certainly wasn't done with any bad intentions.
> Its biases are simply acknowledged and plainly described.
Not at all. The author attempts to show the residue of white/pink/brown noise and of a couple of famous test clips (the ones that could be labelled as European). The author makes no attempt to compare or show that MP3 has worse artifacts on non-European music. Great job there. That's why I find the racial accusation to be out of place.
I should note that I've done subtractive analysis of MP3 and Vorbis before, so I've seen the results of my test samples as well. Briefly, I found that MP3 has more tonal residue, whereas Vorbis sounds more like noise. The post-masking residue on Vorbis is stronger.
Never mind that the Chinese still have not completely made up their minds about exactly how to represent their language in digital form. It is ludicrous to think that in the 1960s they could have somehow anticipated how Chinese people wanted to communicate in digital form, and just have support for that. Essentially invent Unicode first.
Of course the authors don't understand any of that, they just saw that there were limits of how Chinese characters could be represented digitally (this was in 2012) and they took an ignorant and antagonistic tone, and threw out accusations of a racist conspiracy by ISO and ASCII. And this paper was part of a curriculum at a major university.
Accusations like this are almost always wrong, and actually damaging. In today's climate, it is especially damaging to throw out accusations of racism.
I can not imagine how the designers of MP3 could have satisfied this person. The had to start somewhere. They had to start with what they knew, and work their way out, which they did.
Text started out being only upper case only, and then lowercase and extended western characters, and then more complexity was added over the years. Photo codecs started out calibrated for Lenna[1], so that they would have a common reference, and worked from there. Were they supposed to be able to perfectly capture every skin tone, all at once, the first time? Should they have started with a picture of a non-white person first? What ethnicity then?
This kind of antagonistic critical mentality is toxic, and should just be dismissed without consideration. At this point, if someone says that some technology is racist, I assume they are a grifter and ignore them.
Your comment was good until you got into an unfounded and reactionary "reverse racism" accusation.
(The Wikipedia article on Tom's Diner says anachronistically that the quote appeared in Business 2.0 magazine. This has been present on the page since the very first revision [2].)
[1] http://web.archive.org/web/20001003052745/http://www.ecompan...
[2] https://en.wikipedia.org/w/index.php?title=Tom%27s_Diner&old...
Even at the time of writing of the article, MP3 was being replaced with much better codecs. It could be said that when the codec was being developed in the 90s, they chose a limited set of music to make the problem easier to solve - which is fair given that decoding was at the limits of processing power at the time. But it's likely they stuck with music on the chromatic scale, and instruments with unusual timbres and microtonal variations, where exceptions to this were usually outside the western canon, though I think that has changed since the 90s with the advent of more sophisticated electronic music and more experimental artists.
I found the report on the listening test, but I could not find a list of the 20-second excerpts they used:
https://sound.media.mit.edu/resources/mpeg4/audio/public/w14...
However like Jpeg this is a tradeoff that we are used to. in the same way that tape has frequency inbalance, records just sound shit (sorry I know you like it, but they really don't sound better. they are great for evoking a feeling, but not for fidelity. Its lomo but for sound)
I don't like the distortion that 128k mp3 introduces, but I imagine that in years to come it'll become a fetish for certain types of audiophile
This hasn't happened in almost 30 years of mp3s being a household format. The only feeling low resolution media evokes is frustration.
This is basically what's happened with vinyl records at this point, right? People have a nostalgia for their sound, even though by any scientific or technical measure a vinyl record represents music/sound less faithfully than, for example, lossless CD audio...
While we now have better technology, people prefer what they are familiar with.
Shrill low-bitrate, lower-fidelity digital sound just doesn't evoke that feeling in us. The same way no one is asking for RealMedia videos on their home theater system, but prefer film grain on video and 3D CGI to look "messy" (another analog noise characteristic). We humans love noise apparently! I guess evolution built us for a noisy and messy world and digital sound just seems off to us. We'll make an effort to get closer to analog warmth but do out best to flee digital shrill.
I remember the first time I saw 60hz video on a flat screen TV. I thought the motion looked so incredibly unnatural it was almost unwatchable. Now, it just seems normal. It had nothing to do with 60hz video itself - it had to do with the expectations built up by years of watching 29.97hz interlaced NTSC video...
CDs and/or Streaming services are victims of the "loudness" wars, resulting in music with low dynamic range. Since vinyl is often mastered by someone who specializes in the vinyl and is rather niche, they have full freedom to make it sound as good as it ever will be.
That's also one reason why some people prefer Tidal or <insert HiFi music services of your choice> because sometimes they simply use better masters. At least, that's the only reason that can't be chalked up to placebo
I'd never deliberately add tape hiss to a track, but when I hear a song with that hiss it just takes me back.
Vinyl does deliver that Crashhhhh, however noisy it may be.
I have noticed people vary a lot in both their sensitivity and preferences.
also https://news.ycombinator.com/item?id=9049196
and then in 2017 https://news.ycombinator.com/item?id=16034547
Apple generally has a very strong case of NIH.
MP3, FLAC (and I assume AAC) requires floating point math to decode.
https://xiph.org/flac/changelog.html
Come to think of it, that still might have been one of the historic reasons for ALAC's creation. Its development started before.
For listening purposes lossy codecs are all capable of being perceptually lossless for trained or untrained listeners, including MP3. MP3 gets a bad rap because low bitrate mp3 does sound awful. 320kbps or higher is sufficient for 99% of listeners, including audiophiles. Most listeners can't pass a double blind study in pristine conditions, let alone average ones.
It's a lot like 192kHz 32 bit recordings. 96kHz audio bandwidth and almost 200dB of SNR has a purpose, but that purpose is not for day to day listening. Your ears, brain, speakers and amplifiers cannot reproduce or perceive the extra information.
It's also the case that MP3 encoders got better over time, and if you're considering the sound of a 128kbps (or 96kbps) MP3 encoded by the Fraunhofer encoder you got from IRC in the late 90s; it's just not going to be good.
Back before mp3s became ubiquitous, you could rip a CD to .wav files and it would basically sound identical to what was on the CD. But then you’d also be using some 600MB versus the same CD sounding not quite as good but only using like 40MB and at that time that was a lot of disk space to use.
Also, if you want music that was recorded in 5.1 surround sound, sometimes the surround version is only available from lossless sources, while lossy sources sell only the stereo downmix.
> sometimes the lossy versions of downloads have "loudness wars" compressed dynamics
The loudness war has been over for years, and it was precisely due to lossless CD audio that it existed in the first place.
Have you seen the size of disks and network bandwidth recently? There is absolutely no reason to bother with lossy audio compression.
They never supported Ogg Vorbis, and they don't support it's successor Opus now (except in Safari where is required as part of the WebRTC spec.)
Fortunately, being lossless it's safe and easy to convert between FLAC and ALAC (or vice versa). FFmpeg can do it, and also supports Apples "CAF" (Core Audio Format) as an intermediate format if required (e.g piping between different encoders) - that's handy for keeping metadata intact between formats.
https://forceincmilleplateaux.bandcamp.com/album/most-beauti...
Using that as a starting point, I found this:
https://web.stanford.edu/~yzliao/pub/master_thesis.pdf
>"The estimation of the frequency and phase of a complex exponential in additive white Gaussian noise (AWGN) is a fundamental and well-studied problem in signal processing and communications. Its numerous applications include carrier recovery in a communication system [1], determination of the object position in radar and sonar systems [2, 3], estimation of the heart rate of a fetus in biomedicine [4], and carrier synchronization in a distributed beamforming system [5]. Regardless of the application, poor estimation can lead to disastrous results. For example, in communication system, with the poor carrier frequency estimate, the down-converter may not be able to demodulate the passband signal to baseband [1]. In the smart antenna system and speech processing system, a poor phase estimator may cause the system to fail to identify the direction of arrival of the signal [6, 7]."
Was thinking in the shower just the other day about genre-specific compression schemes. Could you get significant improvement if you knew something like the BPM or the spectral profile of common instruments ahead of time? Or is production too inherently complex for this to be worthwhile? Obviously no two hurdy gurdy recordings will sound the same, even though your algorithm knows wtf a hurdy gurdy is.
White noise could be represented by a grid with frequency. Or perhaps use vectors with less frequency precision. That way the same mechanism can be used for both types, but with frequency re-scaled for noise to reduce data size. A given noise vector would represent an approximate frequency range rather than a specific frequency.
A frequency vector would probably need about a quarter note step in precision for a typical compression level (but adjustable). If it allows interpolation between vector nodes[1], then it can handle frequency tremolo fairly well. A vector node can have flags saying whether to interpolate frequency and/or volume, or just do a direct step jump, which ever best fits the original. (Noise vectors can also have interpolation flags.)
The noise vectors may only need a half-octave or octave range of precision. I'm guestimating only 4 to 6 bits are needed for the default precision, whereas the frequency vectors need around 11 bits.
As far as how to process the sound to produce such vectors, the encoder would have to "look" at the time/frequency plot, and divide artifacts (areas) into frequency "lines" and noise lines. I'm not an expert on such algorithms, but it probably can be refined from experimentation. Maybe make a rough guess start, and use a genetic algorithm to breed a vector set that best recreates the original, given a data size constraint.
If there's a lot of sound going on in one spot, then precision can perhaps be reduced for that spot. Human ears can't typically isolate details when the rock band is going all out, for example. Maybe give more precision to the loudest vectors, but slack on the lessor ones to save space.
[1] I'm thinking of segmented lines. Each segment node would contain a frequency value, volume value, frequency interpolation flag, and a volume interpolation flag. The range and meaning of each value would depend on the type (voice vs. noise) and precision specified for that line. "Voice" vector may be a better term than "frequency" vector.
Does anyone have a full list of the recordings they used for these listening tests?
For example:
Audio Clip 1 - this is what I might first think of as random noise.
Audio Clip 2 - is this also random noise or is there some kind of pattern to this?
TLDR: If you need to go lossy, use Opus. At higher bitrates, it's the only lossy codec that has a flat frequency response.
The closest thing to an objective analysis tool is PEAQ which is also pretty terrible, but its universally terrible so that makes it useful as a benchmark.
That's about as objective as you can get.