... I have no idea how to do that. Even a toy one.
... I have no idea how to do that. Even a toy one.
Source code for a lot of early encoders is available to study. Probably easier to understand than the modern implementations. LAME started out as a set of performance patches against the ISO sources until eventually being rewritten from scratch.
Toy MP3 encoders are not horrifically complicated, though ones that sound decent are. It's astonishing how much improvement there's been over the past few decades even with the same underlying format, modern 128kbps MP3s don't even make your ears bleed.
1. Download liblame
2. Thank god and Richard Stallman for open source software
3. Have a cup of coffee.
The psychoacoustic model took more time on its own to code than the PCM splitting, MDCT, windowing and Huffman coding.
It’s a fun project however painful. There’s a MP3 encoder from the early 90s floating around that you can use as a base if you can fix the legacy code. I can’t remember the name though, sorry.
(Coincidentally, these are the steps detailed in the article, in the opposite order.)
... snark, but not really? Todays hammer of choice is machine learning and AI.
It wouldn't be able to measure "warmth" or "heart" or "anger" or anything like that. It would be able to tell you which model most closely matches the original at the frequencies you can hear. Once you have a metric, you can use it as the basis of a parameter space search. Not exactly machine learning, but maybe some preliminaries.