AudioFlux: A C/C++ library for audio and music analysis
github.com
github.com
- https://github.com/marsyas/marsyas
- https://github.com/ircam-ismm/pipo
- https://github.com/flucoma/flucoma-core/tree/main/include/al...
What kind of music are you trying to transcribe?
Feel free to email me.
https://github.com/Music-and-Culture-Technology-Lab/omnizart and https://basicpitch.spotify.com/
They work better if you apply some source separation before (e.g, https://github.com/sigsep/open-unmix-pytorch, https://github.com/facebookresearch/demucs, or https://mvsep.com)
Still, I think the best results are from proprietary models (specifically https://www.ableton.com/en/manual/converting-audio-to-midi/ and https://www.celemony.com/en/melodyne/what-is-melodyne)
1) Tuning hyperparameters of your audio preprocessing is a pain if it's a preprocessed CPU step. You have to redo preprocessing every time you want to tune your audio feature hyperparams
2) It's quite common to use torchaudio spectrograms, etc. purely because they are faster (I can link to a handful of recent high-impact audio ML github repos if you like)
3) If you use nnAudio, you can actually backprop the STFT or mel filters and tune them if you like. With that said, this is not so commonplace.
4) Sometimes the audio is GENERATED by a GPU. For example, in a neural vocoder, you decode the audio from a mel to a waveform. Then, you compute the loss over the true versus predict audio mel spectrograms. You can't do this with these C++ features. (Again, I can link a handful of recent high-impact audio ML github repos if you like.)
Again, I just don't get it.
Yes please :D
https://github.com/descriptinc/descript-audio-codec/blob/mai...
https://github.com/NVIDIA/BigVGAN/blob/main/loss.py#L23
https://arxiv.org/pdf/2210.13438 (the github repo doesn't include training, just inference)
It is INCREDIBLY common to use multi-scale spectral loss as the audio distance / objective measure in audio generation. They have some issues (i.e. they aren't always well correlated with human perception) but they are the known-current-best.
Anyway, in audio ML what is very common is:
a) Futzing with the way you do feature extraction on the input. (Oh, maybe I want CQT for this task or a different scale Mel etc)
b) Doing feature extraction on generated audio output, and constructing loss functions from generated audio features.
So, as I said, I don't exactly see the utility of this library for deep learning.
With that said, it is definitely nice to have really high speed low latency audio algorithms in C++. I just wouldn't market it as "useful for deep learning" because
a) during training, you need more flexibility than non-GPU methods without backprop
b) if you are doing "deep learning" then your inferred model will presumably be quite large, and there will be a million other things you'll need to optimize to get real-time inference or inference on CPUs to work well.
Is just my gut reaction. It seems like a solid project, I just question the one selling point of "useful for deep learning" that's all.
Can you start by suggesting what you task you want to do? I'll throw out some suggestions, but you can say something different. Also you are welcome to email me (email in HN profile):
* Voice conversion / singing voice conversion
* Transcription of audio to MIDI
* Classification / tagging of audio scene
* Applying some effect / cleanup to audio
* Separating audio into different instruments
etc
The really quick summary of audio ML as a topic is:
* Often people treat it audio ML as vision ML, by using spectrogram representations of audio. Nonetheless, 1D models are sometimes just as good if not better, but they require very specific familiarity with the audio domain.
* Audio distance measures (loss functions) are pretty crappy and not well-correlated with human perception. You can say the same thing about vision distance measures, but a lot more research has gone into vision models so we have better heuristics around vision stuff. With that said, multi-scale log mel spectrogram isn't that terrible.
* Audio has a handful of little gotches around padding, windowing, etc.
* DSP is a black art and DSP knowledge has high ROI versus just being dumb and black boxy about everything.
The point is, ship it.
Seriously, nobody is lugging a GPU around to interact with their most frequently used micro-computing platform, their headphones, which right now, already represent a new and extraordinary era of "accelerated component" market expansion.
The 7 microphones in your earpiece, and the 6 speakers pushing air into your head, are not quite as close to the GPU, as they need to be, perhaps .. but they already have a DSP, and there is already a silicon battle going on among the vendors.
>You can't do this with these C++ features.
Yes, and I think the point in the end, is to use AI to write better C++ code, and design better, cheaper, smarter silicon, as always (and actually ship it) ..
Not to say that big-AI shouldn't have audio analysis as a compelling sphere of application, but more that, until the chips arrive, in-ear AI is less of a specification/requirement, than in-ear DSP.
We don't need AI to isolate discrete audio components and do things with them, in-Ear. Offline/big-AI, however, is still compelling. But we don't yet have GPU neckbands ..
/extremelypedantic
To make it easier for those skipping English classes
"A forward dash can be used to state alternatives. A sentence that uses a forward slash in this way can be read to mean that any or all of the stated words could apply."
https://www.thesaurus.com/e/grammar/slash/
And from the world of reference C and C++.
"The C/C++ Users Journal"
https://en.wikipedia.org/wiki/C/C%2B%2B_Users_Journal
"Visual Studio C/C++ IDE and Compiler for Windows"
https://visualstudio.microsoft.com/vs/features/cplusplus/
Random job post from Microsoft,
"Perform software development in C/C++, Python, and other languages."
https://jobs.careers.microsoft.com/global/en/job/1752991/Pri...
Random job post from Apple,
"Develop/maintain bit-accurate function C/C++ model for hardware verification
Develop/maintain cycle-approximate perf C/C++ model for performance analysis - Analyze model
Excellent C/C++ programming skills"
https://jobs.apple.com/en-us/details/200448639/graphics-mode...
Random job post from Google,
"4 years of experience coding with one or more programming languages (e.g., Java, C/C++, Python)"
https://www.google.com/about/careers/applications/jobs/resul...
Random job post from NVidia,
"Strong C/C++ programming skills"
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCar...
More examples from WG14, WG21 members, C and C++ compiler vendors can be provided.
Yes, Visual Studio supports both C and C++, but those are, in fact, two different languages.
You'll struggle to find the links you promised at the end because C and C++ are run by two different groups, meaning you won't be able to link us to single sources.
From Herb Sutter, a name that you might know what relevance it has for WG21, I hope.
"Keynote: Safety, Security, Safety and C / C++ - C++ Evolution"
A library is one or the other. Talking about safe systems languages where C and C++ share memory safety issues is very different from promoting a library. Thanks for the opportunity to clarify here. You’re confusing context with lack of technical precision.
We both made it quite clear where we stand, so there is hardly any value pointing out uses of C/C++ expression by other key WG14 and WG21 members, papers or products.
"C/C++" has meaning in some contexts and reveals ignorance when used out of context. The post title here uses it incorrectly, but yes, there are ways to use it correctly. We disagree on that because you can't tell the difference in the two. So it is, but the actual explanation of usage is there for others who do care if they are perceived as non-technical in technical environments.
Definitely doesn't count as _lying_, but still underwhelming.
not any C, only the C++-compatible subset.
int* foo = malloc(sizeof(int));
has never worked in C++ for instance while it's valid C. Code that worked is code that people actually did effort to express in a way compatible with a C++ compiler.As far as compatibility and “history” the languages are different enough now. There are both: features in C that do not exist in C++, and code that is conforming C that would be UB in C++. Saying C/C++ (for real) is usually a dumb target when it’s better to pick one and settle with that.
If it’s C, just say so. Everyone knows what extern C is, you don’t need to confuse.
#ifdef _cplusplus
#include <iostream>
#define print() int main(){cout << "Hello world! -- from C++" << endl;}
#elif (defined __STDC__) || (defined __STDC_VERSION__)
#include <stdio.h>
#define print() int main(){printf("Hello world! -- from C\n");}
#else
import builtins
print = lambda : builtins.print("Hello world! -- from Python")
#endif
print()
Some python code works in C and C++ as well but people don't group them together and call Python/C/C++