Adventures of developing an MP3 decoder in Python
portalfire.wordpress.com
portalfire.wordpress.com
I initially wrote it in floating point C (compiler translated float to fixed point) and it was damn slow. I then rewrote everything to use fixed point arithmetic only and dropped down to about half a second of decoding for a second of data. That wasn't still good enough because the DSP had to take care some other stuff like ide drives, dma, fat, dac, usb , display drivers etc, so I rewrote all critical parts in assembly taking advantage of the chip's build in dsp functionality and dropped to 100ms of decoding for 1s of data.
So it might be helpful to start from a pure python implementation then move your way down, you'll still learn a lot.
> I finished the decoder and it can produce wav file output from an mp3. The good news is that is works, the bad news is that its 34 times to slow to play the audio runtime.
I think it's great to try learning something in a language that makes experimentation easy. Looking forward to see how he goes with the optimisation.
My hats off to him.
However, I'd really like if at least something that's now a part of Numpy (which I believe is now even more a language extension than a library) would become the part of the Python language.
But numpy brings you abstractions on top of python which are useful and even more high level than python, and that's what is interesting: abstractions which do make your code faster (when applicable, of course).
From the blog's profiling posts, a lot of time seems to be spent into MDCT and polyphase decoding: both could be coded quite efficiently in numpy (I hope to add MDCT itself in scipy, actually).
About being part of python: I don't think it would be a right move (and has been given up from both python and numpy), because numpy is quite big, and depends on quite hairy dependencies if you care about speed. I don't think it is fair to say that numpy is a language extension - numpy feels more natural than say twisted from a python integration POV. But I am a numpy developer, so maybe not the best judge.
Maybe having some default build to have the main functionality and have some way to use the hairiy dependencies otherwise? As you say:
> Numpy brings you abstractions (...) which are useful and even more high level than python,
That's what's a loss for Python as the language. I don't expect to have the access to every FORTRAN-coded scientific library routine in Python by default, but being able to use the right abstractions in the langauge and not through something like the patch... maybe that can happen once?
I understand, of course, that the current state is easier for you and for Python developers.
Still, LuaJIT is easier to obtain and run, as far as I know.
The tests I've done were only inside of Firefox 4 Beta and I've got 3.4 sec for 5M n-body on my slightly (30%) faster computer. I remember that traditionally the standalone builds of Mozilla JS were not giving the same results as the one in the browser.
A message for Mozilla devs, in case anybody reads: be aware that people do try to measure how fast JS is from the command line and let them get the really last results! And use http://shootout.alioth.debian.org if you want to fundamentally improve your implementation of JS.
They can do better than use the benchmarks game - Kraken is one attempt they are making to do better. http://krakenbenchmark.mozilla.com/index.html
http://news.ycombinator.com/item?id=1754881
The benefits of the benchmarks game: allows comparison with other languages (like Lua) whereas as far as I understand their Kraken is JS only and browser only. Thanks a lot for your good work!
Also: you might want to try Cython, it compiles a Python-like language to (reasonably platform-independent) C or C++ code, with very good Python interoperability.
> I’m going to write a python based MP3 decoder. Yeah, its a bad idea from a performance standpoint, but from a learning/documenting perspective, it is quick and easy.