Whipper: Accurate Audio CD Ripping
github.com
github.com
I guess it's great to have a *nix tool for the rare occasion of ripping a CD.
More details here: https://wiki.hydrogenaud.io/index.php?title=Whipper
Of course, those are mostly oddities that virtually no one cares about in a sonic sense, but it’s still a piece of musical history.
Kind of makes me wish whoever sued what.cd offered them a sort of clemency deal: keep the site up until everything has been archived by a legit third party (that wouldn’t share it, but so everything has at least been preserved), then shut it down.
Instead they burned down the musical library of Alexandria..
I'm mostly thinking of those interstitial lead-in thingies on live and concept albums. E.g., if you let a CD play from track 1 to track 2, there may be 5 seconds of crowd noise leading in to track 2. CD players would show this as track 2 at timing "-5", then count down to zero to get to the actual beginning of track 2.
If the user skipped directly to track two, this interstitial thingy wouldn't be played.
Edit: wording
Also: same question for what.cd ripping process. IIRC they broke the tracks down into different files.
It pretty much mirrors actual CD structure of TOC (containing all necessary timestamps and metadata) and continuous sequence of frames with audio data
I was curious about this as well, and the answer was "no". I meticulously followed EAC setup guides for three drives in EAC, I used the recommended gap settings, the results were completely verified by AccurateRip, I was storing the results as a CUE sheet and single WAVE file, all drives would produce the same file, I was using EAC to burn the CUE sheet and WAVE file back to new Verbatim CD-Rs, re-ripping was done with the same setup in EAC, and the files still didn't match. I've been meaning to dig deeper and compare the files in binary, but haven't gotten around to it.
I now rip each of my CDs to a single FLAC and cue sheet (using XLD on Mac), on the theory that it's the most accurate way to archive a disc. However, I haven't found anything that offers the same accessibility to such a collection as iTunes did. I look through folders, and drag a few cue sheets at a time into Foobar 2000.
Any tips?
While I rip all tracks separately the FAQ states: „Swinsian also supports albums ripped as a single file together with a cue file, and FLAC, Ogg Vorbis and WavPack files with embedded cue information„.
There is a free trial, maybe this works for you.
I don't know where they're located but a refurbished Mac mini from Apple is $589. AWS EC2 M1 Mac instances are ~$16/day ($0.65/hour, 24-hour minimum).
https://www.apple.com/shop/product/FGNR3LL/A/refurbished-mac...
https://aws.amazon.com/about-aws/whats-new/2022/07/general-a...
https://aws.amazon.com/ec2/dedicated-hosts/pricing/#On-Deman...
It seems like a lot of extra steps to play music but when my library started getting huge it was becoming a pain to organize and sync between devices.
I won't argue with this but I think there needs to be a qualifier - the most accurate way to archive a disc while using compression ...
Given that an audio cd is encoded in WAV/PCM and we have a WAV file specification, I think ripping to a WAV file remains the gold standard:
"... digital audio extraction software can be used to read CD-DA audio data and store it in files. Common audio file formats for this purpose include WAV and AIFF, which simply preface the LPCM data with a short header ..." [1]
I like the idea that ripping to a WAV file means I never have to rip that disc again.
[1] https://en.wikipedia.org/wiki/Compact_Disc_Digital_Audio
Edit: I only mention this because I can't think why the extra "while using compression" qualifier is relevant except for information loss during compression.
No, but if that is the definition being used, neither does WAV store that data either.
> The audio will be preserved but there might be interesting details on the CD like easter eggs, copyright text, etc. and I'm not sure FLAC would capture those.
Flac does allow for storing a lot of metadata in the flac file itself, including: CUE sheet, a picture (image file), and arbitrary tags (key value pairs). So much of those could be stored in the flac file, but obtaining them (Easter eggs and the like) is dependent upon the program ripping the CD, not the file format storing the result of the rip (flac does not do CD ripping itself).
I meant versus a raw binary dump of the CD.
This was done to both increase play time and discourage copying.
So since you’ll never get the raw redbook data off a cd, there is no reason to prefer WAV over FLAC for CD audio.
I doubt copying figured into the decision process (at least not a deliberate "anti-piracy" thought process).
The CD was released in 1982 [1]. This was only one year after IBM had released the original IBM PC, which came with one or two 320KB 5.25 inch floppy drives. The IBM PC did not have an official "hard drive" variant from IBM until the PC XT [3] released in 1983 (one year after the CD) and the hard drive option was a whole 10MB. When the target was replacement of the Vinyl LP player and compact cassette deck, and when their target customer base might have had 10-20MB of total storage in their personal computers (if they even had a personal computer), it would seem unbelievable to Sony and Phillips that those customers might even be capable of copying a disk holding somewhere upwards of 850-950MB of digital audio data.
And of course in the above I am overlooking that in order to "copy a CD" (as in make a digital copy) one first requires a CD drive that can interface with a PC in a digital way, and such drives did not appear until some years after the CD's release.
So with an amount of data on each disk that exceeds end users available storage by orders of magnitude, and no "CD drives interfaced digitally to computers" at the same time it seems unlikely that Sony or Phillips ever even considered that end users would be able to digitally copy CD disks.
[1] https://en.wikipedia.org/wiki/CD_audio
So if you really want to save an exact copy of the CD, you will have to find something better than wav to save the data. There are a few programs that can do an image to ISO format, that might be what you want. I guess it's a bit harder to find a player that can handle the audio tracks from an ISO though.
(Memory is a bit vague on formats and names but this is the general idea.)
It’s expensive but Roon does this at a level far beyond iTunes: https://roonlabs.com/
Using whole files rather than individual tracks is just asking for problems with mainstream players in my opinion.
However if you're serious about your music library I recommend Roon. Not cheap but a great solution.
Otherwise, I prefer methods that get me CD->FLAC transfers in a few minutes.
It is open source now just like FLAC, and actually has built-in support in Windows and many Linux music players despite being an Apple format, making it somewhat more universally compatible than FLAC (which isn’t supported in iOS or macOS).
> something like Windows Media Player to cover all that much either
Well, in Windows 11, Windows Media Player does support ALAC. Heck, it will even rip your CDs and encode them in ALAC.
Also I could be wrong, but don't most trackers still accept rips that they consider inferior, but then they could be trumped by someone else that made an EAC rip?
Or XLD: https://web.archive.org/web/20161123023617/http://whatcdinfo...
And it's now packaged with many Linux distro. At some point it wasn't packaged with Debian and I simply couldn't get it to build so... For a while I ran Fedora on an old PC only to get whipper! Nowadays whipper* ships on even Debian / Devuan so life is good.
# $1 = output file
# /dev/sr0 is assumed
cdrdao disk-info
cdrdao read-cd --with-cddb --device /dev/sr0 --read-raw --datafile $1.cdrdao.bin $1.toc
dd conv=swab if=$1.cdrdao.bin of=$1.bin
# To play the BIN: mplayer -demuxer rawaudio BIN"
I'm not an "audiophile" (perfectionist seeker of equipment) but I enjoy listening to classical and (some) jazz compositions and performances as they were conceived, not as "tracks".Extract data to wav:
cdparanoia -B
Compress all wavs to mp3 (in parallel, takes like 1 second, crazy!): parallel lame --preset extreme {} -o {.}.mp3 ::: *.wav
Clean up folder rm *.wav
Import into mopidy library and lookup/apply metadata with beets beet import .
sudo su - mopidy -s /bin/bash -c 'mopidy --config /etc/mopidy/mopidy.conf local scan'
(I actually have a bash script for that, obviously)Done. Next step is to go into Mopidy-Iris UI and hit play and it plays all around the house thanks to snapcast. Garsh darn glorious.
I honestly don't understand why mp3 is such a sticky codec. It's long been surpassed by others.
I challenge everyone who feels strongly about this to actually bust out abx and do some listening tests comparing flac to LAME-encoded extreme MP3s and prove to yourself that flac matters at all on your equipment. And then share your results with us if you want!
https://manpages.ubuntu.com/manpages/xenial/man1/abx.1.html#...
https://en.wikipedia.org/wiki/ABX_test
But yeah I should put the command to encode to flac there for people who prefer it!
But generally I agree with you, if you're happy with how your setup works and sounds, MP3 is no great sin.
As for what the commenter said, firstly, MP3 was synonymous with digital music in most of our minds for so long; second, even if modern equipment handles all the better codecs, a lot of us still have memories of times it didn't in the past. Third, I think some of the tooling around metadata is not as developed or ubiquitous for, say, an ogg container. There are ogg comments, but ID3 is better supported.
I think this is probably the main reason for mp3's stickiness.
> second, even if modern equipment handles all the better codecs, a lot of us still have memories of times it didn't in the past.
Sure, and your dvd players didn't always handle x264. Things change :). It's been probably 10 years since new audio hardware had trouble with non-mp3 media.
> Third, I think some of the tooling around metadata is not as developed or ubiquitous for, say, an ogg container. There are ogg comments, but ID3 is better supported.
Granted, I think ogg will have worse support for equipment. However, I'd expect that aac in m4a will end up with the same level of support as mp3s do today simply because it's a lot more common than opus (and older).
> MP3 is no great sin.
It's not, I just don't like seeing generally superior tech getting sidelined because the inferior tech is more familiar. Perhaps that's a sin on my part :). I admit it probably doesn't ultimately matter if your music is 1MB vs 2MB.
(After typing that, I realize that calling it a "black box" is a pretty good pun on MP4 box formats..)
If you aren't trying to save bits, then why use a lossy codec at all?
That's more my point. MP3s will be smaller than flac, for sure, but opus or he-aac files can be even smaller (half the size or more).
I still encode to plain old 128k mp3. Here's why ...
First, I keep the WAV originals and those are what I listen to on my music system, in my music server, etc.
Second, if I am using the mp3s it is because I am going to some unknown place to interface with some unknown tool to play these - let's just make life simple and use something that will work everywhere - even the dumb creative audio bluetooth adapter that was in that airbnb that one time ...
Finally, 128k mp3 is typically a 10:1 compression ratio and makes size and space "budgeting" easy. It's easy to remember.
One other thing:
When I export my ripped CD wav collection to mp3 I also compress the filenames - I flatten to ASCII 256 and truncate the filenames to 64 characters, etc. LOTS of car audio interfaces just puke when they hit weird unicode characters or can't display long filenames ... it creates all manner of havoc.
FLAC is well-supported. But, PCM RIFF WAV files are triv-i-al. Any CS101 student can write a parser. They don’t need decoding.
2:1 lossless compression is nice. It’s quite a technical feat.
Meanwhile, as an old software engineer, I focus pretty hard on technical simplicity. And, the complexity ratio of PCM vs. pretty much else everything is a a very, very small number. Very small :p
As I understand it cdparanoia does not such check. It's "paranoid" by its own but you're not verifying that your rip is 100% perfect.
But then if you then convert it to mp3 and discard the lossless files anyway, I take it you don't care much about data integrity.
whipper and EAC do serve another purpose than cdparanoia (I think, btw, that whipper is something uses cdparanoia under the hood) and I'm pretty sure that people who do care about bitperfect rips do not then go and convert their files to mp3s.
FWIW a .flac file (now a .wav but a .flac: that is lossless and compressed) is about twice the size of a 320 kbps mp3 I think.
Seen that songs are tinier than tiny files compared to modern standards I'm totally fine playing the .flac files I ripped using whipper.
It's both my audio "source" and my backups.
It’s funny how a CD is more like a tape with contiguous content and an index, rather than a file system.
https://en.wikipedia.org/wiki/List_of_albums_with_tracks_hid...
I bought hundreds of CDs from 1990 through to about 2005, some of them on that list, and I never had a clue about this technique.
I knew about hidden tracks at the end after a silent gap - e.g. Nevermind had a hidden track after 10 minutes of silence at the end, was fun to schedule it on a CD jukebox in a bar...
I own music only discs in each of these formats. I think the last Blu-Ray Audio disk I bought (Yello Point Dolby Atmos release) was shipped in 2021.
Granted, they are fairly sparse. That disc was an import.
Granted, it's not perfect...
To compensate for this, data CDs include intermittent data on the spiral that says "you are here", but audio CDs don't do this. The only way to ensure that you haven't either skipped over a piece, or duplicated a piece, is to rip the CD in overlapping chunks, and then compare the overlaps.
Edit: honestly I don't get it. Do you export all raw images from your camera to PNG too? JPG is good enough, and single-pass CD rips are just fine too.
No, i keep them in RAW and develop a few for albums and print and such.. No, JPGs and PNGs are not good enough, the loss is REAL when you try to actually develop the picture.
I must admit I _REALLY_ don't get this "good enough" mentality.. Like.. Here's this way to get a more precise representation of the music you bought, and you'll go "nah, that's not for me, I prefer whatever random data I happen to get on the first go".. for what? why ?
If you really want to have a different product on every playback, I guess vinyl is the way to go :P
The purpose of music is to be perceived. There is no perceptible difference, therefore there is no difference. CD Paranoia/etc are a waste of time, these tools make ripping a CD take several times as long for no tangible benefit.
This is often not the case. We're not talking about subtle interpolation differences. Depending on the drive, and the disc, bad rips can be full of music skips, clicks, and pops. It's really annoying to find this out days later after you ripped something. Better to get it right first time.
Data CDs have an extra layer of error correction. Audio CDs have less error correction because small bit errors are not a big deal for your listening experience. Most CD players quietly interpolate over small errors in a way that you probably don't even notice.
But of course it makes a difference if you want to make high-fidelity copy.
That's why there always have been copy apps (cdparanoia on Linux, EAC on Windows) that do overlapping reads, and then detect and compensate these problems.
You do realise, assuming you ripped to a lossless audio format, your cd rips are 8-12x more accurate than anything streamed of Spotify?
AccurateRip is a pretty important feature imo. Like at least I know someone else in the world made a rip which was exactly the same as mine regardless of the optical drive we used. Overkill? Maybe but there is some kind of "safeness" to it
How did you calculate a factor by which lossless is better than lossy?
>...How did you calculate a factor by which lossless is better than lossy?
I specifically typed accurate, not better.
>...How did you calculate a factor by which lossless is better than lossy?
>I specifically typed accurate, not better.
Yeah OK, that's just wording. Which two factors did you specifically compare, that gave you the 8-12x figure?
Maybe in theory. In practice, the difference will be very hard to hear (for most people, in most scenarios). Have you ever done an ABX test to determine whether you would be able to tell the difference between lossless and Spotify's quality?[1]
I did, a while back, and while I was able to tell the difference somewhat reliably for music I knew well (and only for that kind of music), the effort and time I had to spend on finding the minute differences, even with high-end equipment, convinced me that for everyday listening, Spotify was completely fine for me.
One example mentioned was that drives generally don’t read the precise samples you ask for. https://web.archive.org/web/20160528213242/https://thomas.ap...
It's not meant for copy Audio CDs but CD-ROMs
That's why CDRDAO or CDDA Paranoia exist for decades.
https://web.archive.org/web/20160528213242/https://thomas.ap...
Is this false?
I have also noticed that, sometimes, ddrescue has trouble ripping dvds, chocking out or producing a low quality rip, that will play perfectly on the same hardware with VLC
User friendly is a matter of opinion.
- Detects and rips _non digitally silent_ [Hidden Track One Audio](http://wiki.hydrogenaud.io/index.php?title=HTOA) (HTOA)
It's unclear how this is used -- is it for secret data or is it for hidden music?
[1]: https://en.wikipedia.org/wiki/List_of_albums_with_tracks_hid... [2]: https://en.wikipedia.org/wiki/Hidden_track
* Rips tracks twice and ensures the checksums match (retrying if they do not)
* Compares the track checksums against the AccurateRip database
I think this gives much more confidence that the result is correct. The tags seem a lot more detailed with whipper too.
It's a very nice workflow and avoids a lot of cleanup ...
I like to think it's a trend?
Been using Asunder to rip, encoding as FLAC with max compression.
Playing back usually with Amarok but it's buggy for me these days so I'm not satisfied. Rhythmbox is good, too, but failed me by not inhibiting sleep during playback. I don't have a really solid go-to music organization / playback app right now but open to suggestions.
Also mentioned elsewhere in this thread: https://pilabor.com/blog/2022/10/audio-cd-ripping-hardware/
Then just run whipper with: `whipper cd rip -L eac`
Maybe it's just me but when I see Docker with a project like this I already zone out. This is still just based on cdparanoia/cdrdao [0]. I wish people would just push single portable binaries instead of starting the whole process with Docker (especially when the source code is alreaedy available)
0, https://github.com/whipper-team/whipper/blob/develop/whipper...
0, https://github.com/whipper-team/whipper/blob/develop/whipper...
It's undeniably useful and pragmatic, but its existence should be a source of shame, a constant reminder that we failed.
As a user/self-administrator, I really do not appreciate a developer throwing a ton of incidental complexity over the wall. Docker is basically a scaling up of the "works on my system" cop out.
I get that there are a lot of oddball distros, but it would seem that a policy of only digging into bug reports on specific well known distros would be more appropriate than basically forgoing the entire concept of a distribution in favor of what are essentially huge static binaries.
It's especially a problem when projects go nuts with this Dockerization anti-pattern, and become actively hostile to distributions shipping plainly administerable versions of their software (looking at you Home Assistant).
As an example alpine has a smaller default stack size. It’s not obvious if that stack overflow is expected or not.
I'm sure Docker's great if you want to guarantee that the thing inside can't access the host system, using all kinds of kernel mechanisms, but that's generally the opposite of what you want for application software.
The main advantage of Docker for application distribution that I see is that it's lazy. You don't have to worry about keeping track of what's required, you don't have to fiddle with paths, you just hack until something works and then ship it. That's fine if you prioritize your time over your user's time.
Still not a single binary but as you note with it being written in python and based on cdparanoia, etc how would that work?
It's based on python with relatively obscure requirements[0] that also calls out to system binaries. Looking at the Dockerfile[1] it is built with specific revs of component software to work around various issues. Take a look at the build docs and you'll see just how many existing projects (python and otherwise) it takes to deliver the end result.
IMO Docker is one of the "best" and most straightforward ways to package up all of this with the end result (as usual) putting any Linux user two commands away from ripping a disc.
[0] - https://github.com/whipper-team/whipper/blob/develop/require...
[1] - https://github.com/whipper-team/whipper/blob/develop/Dockerf...
Cosmopolitan python is a good starting point: https://ahgamut.github.io/2021/07/13/ape-python/
> IMO Docker is one of the "best" and most straightforward ways to package up all of this
No. Using containers has its place. But would you use containers on hello world?
Likewise, if you've got to use a container for cd ripping tool, you've failed somewhere.
This is nothing like hello world.
Full disagree. Ripping a CD in 2023 should be as complicated as hello world.
> The post you're replying to gave several reasons
Let's go through them!
> Packages are available for just about any distro
I use windows.
> It's based on python with relatively obscure requirements
So include the modules along with the cosmopolitan python
> that also calls out to system binaries
Put that logic into the cosmopolitan binary: if Linux then do this, if Windows then do that etc
> Take a look at the build docs and you'll see just how many existing projects (python and otherwise) it takes to deliver the end result.
Do the same to these extra projects.
> most straightforward way
The lazy way? Yes.
Like if you can't be bothered to implement basic features, have a billion dependencies. If as a consequence you code is unstable, put it into a container, and orchestrate.
Here's a simple C equivalent: if you have memory leaks because you know malloc() but don't know about free(), the right solution isn't to kill and respawn when you go above some memory quota, but to learn about free().
This project is at least nine years old with over 1,600 commits. There are 105 open issues.
If you bothered to spend a few minutes to look at the source and open issues you'd realize how complicated and difficult it actually is to enable as many (Linux) users as possible (with cheap, out of spec, and shoddy hardware) to make close-to-perfect rips (with metadata, in multiple formats, etc) of nearly any compact disc produced over the last four decades.
Finally, as someone who has created and contributed to open source projects calling this project "lazy" is downright offensive and completely unfair. Please feel free to spend your personal time and effort to create something better. I'm sure the people who successfully use whipper everyday will be anxiously awaiting and rejoice at the release of your perfect implementation.
Then, when it (never) appears someone like you will be here trashing a design decision or compromise you made. Or, as the saying goes, I'm sure the maintainers of the project would appreciate your pull requests.
Valid criticism and debate is great (and beneficial) but your comment and attitude go way too far.
Please try to put yourself in the shoes of people who donate their time to actually produce something of value and utility (for free) only to have keyboard warriors come out of the woodwork and call them lazy.
Making stuff work together isnt always easy. Having a predictable easy to manage unit can be really nice, offload the particular ecosystem concerns & get folks to "just use it" places.
For most Docker instances I run I have a simple script (or compose file) which does everything, and by reading it I know exactly which configuration tweaks I did and where the data is stored. No more forgetting that one line in some non-obvious file in /etc that made everything work. Backing it all up is trivial as the volume mapping makes it clear where the important data is stored.
Of course, if the project doesn't have any dependencies and doesn't require tweaking /etc settings to work, then sure a single self-sufficient binary would suffice. However this project relies on Python, so that's probably not gonna work.
Do you have a specific project in mind that does DVDs?
Or, if building from source is desired for whatever reason, it should be "Clone the repo. Run `cargo build --release` or w/e.
Snap actually uses the same container technology as Docker.
(Not a fan of Snap, though)