Audio CD ripping – optical drive accuracy listing
pilabor.com
pilabor.com
(HN discussion: https://news.ycombinator.com/item?id=17649374)
I added this to the article along with some other things, I took from this discussion. Thank you HN for providing more details!
Tightly controlling how data can be submitted allows that accuracy to be maintained.
I forgot that CTDB (http://db.cuetools.net/) exists, which is is an alternative to AccurateRip. CUETools is open source Windows software to rip CDs. Instead of just providing a checksum of the track, it provides error correction information. So instead of just getting that you probably have a bad rip, and keep getting a bad rip, it's possible to correct the rip. EAC has a CTDB plugin that's installed by default, whipper currently doesn't support it.
AccurateRip is not something from EAC, it's from dbpoweramp.
I wrote this article mainly for myself... I just converted the dbPowerAmp stats into a somewhat searchable JavaScript table and added some useful notes I found out over the last days. I did not expect HN Top 10 for a whole day with this... :-)
I cannot unhear the Flintstones now...
> plays the right version and track listing
Just try to get ahold of Rust in Peace today. The rerecorded album by Dave's Megadeth cover band sounds terrible, and the original is out of print and not available from any of the legal streaming platforms. I'm considering stocking up on original CDs and pressings, and storing them in different locations. I'm not even joking.If you really want the right version and track listing, you'll have to preserve it yourself.
Anyhoo, the official site had all the songs (5) for download so I forged my age to access the mp3s and I still listen to them, to this day.
I just wish I knew more about music formats back then, I have no idea if there were FLAC versions available. My annual pleas to the Tanqueray and Music Beast Twitter accounts remain ignored.
Alas.
In each instance I've come across this problem (mainly various Bob Dylan albums so far), it turned out that on the actual CDs, the affected portion had been mastered as part of the pre-gap (i.e. the bit where a hardware CD player will show a negative timecode counting back up to 00:00), so as silly as it sounds, it almost feels like they themselves simply ripped the CDs in order to upload them for digital downloads and streaming, and somehow managed to mess up the process and discard the pre-gap bits.
Though unlike your example at least in that case the CDs are still in print, so for now you're not completely stuck there…
why do you feel that sounds silly, as it sounds exactly like what I would expect them to have done. it was probably some intern tasked with the job as well. it's not like they were going to go back to some mastering format to have the streaming files created for the entirety of the back catalog.
I bought this album on Google Music, which got shut down and I didn't copy it off in time. Now I can't find it online.
https://www.hackster.io/mark-hank/sonos-spotify-vinyl-emulat...
and I think I found that via:
https://www.raspberrypi.com/news/who-needs-vinyl-records-whe...
FLAC is 100% identical to the original audio CD sonically, it's a lossless format. See it as a ZIP file for WAV audio files.
I've recently come across that problem at least twice – I learned about a potentially interesting album done by somebody, but it had already disappeared from Bandcamp without so much of trace. In one case it turned out I had missed the album by only two months, because the album page still appeared in Google's search index cache with a corresponding last indexed-date, but of course that's not much use as far as getting the actual music is concerned.
Ultimately, unlike the OP I luckily found at least the one album I was most interested in "elsewhere", though it's still a bit of a shame that it's no longer regularly available and it'd absolutely would have merited paying for it.
https://web.archive.org/web/20130825015207/http://dombrovski...
As long as someone is not “distributing” (AKA uploading to any user. AKA torrents are off the table.), then an argument can be made that the user is format shifting using the internet as the format-shifting tool.
Even though the US (and now other countries) are a bit of a nightmare for fans of copyright works, I always thought we’d end up with more tools that “read” a physical item, then download the bit-perfect copy.
In general, error correction adds a lot of time. My sense is it will re-read the damaged sectors several times and take the "most common" reading. Then, it compares that with the external full-track hashes to establish correctness. So, it has to do a lot of seeking to read and re-read a damaged portion, which can take minutes for "secure" rips in EAC.
Except when intentional errors are used for copy protection.
The gist is that you need to clean your CD's meticulously before ripping them!! Also, clean the CD-ROM player's lens using a cotton swap and lens fluid if you can (wrap the cotton swap with non-abrasive lens cleaning tissue).
If you do that, there will be no audible difference between the ripped version and the data on the original CD (providing you don't have some setting turned on in your ripping software that does some kind of processing of the data). You don't need special software to get a clean rip. I've been doing it this way for decades.
The signal processing IC's do set flags on their pins to indicate whether interpolation was done, but this isn't handed to device driver AFAIK and certainly not passed on to the client software reading the data.
https://en.wikipedia.org/wiki/5.1_Music_Disc
I have a 300 disc Sony changer that I am filling up with DTS discs.
The normal interpolation algorithm in CD players is not going to work right for a DTA stream so the quality of the stream is excellent.
This guy dumped the digital output of two identical CDs into a digital recorder and found the output was bit-by-bit identical
https://www.youtube.com/watch?v=f-QxLAxwxkM
For the life of me I can’t figure out why Techmoan hasn’t demoed a DTS cd on his show.
For data discs (or more accurately, the data portion of a disc, because some discs have other types of content on them too), the drive essentially functions as a standard random-access block device, just like a hard drive or floppy disk. The error correction etc is still hidden from you, but as long as your entire disc is just data, there's an easy, obvious thing to dump: the bytes of that block device.
On top of that, the audio extraction mode (in the drive) is usually optimized for constant-time delivery rather than accuracy; if an audio sector is damaged, most players will opt to deliver inaccurate data rather than stop the audio stream altogether.
On a data CD, those 2352 bytes are split in 2048 data bytes, plus an additional 4 error detection, 276 error correction, plus some other bytes including an address. So there is an extra layer of error correction.
Additionally, there is no way to know that you've passed that threshold -- the drive will just return bits, with no way to guarantee that they're correct.
A better description for this is "passive" error-correction. To try to get a 100% bit-perfect rip, it's preferable to start off with a drive that you know minimizes errors in the initial read process.
That just goes to show the lack of design foresight. They were still thinking in terms of grooves carved in plastic or a waveform on tape.
My guess is that you couldn't fit a lot of information on it, what with the redundancy to make it perfect.
Real media like CDs usually have 'burst errors' like scratches that give a whole bunch of successive bit errors; however, there's an interleaving process to "spread" the errors across blocks, i.e. spread the information so it will resist a scratch. In an absolute way, indeed it's impossible to guarantee almost no errors (although you can guarantee almost no undetected errors) -- simply because your CD might be ruined in a way (which I guess isn't all too unlikely). By setting your rate R to a reasonable level above your drive error rates, you can get almost no errors (as few as you like) within those bounds (and fail above).
Here's a typical error rate curve for a moderately large (255b) code: https://en.wikipedia.org/wiki/Reed%E2%80%93Solomon_error_cor...
That is, however, error detection and not error correction.
Once an error is detected you can apply 'correction' to the nearest probably correct value .. which only 'works' for errors under a threshold.
Hence the assertion in comment above yours.
A 32 bytes hash for every 32 bytes? Then you double your storage requirements.
There are actually WAY better algorithms for this than a hash (a hash is meant to handle secrets and cryptography - it's the wrong thing for this purpose).
ECC codes can tell you if the data is bad - and you can choose how many bytes this covers, and that can also fix bad data, again, you can choose how many errors per byte you want.
If you are curious: https://en.wikipedia.org/wiki/Error_correction_code
The rest of your comment is good, but I wanted to quibble with this: hashing is useful for a very wide range of things, and are one of the main building blocks of modern algorithms (hash tables being the best know).
Interesting ! At what point with today's high-capacity-disks (DVD etc), can we just say store the data x3 and take the value of the closest two ?
If the only thing we really want is "high-audio-accuracy" ?
A N bit checksum can at most distinguish 2^n different bitstreams as "100% accurate", and that is assuming there are no transmission errors in transmitting the checksum itself. This is proven trivially by the pigenhole priinciple.
And for most media, the checksum is itself transmitted over a noisy channel like the rest of the data.
Read your link carefully.
Regular error correction + checksum is a pretty solid implementation of that.
Another way to see it - care to list any error code used anywhere with zero error rate?
Read the paper. Shannon write about this in his 1956 paper and subsequent work analyses further, but it's theoretical, and probably not practical for any real world error codes.
So if noise in the channel could turn a 0 to a 1, or a 1 to a 0, then there is never a positive zero-error rate.
This is explained in the paper, page 9, second column.
So this is an interesting math question for channels that don't occur in practice, and has resulted (as far as I can tell) in no working codes usable even in labs. It certainly does not apply to transmitting binary data over any noisy channels or media, which is where most if not all error correction codes are actually used.
You have a skewed perspective of growing up with ubiquitous computing and cheap ICs. This was very much not the case in the early 70s when CDs were designed. Trying to simultaneously design for high fidelity audio and a then completely theoretical demand for data storage would have made no sense and would be unjustifiable scope creep. The spiral groove of a CD is objectively better for a low cost linear playback device - even though it is obviously worse for random data access (though modern DSP mitigated this 15 years later). Engineering is about trade offs.
There’s no binary best effort vs dead sure - it’s a matter of degree, unlike the mistaken GP post, CDDA and CDROM both use RS error detection and correction, just the former has less of it - a CDDA (Redbook) can store about 15% more audio than would be possible with the equivalent WAVE files on a CDROM - at the expense of some redundancy, but it still has very robust error correction - otherwise every little tiny piece of dust would cause a skip. And for audio playback - trying to fill in a best guess is the right thing to do - rather than just making the thing spit the disk out with a failure.
This is not quite right. CDs use a forward error correcting code, and of course that means you can detect errors (how can one correct errors if they can’t detect them?). The CD drive most definitely can distinguish between an error free signal and a marginal one - both at the analog level and the digital level.
Most CD drives firmware are a bit inconsistent in their ability to correctly provide error correction information but this is not a fundamental issue with the technology and far from “no way to know” - relying on the drives error detection facilities is what the C2 error detection option in EAC does for instance. There are some drives that do well with this.
https://wiki.hydrogenaud.io/index.php?title=EAC_Drive_Option...
Ripping software like EAC or whipper do verify, using an online DB of rips, that your rip is 100% bitperfect. So ripping "as accurately as possible" ain't a concern: your rip is either 100% correct, bit for bit, or it's not and you should start over.
If you're on Windows use EAC while if you're on Linux, use whipper. I use whipper and it's nowadays stock on many Linux distros (Fedora, Debian, etc.).
> but then I'll finally be able to throw out all these CDs.
Well I think that legally you need to keep the CDs if you want to keep your rips (it may depend on the country but in Europe you're allowed to rip CDs you physically own).
Just clean your CD's carefully and meticulously before ripping. If possible also clean the CD-ROM player lens using a cotton swab wrapped in non-abrasive lens tissue (available at photographic stores) and dipped in lens cleaning fluid.
Check that your ripping software doesn't do any other processing. For example the 'freac' ripping software has an "Enable Signal Processing" menu item which should be turned off.
Anyone has some idea what is this measuring?
Bit accuracy (way too low)? A full disk read 100% error free? The actual accuracy really depends on the quality of the disc, is this measured with a reference disc?
Wait: I'm pretty sure that's not how it works at all. No matter the drive, there's only one correct rip. And two people won't, by coincidence, happen to read the same disc with the exact same error(s).
So when you rip and compare to the AccurateRip DB, you know you have the correct rip if it matches the one already in the DB.
I don't think AccurateRip "considers" rips correct not: to me it knows with 100% guarantee. Which is the entire selling point of the AccurateRip DB. And which is why it's called that way.
Depending upon the quality of the drive mechanism itself, and how it reports errors to the ripping software, you can end up with incorrect data more or less often.
There is "analog data that is within a threshold of the bit that we decided".
Hard drive or CD doesn't read "1" or "0", it reads say (after normalization) 0.823V. Or 0.654V. Or 0.983V. The way circuit is designed decides that above some value its "1", below that is "0", but physical deterioration and material variance will change that value and it will differ bit by bit too.
It's all good when signal is strong and say all the 1's are > 0.8, all the 0's are < 0.2 but when it starts to blur (again, bad media, degradation etc.) you get bits that can be 1 one read, 0 another. And having "tri state" input that can say "ok, it's between 0.4-0.6, we can't tell whetehr that's 0 or 1" is generally rare.
Does EAC today offer any advantage over cdparanoia?
It's a pretty easy workflow. EAC to rip. FlacSquisher to convert to MP3 and OPUS. Then copy the files where they need to go and delete them from my laptop. CD then goes into a bin in the garage until I need it again or want to admire the artwork.
- 192 kbps CBR mp3 - this was supposed to be good enough (it is) and I had an iPod with only 20 GB so I went with it
- V0 VBR mp3 - I told myself I could hear the difference on my now slightly fancier headphones, and this would not kill my OiNK and later What.cd ratio as fast as lossless
- FLAC, if possible 192 KHz and 24 bit - now I got into really expensive headphones and I only downloaded the most verified of torrents (if possible vinyl rips) or I ripped using all the EAC/Paranoia guides out there. Played music via Foobar2000 with some plugins to minimize “interference” in the audio path. Had a full DAC + Amp setup
- Got a job and kids
- 256 kbps VBR AAC, via Bluetooth to my AirPods Pro - I don’t have time anymore to fret about audio quality, streaming is so super convenient and let’s be honest, above a certain baseline no one hears the difference anyway…
Ages ago I made custom cdrom drivers that would purposefully randomize data returned from reads (cd copy protection). EAC would rip right through it without skipping a beat and reconstruct the correct audio. The only way to defeat it was to never return all of data, randomized or not.
Would like to know how it's seen in the ripping community https://wiki.ubuntuusers.de/abcde/
"There is no "digital data" in physical reality.
There is "analog data that is within a threshold of the bit that we decided".
Hard drive or CD doesn't read "1" or "0", it reads say (after normalization) 0.823V. Or 0.654V. Or 0.983V. The way circuit is designed decides that above some value its "1", below that is "0", but physical deterioration and material variance will change that value and it will differ bit by bit too.
It's all good when signal is strong and say all the 1's are > 0.8, all the 0's are < 0.2 but when it starts to blur (again, bad media, degradation etc.) you get bits that can be 1 one read, 0 another. And having "tri state" input that can say "ok, it's between 0.4-0.6, we can't tell whetehr that's 0 or 1" is generally rare."
And it’s not like audio is actually high fidelity. Any DAC will introduce a lot of errors. And the original recordings are not perfect either. Microphones, loud speakers, amplifiers, that’s a lot more errors than one wrong bit too.
https://audiophilestyle.com/forums/topic/61633-intel-cpus-do...
The data is digitally encoded 16-bit PCM, and already has builtin CIRC error correction, there is no "quality" or "accuracy". Your software/hardware either correctly ripped the data or it didn't and this can be trivially verified with crc32 against CUETools or AccurateRip or whatever.
The data seems to be based on telemetry collected for a CD ripping software, which I'm slightly more inclined to believe is real than some mumbo jumbo about vinyl warmth or whatever.
> Drive: HL-DT-ST - BD-RE BH14NS48 (116 users): Submissions: 10260 accurate, 23 inaccurate, 99.7763 % accuracy
23 of these submissions come from a broken drive, or the disc is damaged to the point a correct read cannot be made. It's not that the drive is producing 99.7% accurate data, there are 10260 100% accurate submissions and there are 23 wrong submissions.
It's a binary state, either the drive is functional and the CD in decent condition, or the drive is broken/CD is damaged. Some drives may have a higher failure rate than other drives, but every drive is either 100% accurate or busted, there is no 95% accurate drive. That's just a broken drive.
Is an entire track inaccurate if a single sample is inaccurate? In that case then yeah, rather than measuring some level of audio purity, this is measuring your likelihood of getting a perfect rip before taking CD condition into account.
No. How good the drive is at reading marginal discs varies a lot between drives.
The problem is, these populations are not identical disks, and the error rates are low enough that we're not really sure we're seeing a good metric of this.
CD: The Weeknd - The Highlights
Drive: Pioneer BDR-XD05TB (not listed in my article)
Accurate Rip reported errors on 3 Tracks.
Ripping with hp GUE1N
flawlessly worked on the first try. Although I also doubt that the error would be audible, but I have a better feeling if the hashes are correct.If your pressing doesn't have a good sample size of hashes to compare to then you use a more secure ripping method that generally makes use of the accurate stream feature (most drives manufactured in the last decade or so have it) it's a slower but more repeatable process of reading a red book CD intended to suppress jitter and alignment errors. If you get through an accurate stream rip and you encounter no read/C2 errors you're probably ok, but if you really have no/few samples to compare against you may be paranoid and rerip the disk a few times to make sure you get repeatable results or even rip the disk with another drive all together.
I remain suspicious of the cargo cult of EAC as doing something substantially different than cdrtools or anything else that knows the ins-and-outs of the redbook format
That would be useful.
If you have a damaged disc, a drive can be better or worse at resolving "marginal" parts of the disc, which will change the data. Whether or not this makes an audible difference is an exercise for the reader, of course.
CD-ROM doesn't have this problem because it introduced a different error correction mechanism with more overhead. So in that case it is "either it ripped correctly or it didn't"... at least until you start getting into mixed-mode discs, or games that stored their FMVs with the weaker error correction mode to save space, or literally anything with copy protection.
It's been a while since I've looked at the standards in detail but I believe that involves another layer on top of the layer of error correction that CD-Audio has.
Why not? Shouldn't it be non-destructive?
So the multiplier is 65535/2000, and the intermediate result is [0, 32.7675, 98.3025, 65535]. But because the output must be integers, so decimals will get rounded off.
It is possible to restore the original numbers exactly if you know the scaling factor, but it's an extra complication that is not worth the hassle. It's better not to burn the normalization into the CD rip; it's better to do normalization on the fly at playback time.