I need to copy 2000+ DVDs in 3 days. What are my options?
reddit.com
reddit.com
> Maybe you should disclose your city or at least state/general location
Washington DC, specifically College Park, even more specifically The National Archives at College Park.
> What is the content of the discs?
Years ago in partnership with the National Archives Amazon.com digitized a tremendous amount of film for the National Archives, with the catch that the material cannot be freely disseminated online by the National Archives until Amazon breaks even on their digitization investment. It’s been years now and for a variety of reasons - many of which are Amazon’s fault - there still exist a solid number of discs that Amazon hasn’t sold even one of. (In part because amazon hasn’t even had them all available to sell. Pretty ridiculous).
The flipside of the catch is that the DVDs can be viewed and even copied for $0 on site by researchers. I am doing research there and I have have a research pass. I was looking at many of these titles yesterday. It’s time to set these national treasures free.
https://www.archives.gov/press/press-releases/2007/nr07-122....
For all the helpful advice offered both here and on Reddit about how to do this, I wish there had been more time asking if this was something the OP should be doing in the first place.
I completely understand the frustration if this is indeed the above set of films and it's been over eleven years since the digitization agreement and files have not yet been made freely available to the public.
That said, planning to use a researcher's pass to "set these treasures free" and blaming Amazon for not having generated more revenue from these films makes it sound like the OP thinks he knows better than the staff of the National Archives how to best care for these assets and that he knows better than the folks at Amazon how to turn a profit. To me, it sounds more than a little arrogant.
A big reason the National Archives enters into agreements like this with companies is that digitization, especially on their scale, is expensive. If it weren't for agreements like this, Amazon would only want to digitize the films for which they knew they could turn a profit and the vast majority of the collection would sit un-digitized and be at risk of loss. The tradeoff is between the immediacy of access vs the number of assets digitized and the team at NARA made the decision that it was better for the American people to have more assets digitized.
I'm worried that the next time NARA is in talks with someone about a digitization agreement (for example, if there's a large number of early jazz audio recordings on 1/4" reels and Spotify is interesting in paying the cost of digitization in exchange for 2-3 years of exclusivity) that the company will point to this example and say "didn't you just let a researcher publish the entire collection Amazon digitized? How can you assure us the same thing won't happen with these recordings?" The result will be the National Archives clamping down on researcher access. I think that would be a net loss for everyone.
Since amazon was only looking to get its investment back, but not a profit, then clearly the intent of both parties was to eventually have it release to the public. Whoever on amazon’s side pushed the deal presumably thought they would be sold, but it clearly didn’t happen. It’s much more likely this has been sitting on a backburner somewhere, left as some forgotten plan, than some kind of long-term strategy to induce sales.
Thus, by virtue of an ill-written contract, we’ve entered a situation that no one wants. Amaxon no longer cares about it, the museum presumably prefers releasing it, and the public can only benefit from it.
By virtue of that same contract, there’s an escape hatch that might bring us back to a state where everyone is content.
It would make sense to exploit it.
Ofc, this is assuming Amazon doesn’t care. But amazon is a company, and the larger a company is, the less distinguishable it is from a government. And governments certainly have control over things it has collectively forgotten about, as it only rarely operates a single, like-minded, cooperative organism.
And amazon is indeed a very large company. It would hardly be unsurprising for Amazon to not even be aware this contract still exists.
Or maybe the National Archives will stop making deals where free distribution of the material is unlikely to happen for decades and will instead open up greater access to researchers.
Or they could make the same deals but put a deadline on them to prevent the exclusivity from sitting in limbo forever. That would at least give Amazon an incentive to distribute the material rather than just squatting on its monopoly.
And yet there are researchers willing to do this quickly and for free, so why do we even need Amazon in the loop?
> The flipside of the catch is that the DVDs can be viewed and even copied for $0 on site by researchers. I am doing research there and I have have a research pass. I was looking at many of these titles yesterday. It’s time to set these national treasures free.
-- https://old.reddit.com/r/DataHoarder/comments/a6fkpm/i_have_...
GP was just wondering out loud about the content itself, not the format of said content.
>AWS evaluates applications to the AWS Public Dataset Program every three months.
>If we bring your dataset into the AWS Public Dataset Program, we will cover the costs of storage and data transfer for a period of two years
yeah i don't think that's public,more like a rental service. i wouldn't bother working with that unless i had free AWS access through corporate
Can you detail how one easily gets a 1 gigabit uplink installed temporarily in the National Archives?
I've had telcos drop fiber into the middle of fields in the middle of nowhere, so it's probably possible. But it feels like there's a lot of complicated bureaucracy to navigate, and the researcher OP is time-restricted.
No one got the joke. I’m suggesting that the DVD’s should be submitted as an AWS public dataset, because Amazon is the company that digitized them in the first place, and the content is apparently public domain.
If anything is out of your few commands, just save work elsewhere, delete the project, and download a fresh copy.
not ridiculous, good lawyer preparing the deal
Honestly it's always been pretty easy to ignore it since it's inception unless you are running a counterfeiting operation.
I was looking into this sort of thing recently for a project. Very interesting technology space with some genuine direct benefits coming out of the application of machine learning.
Someone's got to build that metadata. Especially back in 2009, you couldn't just download it from some public database. Databases have to be built.
Was the postal stream the limitation on throughput?
Aside: Re-encoding to a tighter format at that time wasn't too bad, but still relatively slow IIRC.
Organizing the DVDs by length would allow optimizing the loading/ripping process to assure minimum time lost waiting for the operator's hands to be free. This kind of planning makes for an interesting project.
It might make an interesting crowd-funded project, if it's reasonably easy to get a research permit. Plan it out, go in with the hardware, come out with the images. Use all the error correction opportunities you've got available.
Do a web search for "bulk dvd ripping" (without quotes) and you'll find lots and lots of discussion and advice, including some about building a dedicated DVD ripping rig. MakeMKV gets good press, in my very quick read of a few posts.
And there's always the option of crowd-funding to raise the exact amount needed to pay off the break-even for Amazon's investment. I can't imagine they'd fight back too hard when looking at a large check vs. a non-performing asset, unless Bezos personally never intended to let the footage go free.
The data rate falls to the floor as soon as the access pattern isn't sequential though, which if you are using multiple readers it won't be. While an OS might be bright enough to organise data flowing out of write buffers so it isn't as random as it could be there is a limit to how far they will go with this because they are general purpose OSs and optimising for multiple bulk streams will punish more interactive activity. If you have a tool they bypasses the OS cache and works in large enough blocks you might see better results except if the write activity from each lines up at which point this will make things worse.
Pulling the data off multiple DVD drives onto an SSD, swapping to output to another once near full to continue while its contents are dumped sequentially to cheaper-per-Gb traditional drives, would probably be the way I'd suggest.
In fact, you would get away without swapping between two SSDs: the read activity pulling data from the SSD to a traditional drive is unlikely to have much effect on the write performance for the data coming off the DVDs unless you have a great many readers in one machine. If doing this all relatively manually, to reduce manual steps once a DVD copy is complete add it to a queue to be moved using something like https://en.wikipedia.org/wiki/TeraCopy so you don't have to worry about manually coordinating the SSD-to-cheaper copy operation to keep it sequential.
Assuming 15 minutes to read each disk (it is a long time since I pulled data off a DVD in bulk so this is guess work based on old memory of it taking a little more than 10 minutes to read a full DVD9 disk, and rounding up to 15 to allow for manual process inefficiencies and some disks being slower to extract due to condition causing rereads, etc) you are looking at wanting 21 or more drives constantly on the go to get the job done in 3 solid 8-hour days (2,000x15/3/8/60 = 20.8). Five laptops each with an internal SSD (128G+) to extract to, five DVD readers on USB3 to extract from, and a 4+Tb spinning disk (also external) to finally write to, might do the job and have the space (2,000x8.5/5=~3.5Tb output per laptop). You'll need a powered USB dock/ for each laptop instead of a passive hub, and you are going to want to add more of everything to allow for the possibility of device failures.
Of course significantly less resource is needed (or you get more contingency time (and/or spare kit to deal with failures) from the same resource) if most of the media is DVD5 and/or not full disks. I've assumed the initial three days is just for obtaining the content - I've not accounted for any other processing (such as indexing and transcoding) or further distribution.
More interesting is that even "sequential" read/writes already have seek times built in because HD's aren't spiral track, so head switching, and track to track seek (and the associated rotational/finding the servo track) are inherent in sequential IO perf. So most filessytem placement/schedulers aren't going to place 3 files being written at the same time on opposite sides of a disk, so those head switch times and track times have nearly immeasurable increases because the drive itself is also storing a large part of a track write and moving 3 tracks and a head, is basically the same as just moving a head.
just have your script queue up a few in /tmp before moving to the mounted storage.
pretty easy & now you have multi level buffering caches that Linux knows how to work with efficiently & that can nearly guarantee sequential writes
The government has to explicitly permit you to do research?
What the fuck is wrong with this world?!
My question is not necessarily related to OP. But can anyone tell me why my ripped DVDs could not play in DVD software properly when there was jitter or corrupted data, but it could play properly when in the DVD drive? When I got lucky, the ripped DVDs would play fine. Otherwise, it was play fine for a few minutes until it ran into the first bit of corruption. Confused the heck out of me.
This used to happen to me a few years ago, and it was overwhelmingly just a sign of a marginal drive. For about 10 years, I was replacing drives about once a year. It doesn't happen as frequently since I started using bluray drives (usually LG) to rip DVDs and a piece of software designed for "piracy".
Anyway, two software things to try, are lower the ripping speed to 1x-2x, and make sure to leave the drive retry on bad sectors enabled. Most decent ripping software will have options to control the error correction/retry logic. The problem is some copy protection schemes leave bad sectors on the disks, and this will hang a lot of the "honest" ripping software that doesn't know how to deal with it.
They may not have all the latest whiz-bang features, but they are rock solid, and are USB powered. I have two at home, and use them to rip DVDs, and have never had a failure in I don't know how many years.
New they run $79 from Apple. Or you can get a used one off shopgoodwill.com usually for around $15.
The solution: a wet polisher for DVDs. Luckily, my local used-game shop polishes disks for $3 each. They use an [Elm Eco AutoSmart][1] and it makes the discs like-new. I have never seen a disc it could not fix. You'd have to ask/search around locally to find one near you... I think some libraries have these machines, too, to maintain their DVD collections.
Five year embargo on the National Archive releasing whatever Ancestry chooses to digitise and host on their proprietary database. This model of digitzation appears to be the principle way that the records are digitized now: https://www.archives.gov/digitization/principles.html
I assume the film is still around in boxes, so event he DVD window drops someone is still free to go digitize them again, but given how fragile film is that's touch and go.
It's no different than if you went to the National Archives and scanned something; you would not be required to distribute the scanned version.
However, apparently part of the deal was that the National Archives would get a digital copy that they could lend to researchers on-site. So this individual wants to copy those digital versions, piggybacking off Amazon's efforts. He/She's probably allowed to since they are public domain.
Archives preserve and disseminate. Budgets usually make this a challenge. They probably (rightly) view the preservation of content on decaying media as a big win.
Thinking 15 minutes per you have 30k minutes. 1 person can get 4 done per hour - let’s think 2 12 hour shift you just need about 25 people + their laptops.
If you need the equipment it could be rented by someone like this that does it for trade shows https://meetingtomorrow.com/austin-computer-rentals
So, I slightly modified Knoppix and started a netboot server, and made a dvd ripping cluster.
I didn't end up doing exactly this, but it's worth trying DVD::Rip. It specifically has support for ripping clusters[1].
But it doesn't even have to be something as complicated as a cluster, just as long as you parallelize things. How many laptops are you allowed to bring with you? Could you bring five laptops and 10 dvd drives?
Two HP towers, each with two optical drives, I spent nights and weekends ripping her CD collection. I just barely got it all done before Christmas. She still has no idea how much work that was.
Maybe put out a call to Jason Scott/ArchiveTeam/ArchiveCorps to saddle up. ArchiveCorps (run by Jason) has a Slack team for such Call To Arms.
OP: Have you attempted to renew your research card with the National Archive to get more time?
Better to just prioritize some sample of discs and do the best you can-- consider it a pilot run. If you monitor everything you'll be able to use the pilot run to estimate and propose what it would take in terms of equipment, time and process to do the whole thing in a reasonable time.
That PC sits in a basement now, I think. If anyone lives in Milwaukee and wants to rip some DVDs over the holidays, get in touch, more than happy to loan the machine out.
It's pretty much the gold standard for bitperfect CD rips amongst the torrent community
I thought CDs are digital so whatever is received is exact by definition?
DVDs do not have that characteristic. While the streams will be recoverable, they will also be very visibly broken if the visible data gets corrupted, and that's much harder to repair. We are much more sensitive to visual corruption in general, and the way the compression works can make even small errors loom large on the screen. (As the compression gets more sophisticated, this gets more true, which is why DVDs generally just have fairly confined block errors, but modern codec errors can cause significant corruption on the screen, followed by the corruption on the screen "moving" like the video was supposed to.)
Of course it's not 'bitperfect' at that point but then thats why EAC does so many retries.
The only way to get an accurate rip of audio data is to compare it to a known good source. This is why Accurate Rip exists...
Source: I've ripped 4000cds :)
EDIT: This applies only to audio tracks. Data tracks don't suffer from this issue.
That's the only way to confirm you got an accurate rip. It says nothing about how to get the rip in the first place. In my experience, "full of errors" sounds exaggerated. Most of CDs can be easily ripped in 'burst mode' (in EAC) or equivalent in other software and you still get same (exact CRC32) result from AccurateRip DB with no C2 involved (I disabled it), as soon as you set the offset correctly.
Disclaimer: I didn't rip 4000 CDs, only hundreds.
'Back in my day' I would chain FireWire drives together (better than USB/Hubs at the time)
So, at minimum you need one machine more than the number of minutes it takes each machine to rip the DVD. If you also have to burn a second set of DVDs, you will need at least one machine more than the number of minutes it takes to burn a DVD plus the aforementioned ripping machines.
If you're just storing the DVD images after ripping, that part can be at least partly automated, set up a machine as a central server (about 20 terabytes should do it if nothing goes wrong) and have each 'ripper' run a script to copy the image up after the rip is done (which hopefully it can do while also starting the next rip, if not you'll even need more machines).
This is the reason why I never bought a single player or disk, Blu-ray and HDMI/HDCP copy protection went way to far (especially with the chain of custody nonsense). At the end, if the industry wants to fuck legitimate customers like that, fine, I just torrent everything, this way only a single person/group has to figure out how to rip the disk or broadcast, just once. There is no reason what so ever to feel bad about it, they did it to themselves, no informed customer should ever buy into that shit. Its self defense, I'm not buying a TV for thousands for the copy protection to brick it eventually.
The only problem I’ve had is when they try to make it hard by putting a gazillion titles on the disk to make it hard for you to figure out which is the right one. That sucks, but is not insurmountable.
> In the event that a key is compromised and published online, or used widely in any way, that key is depreciated and all Blu-ray’s from that point onward contain keys that cannot be decrypted with the compromised hardware key.
If someone merely provides a title key decryption API, is there any way to figure out which device key they're using?
Getting keys from your hardware is a hassle and I didn’t want to wait for a decryption solution later.
Newer codecs help with that, yes. However, streaming will not be as high quality as a "station wagon full of tapes" i.e. local, in the near future.
My top speed is 8 Mb/s. This is the most prominent provider in my country. My plan is 20 Mb/s, but that has literally never happened. A lot of people in my country use wifi providers with outdated equipment, giving them around 10 Mb/s as well. And lastly, check the page on average speeds of the Internet on Wikipedia.
BTW the absolute majority of customers are sharing their bandwidth and the infrastructure is built with that in mind. If everyone streamed 4K videos, the actual speed would go down massively - of course since the only solution would be to rebuild the infrastructure, that is not going to happen, instead they will lower the speed or introduce FUP.
You can see Czech Republic pretty high in that list - that's because it's average speed, not median. I have a friend that has a 600 Mb/s plan that's 3 times cheaper than mine, but that's not the norm - and I guess it's similar all around the world, either you're lucky... or not.
Exactly.
Maybe you could reach out and borrow one?
From a quick google, it's the VGP-XL1B/VGP-XL1B2/VGP-XL1B3
They are still around but have since switched to a streaming only model. You could reach out and see if they still have their DVD ripping infrastructure, I'm sure they would love to help.
I use to own one in the 90s. I would do work for churches and conferences. I could get a thousand done in a few days. They use to have automatic feeder drives for producing more. No one sells them anymore it appears.
If he has network access, then it might be faster to just upload the entire set directly to another drive over the internet, since most universities have fairly good connections, that shouldn't be any slower then trying to write them all to DVDs.
It's definitely possible with phonograph records.
(Yes, I know.)
It's about 4 GB, which means that if you got a perfect image of just the data of the DVD (one pixel imaging each bit) the size of the picture would also be 4 GB. That means, as far as I understand, a 4 gigapixel image AT BEST, assuming you were able to get every single pixel to capture one bit of data on the DVD. In reality, since DVDs are circular and images are square you're looking at images that would be more like 6 gigapixel if my math is correct.
Today, available cameras are at best 50-megapixels, which means you need at the very minimum about 100 photos of each DVD, but in reality, there is no way you can ever get one bit per pixel, because to begin with the bits are not aligned in rows unlike the camera's pixels, so you're most probably looking at ten times that.
It's an interesting idea but I'm pretty sure it's completely unpractical.
Edit and I realised I counted 4 GB, but I should actually have used 32 Gb as a starting number, as what matters here are the bits, not the bytes, so add another 8x to the amount of photos you need.
Pied Piper could do it. Would use their “middle out” compression algorithm - anyone with a 5.2 Weissman score could do this pretty easily.
Each pit on the DVD represents 1 bit, but need at least 1 pixel to stored. You need 8 bit to store a pixel. So you have write amplification of at least 8 when taking picture of DVD.
It's not even remotely practical but I thought it was an interesting idea. Judging by the way the votes are yoyo'ing up and down, about half of HN readers agree. :)
Then software would have to sort out the mess.
Not to mention how albums would often be spread over two records where as the same album could fit on a single audio CD (let alone DVD).
I'm guessing the GP has never owned any vinyl. I could forgive someone questioning the density of records if he's not familiar with the tech.
I'm old enough to know about vinyl. That's not really the point though - the principle is the same, just with much higher information density. With a suitably high resolution camera, or set of cameras, or video, and some serious magnification it should be possible to photograph a DVD and play back the data from the image. I find that quite interesting, or at least entertaining, to think about. I guess the HN readers who're downvoting the post don't.
Something to consider - technically we can already do it - it's how DVD players work, albeit with a single coherent light source and only reading one bit at a time.
I was talking about user U2EF1 (the "citation needed" comment). I didn't have any particular issue with your comment aside the impracticality of it with current consumer hardware (which I'll address below). But as a concept it's an interesting point.
> the principle is the same, just with much higher information density.
The problem is in the detail. Even with gramophone records, you'd get a very low success rate with a consumer camera. Plus you'd need multiple macro pictures and a precise way to stich them together. At which point it would be quicker to play the record while recording it in Audacity (or similar) in real time. So you're talking a less than 1x record speed and lower success rates to boot; and that's just low density vinyl. When talking about DVDs you'd need to improve the operation by several orders of magnitude in terms of accuracy and resolution.
So I think your comment was interesting from a conceptual point but we're a long long long way off having that level of detail in consumer photographic devices.
> Something to consider - technically we can already do it - it's how DVD players work, albeit with a single coherent light source and only reading one bit at a time.
I think it's a little disingenuous having a laser reading binary reflections sequentially compared to a digital sensor detecting literally billions of analogue reflections (because a camera isn't just detecting the existence or absence of light) in parallel and then forming a precise sequence of digital bits from that. The technologies are completely different, the scales are completely different, the concept is completely different. They're just not equatable.
> I find that quite interesting, or at least entertaining, to think about. I guess the HN readers who're downvoting the post don't.
It wasn't me who downvoted you but if it's any consolation, I've gotten downvoted for factually accurate comments before (let alone impractical suggestions). You just have to remind yourself that it happens occasionally and usually the positive outweighs the negative. :)
Note:it's been a looking time since I have tried to pull DVD files, but if my (rusty) memory serves certain types of DVD readers will just let you copy the vob.
Using dd (Linux) or Disk Utility (Mac) makes this straightforward.
Create a directory and a file on the target machine first a name like /myfile/this file Then to create a bit by bit copy of that file use this command Something like this Dd if=/dev/cdrom bs=1 of=/dev/sdc/myfile/thisfile.iso conv=notrunc