The Internet Archive and Google Books have a significant head start over the largest library in the world
The Internet Archive and Google Books have a significant head start over the largest library in the world
> The Library's content, programs, and expertise are national treasures – we are dedicated to sharing them as broadly as possible. The growth of the Library's digital content, which includes our collections, has increased exponentially every year. We will make that content available and accessible to more people, work carefully to respect the expectations of the Congress and the rights of creators, and support the use of our content in software-enabled research, art, exploration, and learning.
> Exponentially grow our collections
> The Library will continue to build a universal and enduring source of knowledge and creativity. We will expand our digital acquisitions program, as outlined in Collecting Digital Content at the Library of Congress; continue our aggressive digitization program, which prioritizes the Library's unique treasures; and improve search and access services that facilitate discovery of materials in both physical and digital formats.
> We will expedite the availability of newly acquired or created content to the web and on-site access systems. This will mean making improvements to the procedures and tools we use to move content from acquisition or creation to access, which will be critical as we continue to experience exponential growth in the size of our digital collection. We will also improve tools to make it easier for Library staff to enhance content after publication, such as adding additional description or information about the conservation of objects.
To answer these “who am I” questions posed in the article, of unidentified historical photos, we need institutions like the Library of Congress digitizing their records, uniting them with other institutional datasets, and organizing them in a way that we can run facial recognition algorithms on them. The answer to “who are these musicians circa 1930” probably comes from a local newspaper that ended publication 50 years ago.
In case it isn’t abundantly clear, this is going to take decades to accomplish. These institutions operate on timelines measured in centuries, not weeks. They care deeply about problems like bit rot when replacing physical archives. If it is lost to the world in mere decades, it is not useful to them.
The Internet Archive is of little help here, beyond sharing technical information for efficient archival and retrieval of truly massive public datasets.
This is such an important point. Governments have longevity and universal service requirements that others just don't. I am personally a charge-ahead-and-try-shit kind of person. But I recognize that works for me because I can easily say "fuck it" and move on from my experiments. Anything the LOC does they're stuck with, possibly for centuries.
So I'm glad that they proceed at a pace that seems positively glacial to me. That's what success on those timescales often requires.
Just because they administer copyright doesn't mean they simply get to ignore it themselves.