HNHacker News
TopNewBestAskShowJobs

mekarpeles

1,082 karma · joined March 10, 2011

Hi, I'm Mek (https://mek.fyi). I run OpenLibrary.org over at the Internet Archive.

I'm an entrepreneur, an advocate of non-profits, 2022 & 2023 Berkman Klein Center affiliate, and a PhD dropout. I helped run tech for Hyperink (YC W11) and Ark (YC W11) and co-founded Hackerlist.net + Baybo.

I am passionate about non-profits, open source software, GNU/Linux, teaching and education, and spreading knowledge across the lands.

My mission is to enable others and to curate a living map of the world's knowledge.

submissionscomments
mekarpeles··on Keep Our Servers Running
I think archiving the world's books is as valuable as websites. Many books represent years of an author's life writing authoritatively on a subject. And there are a staggering number of books (especially before 1980) that are simply disappearing from the public record.

As book production goes digital and titles are increasingly sold as digital "leases" that customers and libraries don't own (i.e. kindle, libby, etc), it's becoming more and not less challenging (counter-intuitively) for libraries to archive and preserve books into the future.

And that's not a coincidence. And I wouldn't be surprised if we continue to see a steep uptick of similar challenges for archiving the web.

mekarpeles··on Keep Our Servers Running
COVID was a particularly treacherous period for a lot of humans. Replying as myself, not my employer. I’m not offering an opinion about the program, as I was not directly involved; I’m only noting circumstances that I think are often under-considered or misrepresented.

1. The state of access.

Tens of thousands of public schools and libraries in the United States and across the globe were suddenly mandated to shut their doors, leaving many students, parents, and teachers to fend for themselves, to learn from home, and unequipped with access to the physical resources many needed. Meanwhile, thousands of public libraries across the United States who had these resources were forced to closed their doors, meaning tens of millions of books the public had paid for became unavailable. Many of these public good organizations appealed for some way to solve the unique access challenges created by these unprecedented circumstances, which connects to point number 2.

2. Who was involved.

The details regarding how a decision is reached, and whether its made independently or unilaterally or with input and comment, are relevant when forming an opinion. In my experience, many people opine on this topic before researching this point. In 2020, more than 100 public institutions and libraries who were affected by COVID, joined in endorsing some temporary way to offer students, parents, and educators relief by connecting them with resources they had lost access to:

https://docs.google.com/document/d/1vkl3RX4CzpRTQsoG1tsdHC0f...

mekarpeles··on Keep Our Servers Running
Wonderful to see @raybb and @anandology in this thread.

Lots of operational challenges come up when running a service for 14M patrons.

And Open Library in particular has a handful of challenges. 1. It's database has grown significantly (800+ GB) and Anand is right that IO (even on SSDs) is a challenge. The `thing` (infobase/infogami) triple-store design is well thought out and gets us a lot, and any system has to be tuned as it scales to hundreds of millions of rows. One strategy here is being smarter about cache and also shifting some of the load from psql to solr. Rishabh and others volunteers have been amazing assets as we've moved in this direction. Jim Champ on staff has been helping me tune psql, pgbouncer, and some of our high IO crons to improve raw db performance. 2. Limited hardware resources. We're trying to move some of our services within the Internet Archive's kubernetes cluster and we've done a great job migrating towards a world where everything is dockerized. It used to be a very painful process for our team of 3 to handle server ops, upgrades, and networking for nearly 15 manually orchestrated servers. One of the bare-metal racks running much of Open Library is significanly oversubscribed on vCPUs and so moving services off to free space and eliminate steal is critical for us right now. Our main web server (ol-www0) suffers from up to 20% steal and we're seeing a lot of congestion before requests even get to our web nodes (app servers). We have a plan and it takes time. 3. Open Library is still dependent on Archive.org for many lookups -- like book availability (which Ben Deitch has been helping me and Drini move into solr). When there are network issues and a network requests takes 5+ seconds, every web.py worker on that thread grinds to a halt and Ray's work moving us to FastAPI has made a significant impact 4. Solr. Drini has been heroic at restructuring our setup to use replicated solr in a way that has increased performance and relieved some of the pressure on our main cluster. This was a huge bottleneck for us this time last year and we've taken a lot of steps to ameliorate our situation. See: https://blog.openlibrary.org/2025/09/12/open-library-search-... 5. Raw spikes in traffic. We are seeing massive amounts of traffic that slams our book pages, increasing the pain of all the above. It saturates our limited resources, puts more strain on our database, ties us web workers... It makes modsecurity even more expensive. Part of the solutions is being more clever about provisioning, part of the solution is using fail2ban to prevent bad traffic from subtracting from the experience of the patrons who depend on us. Part of the solution is caching and optimizing our database to scale with load.

There isn't just one solution and the same 3 engineers on staff (and the support of a completely stellar community of dedicated volunteers fellows and leads) are doing our best to balance ops improvements with the necessary "product" and design improvements necessary that ensure we're useful to people to begin with.

I hope this gives the world a bit more of a glimpse how we operate and what some of our challenges are. We're an open source project and our goal is to share as many learnings as we can and to build something useful, sustainable, and beneficial for the community at large.

Thank you Ray, Anand, Drini, Jim, Lokesh, Lisa, Charles, and so many dozens more for your tremendous work (present and past) and thank you for being in our corner.

mekarpeles··on The Internet Archive has lost its appeal in Hachette vs. Internet Archive
> it frankly disgusts me that the project's current management has for the last few years had its focus on fighting windmills in court instead of their core mission - preserving our digital history.

Hi, Mek here (speaking as myself). Disclosure that I run OpenLibrary.org at the Internet Archive. I'm sad to hear you're disappointed with how things are going. I share your frustration.

I wanted to join in and +1 one of your comments: the importance of preserving our digital history. Preservation is a core mission of the Internet Archive and central to the tagline, "Universal Access to All Knowledge".

At the end of the day, the reason to preserve cultural heritage is so that it can be made accessible: Eventually. In ways that serve people with special accessibility needs who are otherwise left behind. In formats and environments capable of playing back materials that no longer have available runtimes. With affordances that make these materials useful and relevant to modern audiences.

An important reflection is that a key role of archives and libraries is to preserve cultural heritage by building inclusive, diverse collections, which span topics and times. For decades, libraries pursued this goal by purchasing physical books and, over time, growing and preserving collections of materials that serve their patrons. Not just bestsellers. Weird, obscure, rare research materials about rollercoasters, genealogy, banned books, stories from lost voices, government records.

The shift of publishing to digital [especially how it's done] fundamentally affects how [of if] material may be archived or accessed. It's not enough to assert the importance of preserving culture. One must actively advocate for a future where media can be archived. As Danny suggests (https://news.ycombinator.com/item?id=41454990), this is something the Internet Archive has been acting on since its inception.

What we're seeing today is a shift to digital, designed and led by publishers who are engineering a landscape with new rules where libraries can't own digitally accessible books. Libraries are being offered no choice, no path forward, but to lease (over and over) prohibitively expensive, fixed pool of books, that disappear after the lease period is up. This means libraries have ostensibly lost their ability (first sale doctrine rights) to own, grow, and preserve a collection of books over time... A fundamental ecosystem change that threatens the very function of preservation that you and I so strongly value. Preservation necessitates the ability to preserve. Preservation is a fight for the future and I believe a preservable future where libraries are allowed to own digitally accessible collections of books is a future worth fighting for.

That doesn't mean we should only be looking into the future. Looking at today, the only permanent collections libraries do / can own and preserve are physical. So what other question is there besides: how can libraries make the materials they rightfully own, preserve, and are permitted to lend accessible to a digital society? How may libraries make the digital jump to help millions of physical books enter public discourse, which takes place ostensibly online?

In my opinion, this is the discussion we're having. The Internet Archive continues to preserve millions of documents of all sorts: websites, radio, tv, books, scholarly articles, microfilm, software, etc. A very small team of staff are doing the best job possible to make sure that, not only does our cultural heritage get archived, but that in the future, archives and libraries have the right to exist, be useful, and that there are materials archives are permitted to preserve; that important research resources are made accessible to the public -- especially those who have traditionally been left behind. Someone needs to fight for the future that lets us continue preserving the past.

I'm personally very open to your suggestions on how the Open Library can improve and appreciate you taking the time to share your thoughts.

mekarpeles··on Let Readers Read
Thank you for posting @mdp2021 --

Mek here, program lead for OpenLibrary.org at the Internet Archive.

Over the last several months, readers have felt the devastating impact of more than 500,000 books being removed from the Internet Archive's lending library, as a result of Hachette v. Internet Archive https://help.archive.org/help/why-are-so-many-books-listed-a...

In less than two weeks, on June 28th, the courts will hear the oral argument for the Internet Archive's appeal.

What's at stake is the fundamental ability for library patrons to continue borrowing and reading the books the Internet Archive owns, like any other library.

Please consider signing the Open Letter to urge publishers to restore access to the 500,000 books they’ve caused to be removed from the Internet Archive’s lending library and let readers read.

mekarpeles··on [dead]
Mek here, program lead for OpenLibrary.org at the Internet Archive with important updates and a way for library lovers to help protect an Internet that champions library values.

Over the last several months, readers have felt the devastating impact of more than 500,000 books being removed from the Internet Archive's lending library, as a result of Hachette v. Internet Archive https://help.archive.org/help/why-are-so-many-books-listed-a....

In less than two weeks, on June 28th, the courts will hear the oral argument for the Internet Archive's appeal.

What's at stake is the fundamental ability for library patrons to continue borrowing and reading the books the Internet Archive owns, like any other library.

Consider signing this Open Letter to urge publishers to restore access to the 500,000 books they’ve caused to be removed from the Internet Archive’s lending library and let readers read.

Learn more at: https://blog.archive.org/2024/06/17/let-readers-read/

mekarpeles··on Bookwyrm – A federated social network for reading books
very helpful, thank you. Good to learn international use case is working okay for you (I know we can improve). If you're not on our slack already, feel free to email me @ <mek@archive.org> -- you're welcome to ask questions and weigh so we can continue to try to move in the right direction for you and others.
mekarpeles··on Bookwyrm – A federated social network for reading books
That's awesome. May I ask, for pre-ISBN books, do you typically look up books by title? What do you do with books when you find them? What is your primary use case / reason? Is it as a reference library (of things to read)? Keeping track of reading?
mekarpeles··on Bookwyrm – A federated social network for reading books
Hi, can I ask what books were missing? The more examples we have, the more we can update our bots to make sure these books get either imported or fixed within our search engine.

I completely understand wanting to use a service that has the books you're looking for and would also completely understand if it's too much work to type up the examples. If you'd rather not do so publicly, happy to receive your email at <mek@archive.org> and do what I can to help. Thank you!

mekarpeles··on Bookwyrm – A federated social network for reading books
I disagree that having a popular flagship federated service negates decentralization.

The point of decentralization is not to destroy the ability to centralize (see e.g. git versus SVN and then look at github).

The advantage is that federation and decentralization empower an entirely new set of use cases, archival strategies, development, and accessibility affordances that may have not been possible before. While enabling the town square & metcalfe's law that are advantageous to many people / use cases.

mekarpeles··on Bookwyrm – A federated social network for reading books
Mek here from Open Library, thanks for your kind words.

Here! Check our @cdrini's https://openlibrary.org/barcodescanner

mekarpeles··on Bookwyrm – A federated social network for reading books
Hi, Mek here from Open Library.

I was incredibly lucky to work at the Internet Archive the same time as Mouse and couldn't be more proud of their work on BookWyrm.

Open Library and its network of generous volunteers have (I hope) made a lot of positive progress towards cataloging the books that are out there and making them more accessible to the world. AND it's absolutely the case that our project exists to support innovative projects like Bookwyrm and incredible thinkers like Mouse.

Open Library can't and shouldn't be everything. It's hard enough doing well at one thing. The Open Library team is considering how we may be able to participate within the decentralized ecosystem by offering a BookWyrm instance so readers may have more ways to socially engage with each other and connect around books. If you're interested in helping us try this as an experiment, please reach out <mek@archive.org>!

I appreciate how difficult it is to run a service which gives communities voices (it requires moderation tooling, staff, and so much more). I'm impressed by the thoughtful, impressive, and creative work Mouse has done building BookWyrm and am super grateful for its progress which I see as being a win for the entire ecosystem (an ecosystem Open Library is proud to be a piece of).

Keep it up <3

P.S. the fact that many services like Mastodon or BookWyrm may have large primary servers is not a demerit. The fact that there are smaller local servers, that new servers can emerge over time, and that engineering thought is being put into how data moves through such environments is key to acknowledging the importance of creating safe communities, promoting archival strategies, and enabling accessibility. Many people use GitHub (centrally) and also use Git (centrally) and the fact that many common use-cases have been centralized do not undermine the significance of the times where small, high impact cases are able to succeed because decentralization has made them possible.

mekarpeles··on Serious Language Learning with AI
Am curious what languages are supported
mekarpeles··on Serious Language Learning with AI
Great idea, using AI to Pimsleur-ize anything.
mekarpeles··on Open Library
Readium LCP will be an alternative to adobe ADE.
mekarpeles··on Open Library
Internet Archive is implementing readium LCP which is an open drm standard with many competing readers like Thorium. No adobe ID required.
mekarpeles··on Open Library
Thank you Ray!

https://openlibrary.org/volunteer

We have community calls ever Tuesday @ 9am PT and design calls 9am PT Friday!

We're currently talking about 2023 planning and everyone's feedback is welcome.

- mek

mekarpeles··on Open Library
By buying their books! The books on internet archive have been purchased by or donated to the library, the same way all libraries work across the US. Books are lent using the same 1:1 owned to loaned ratios. Open library is a catalog and publicizes info about authors and promotes their works, irrespective of whether or not lendable titles are available.

- mek (open library team)

mekarpeles··on Open Library
How can we improve the reading experience? Have you tried the Open Library Reader app? Looking forward to your advice!

Thank you!

- mek + open library

mekarpeles··on Open Library
Mek here with the Open Library team. In the past two years we've imported about 8 million modern books, including options to import your Goodreads books.

The community has merged 25,000 duplicate works and cleaned data for another 200k+

We also have a massive search improvement (exact edition search) slated to launch this month which is already on testing.openlibrary.org

If you decide to give it another try please let us know what you think and where we should focus our efforts to improve the experience for you!

mekarpeles··on Show HN: A more social, Amazon-free alternative to Goodreads
this is great work -- if you y'all need more book data (or non-google data) or have any interest in featuring free borrow links to titles which are digitally available from the Internet Archive's digital library, let me know and I'm happy to help.

- mek from Internet Archive's Open Library

mekarpeles··on Show HN: A more social, Amazon-free alternative to Goodreads
mek here from Open Library -- not sure what data booqsi is using but we're very happy to share our catalog's data publicly and freely with the world

https://openlibrary.org/developers/dumps

mekarpeles··on Bitcoin is a Ponzi
I think many people in these comments are bound to talk past each other when the real answer is:

Yes, Bitcoin is a Ponzi scheme (by virtue of people making it so), but that does not, nor should it, imply that the only thing Bitcoin is, is a Ponzi scheme.

Bitcoin is a Ponzi scheme in the same way that sociopaths exist in the game of geeks, mops, and sociopaths: https://meaningness.com/geeks-mops-sociopaths

It's part of the ethos but does not describe the whole ethos, and trivializing Bitcoin to a Ponzi scheme is the wrong answer. As, honestly, is trying to defend that Bitcoin is not a Ponzi scheme.

Irrespective of how value is in practice (or volume) making its way through the system is different from the system itself.

There are several promising elements of Bitcoin which contribute to its value. One is sheer access and connectedness (e.g. Metcalfe's law). People use facebook (its valuation reflects this) or pay for a phone plan because connecting to the Internet is a synecdoche for ubiquitous access. Creating a fabric where people in any country may work for a company, irrespective of its location, is a paradigm shift from what we had previously. Today, working for a US company from Taiwan is a legitimate challenge & deterrent.

There are many such examples, where transparent ledgers for instance may increase trust between parties and give people more confidence in financial robustness. All of these things have intrinsic value, just as the architecture which is PageRank has value for Google.

A common mistake I hear is for people to point to these advantages as if to suggest that it is without flaws or incapable of flaws, when we -- the people using the system -- are quite fallible and driven by prisoner's dilemma-type incentives which may by no means be long-term efficient or equitable (given we're all in the same boat). And I think these social challenges should be noted and not discounted on behalf of "technology being good enough".

So, I think this is a case of "Yes and".

mekarpeles··on Why has no one made a better Goodreads
This is great! Thank you for the interest!

We've had a rudimentary bot system for the past 10 years or so. Let's please talk if you where others would like to get set up with a bot account and write access!

https://github.com/internetarchive/openlibrary/wiki/Writing-...

Contact: mek@archive.org

mekarpeles··on Why has no one made a better Goodreads
Yes!

Here's the issue: https://github.com/internetarchive/openlibrary/issues/5022

Sorry about that, on the docket for this upcoming week.

We're currently staffed by two eng (myself included) w/ ~80 open PRs, 500 issues, just finalized a py2 to 3 migration, and reprovisioned all of our servers w/ docker so we're playing catch-up on some newly surfaced issues.

Please do open such issues so we're aware! It's a huge way to help.

The issue mentioned above is happening because our cron jobs got moved from one machine to another.

Will reply here when resolved.

mekarpeles··on Why has no one made a better Goodreads
Mek here from internet archive's OpenLibrary.org.

Open Library was started by @aaronsw.

We're a library catalog with 3M+ books to read & borrow.

We've been around for 15 years, not going anywhere.

We're open source and non profit: https://github.com/internetarchive/openlibrary

We defend patron privacy, offer free APIs, and release all public data openly: https://openlibrary.org/developers/dumps

Most projects on this page have likely used our data.

We have a Reading Log and several other more substantial features in the works.

Our catalog spans more than 20M works: https://openlibrary.org/stats

You can help! https://openlibrary.org/volunteer

mekarpeles··on Introducing the Open Library Explorer
To my knowledge, every preview includes at minimum the basic front-matter (such as those you describe). There are other mechanisms in place which enable specific controlled use cases e.g. limited page previews for folks coming from Wikipedia citations. I don't know if Open Library is presently taking taking advantage of these enriched previews -- I'll check with our product team to see what improvements to the experience may be possible.
mekarpeles··on Introducing the Open Library Explorer
Thanks for noticing, our primary goal is making sure a minimal experience catalog works (best we can) for folks in lower bandwidth areas, older devices, etc.

I've been meeting w/ John Gilmore (EFF) & RMS to make sure we're also testing, to the best of our ability, w/ privacy badger, librejs, and other tools to ensuring beyond nojs support that we have some modicum of assurance for folks who are willing to use js but appreciate a bit more care.

mekarpeles··on Introducing the Open Library Explorer
If you're able to describe how it's wrong, an issue on https://github.com/internetarchive/openlibrary/issues/new would be appreciated! Thank you!
mekarpeles··on Introducing the Open Library Explorer
Please see:

https://help.archive.org/hc/en-us/articles/360016554912-Borr...

There are multiple options for borrowing, depending on the title, including Adobe Digital Editions for applicable titles.

Page 1 of 6Next →