The Perils of ISBN
rygoldstein.com
rygoldstein.com
Sometimes these have different catalogue numbers or barcodes to distinguish them, sometimes they don't but they're still different. I've seen releases where the only difference is the label in the centre of the LP, or the back of the CD case has a two-column tracklisting vs a one-column tracklisting. Music publisher uses the same code and says it's identical and yet it's clearly not.
Then there's the "recordings" on an album, which even if they're never re-recorded can still end up chopped up, bleeped or remastered. They're not the same sound. MusicBrainz likes to track when they are exactly the same recording (e.g. the LP recording of a song appearing on a compilation album verbatim) and when they're not (e.g. radio edits of the LP recording). And if we're going beyond recordings by one artist of "their" song, i.e. cover versions, or just plain standards, those are "works", with composers, lyricists, and can be recorded thousands of times by different artists...
I greatly appreciate the pedantry and flexibility for noting down when creative works are the same versus where they differ, in relational database form.
[0] https://musicbrainz.org/release-group/1b022e01-4da6-387b-865...
I haven't looked into what their schema is like, but if it's anything like Musicbrainz it will be pretty comprehensive and easy to pull the data you want out of!
I've recently been doing data entry on Open Library... sometimes even worldcat doesn't have an OCLC for an edition, and Open Lib is my fallback. Maybe I should be doing it on Bookbrainz instead.
At this point, it's probably a bad thing that it exists, because if it didn't exist and you were trying to describe how you wish there was a good site that was like an IMDB for books that was editable like Wikipedia, then they couldn't say, "There is! It's called OpenLibrary", or dismiss your criticism by telling you that whatever is wrong with it you can help fix because it's both editable and the code is all open source.
"Officially, ISBNs should never be reused. However, problems can happen if:
- A publisher improperly reuses an ISBN
- A small or self-publisher mis-registers a book
- An ISBN agency error occurred
- A book was published before 2007 and conversion from ISBN-10 to ISBN-13 created confusion" [Source: ChatGPT]
In 2009, I had plans to use ISBNs to distinguish the books in my personal library. But after scanning some ISBN bar codes with a MacBook app, I discovered some codes were associated with different books (the app also pulled the cover art, so it was easy to spot). Never had the time to find out if the bar code scanning was defective (=did not use the check sum) or these were cases of assignment errors, which "shouldn't happen" but have already happened.
There is a certain type of ignorant developer who reused "unique IDs", I've even seen a database in production use where GUIDs were recycled (no joke).
We'd put pricing barcodes on every book in the store, and those were always based on the ISBNs and had the titles and authors printed on them, which was info that came from ISBN lookups either from Bowker's Books-In-Print data or Ingram's data. We'd print the barcodes in large batches and then have to match them to the books based on the title and author shown and verify with the ISBN, so all 80,000+ were checked, and the actual ISBN issues were _extremely_ rare and always from a _very_ small/amateur publisher.
Scanning the "ISBN bar code" then is actually leading to a lookup for digits of the UPC (a totally different sequence).
It's also the case that unless you were a used bookstore, you're fundamentally dealing with different kinds of inputs than someone cataloging a home library.
(By “decent enough” I mean breadth. If you are strictly collecting some genre products from a small number of commercial publishers, you might be in the walled garden where everything just works.)
SBNs were introduced when, in addition to existing mass production, mass accounting and storage management for each item became possible (with computers). Outside of the centrally controlled environments they don't work well, or mean much. Sure, national authorities make enough rules about having proper ISBNs, but they do get ignored.
There are small university/gallery/collective publications that have bigger print runs than “official” books on some specific topic. There are books that are uniquely made or uniquely altered, and therefore can't share the identifier with another item. Most common example is getting an autograph — you probably want to know precisely where you've put the copy of Bible signed by the author, not just any other Bible that looks the same. Some people oppose ISBNs for political reasons, and either ignore them, or invent bogus numbers.
Then there's International aspect. Soviet Union, for example, did not use ISBNs until the very last of its years. There are still many books printed there — including complete works every scholar needs to reference — that never had any ISBNs.
Some works have been published for that last time a century ago. Some of them might had been immensely popular back in the days, but now they are forgotten. Others have been re-printed, but you've managed to get the first printed edition, a small book of then-unknown author. Those also won't have ISBNs.
So the idea itself that any book must be an interchangeable product from the batch in which each item has the same effect, and therefore can have the same identifier, is a bit narrow.
Obviously, professional librarians could instantly tell you that ISBN is merely one of the search markers, and is not the way the inventory is kept.
30 29 28 27 26 2 3 4 5
As I understand it: That would be the 2nd impression, printed 2026. It's designed so the publisher can remove the innermost character(s) for each new impression, which I imagine was practical for printing presses - the type is already set, just remove a couple characters. Therefore the next impression this year would have, 30 29 28 27 26 3 4 5
The 4th impression, next year: 30 29 28 27 4 5
etc.I also included a BZR revision number but that’s more difficult to do with git as it doesn’t really have the concept.
Interesting to read that the reason for the removal was Cat Stevens' apparent endorsement of the fatwa against Salman Rushdie. It seems it was the band themselves that requested it? https://www.rollingstone.com/music/music-news/cat-stevens-br...
I wound up making an account, uploading the info, managing the 29 different reasons a neophyte makes a mistake causing their data not to be accepted, and finally got my CD into the system. This included using a random chinese persons web from the 90s who presumably had come to Australia and bought the identical pressing which appears to be a hyper-local market specific variant of the ones which other (European, American) markets got.
I have massive sympathy for the brainz, because as this article on ISBN and my experience shows, people are cavalier about renewing their 'unique identity' info, when they think they don't have to.
The best way to distinguish an album (after barcode or name/artist, and medium) is number of tracks, and if that's not enough, release year/country. I got my metadata by using their Picard tagger and the CD TOC (as it contains the number and lengths of tracks, it's much less ambiguous), but of course opening every case and putting every CD in the drive is a lot more effort than barcode scanning.
You can use the advanced search syntax if you need to look up multiple fields at once. https://musicbrainz.org/search -> Type "Release", method "Indexed search with advanced query syntax"
For example: barcode:075596073820 AND tracks:11 gives https://musicbrainz.org/search?query=barcode%3A075596073820+...
To me it makes the most sense to index music by its fingerprint. Releases, EPs, etc should just be pointers to that.
BTW, they misunderstood their own example of "Hotel Iris" by Yoko Ogawa when they wrote "the same work is duplicated four times." In fact, those four entries in the list point to distinct works.
One of these is a French publication by the publisher Actes Sud. Translations are not the same work as the original. They are derived works.
But it's true this list is a mess. Another entriy has 3 editions, one in English and two in Spanish, so it's obviously an error that mixes two distinct works.
Which adds yet another layer. Because you still want them to be considered as part of a larger single entity. If you're performing a search, you want to find the single main entity, and then have different translations listed the same way you have different editions listed.
In Openlibrary specifically they should be combined as one work. The editions can store the language and the translator info.
The current grouping is probably because semi-automatic (and some manual) merging is easier for titles in the same language.
https://openlibrary.org/help/faq/editing#works-special-cases
So an abridged or bowlderised or annotated or illustrated version are collected under the same work, even though people might have good reasons to want one over another (the language used and the specific translator being just two important attributes)
But summaries or adaptations or plays and screenplays are not.
There's always gray areas, but note the edition info isn't lost, it just lives in a subordinate position that is linked directly from the work.
I worked on the library systems and one of my innovations was to use the ISBN mapping database of WorldCat to find books with identical content but different ISBNs to help kids find the books on the list.
Over ten years that one SQL join in the code made the kids read an extra million books they wouldn’t have otherwise.
My biggest “bang for buck” in my career!
The ISBN alternatives table was just groups with an integer ID shared by each group. My import process synthesised this from worldcat data, which was more messy.
Regarding the taxonomy of WEMI (work, expression, manifestation, and item), all of them are useful since we are talking about books at different levels. From "I have read Don Quixote", which is about the work (translations are the same), to "My Don Quixote has coffee stains", which is about the item.
https://annas-archive.li/blog/all-isbns-winners.html
AA makes use of a WorldCat scrape they performed:
https://annas-archive.li/blog/worldcat-scrape.html
I'm currently working on an ISSN database, which is periodicals, and is grossly underrepresented compared to books.
Ability of an ISBN search of my collection would have helped me in this case - scanning a barcode is easy enough task to accomplish.
And even if I had a different edition, the resulting title from searching for a different edition would be enough to help me figure out that I should not buy a book I already own.
If you can really remember every single one of your nearly a thousand ebooks that you've bought, that's both impressive and baffling to me.
As a hobby, I do a lot of wood working. Recently I was acquiring some books on wood finishes. I accidentally bought one book three times - not all at once but over a several year period. I realized this while trying to organize.
Mainly because they can have very similar / generic titles, and be by different authors. Since I’m using them as a reference, even if I know the author, I can’t always remember if I own this book by them or maybe I own this book with very similar title by a different author.
I did eventually solve my duplicate book problem by making ourselves a searchable list I can access remotely, so now I can just look it up when I'm at the store.
Being deliberate about obtaining them isn't even remotely related to any of that.
I'd imagine this is also more specific to them being physical items, since it's much easier and obvious to look up an ebook if you just look at a list of files wherever you keep them.
- searching by title, ie: "The last unicorn" will return books across many years, and many editions, and with lots of different titles, examples:
The Last Unicorn (thorndike Press Large Print Science Fiction Series) THE LAST UNICORN The Last Unicorn (40th Anniversary Edition) The Last Unicorn the Lost Journey The Last Unicorn: The Lost Version The Last Unicorn das Einhorn im Spiegel der Popkultur
and then books that have a similar title but are by completely different authors:
The Last Unicorn: A Search for One of Earth's Rarest Creatures
- there's no way to programatically link an ISBN or ISBN13 to all of the other variants of that book across years or editions ( "First Edition", "1st U. S. printing", "6th Printing", etc..) or bindings ("Hardcover", "Mass Market Paperback", "Library Binding", "Kindle Edition", "Audio Cassette", etc..) or languages ("en", "English", "zh", etc..)
- I wrote some code that would consume the 1000 items in the ISBNDB API search results, and attempt to reduce the list of search results based on the the language, the title, and author(s) using Jaccard similarity, and then sorted by year, and grouped by binding, which mostly worked to be able to see all editions for a book, but it's super messy.
Going to have to see if I can use OpenLibrary instead, looks like a great option.
However, within the block publishers can assign ISBNs to different imprints.
The ISBN always indicates the country it's from, the United States getting the biggest block, other European nations and Japan getting their own, with Africa, the Middle East, and so forth all getting a block in common.
See https://en.wikipedia.org/wiki/List_of_ISBN_registration_grou...
For example, compare the most recent edition of 'Straight and crooked thinking' with the one published in 1930.
I "grew up with" a specific translation of Lord of the Rings into Norwegian, for example. There are two. They are very different. But the editions also differ in whether they include the appendices, whose illustrations are used, and more.
Are we talking material plot or characterisation changes?
>English: I will not tell anyone the secret.
>Bokmål: Jeg skal ikke fortelle hemmeligheten til noen.
>Nynorsk: Eg skal ikkje fortelja løyndomen til nokon.
Source: https://www.visitnorway.com/typically-norwegian/norwegian-la...
The title differences are also a good illustration of how different it can be:
Bokmål: Kampen om Ringen, Ringenes Herre (the first one is literally "the battle for the ring")
Nynorsk: Ringdrotten
An example is the name Bilbo Baggins. In the "canonical" Norwegian translation, he's become Bilbo Lommelun. "Lomme" means pocket, and "lun" means snug, warm, or comfortable. It's not literal, but it fits the nature of hobbits well while referencing the "bag" in Baggins", and the connotations comes immediately in Norwegian without having try to deconstruct the name.
In this case, I think the newer "canonical" translation is generally considered unambiguously the best, but people often have favourite translations. E.g. my favourite Scandinavian translation of Walt Whitman's Leaves of Grass isn't even Norwegian, but an old Danish translation which sounds much "softer" (it's hard to explain)
[0] Before anyone says it, I'm sure some bible nerd has numbered them, it's hyperbole.
Then click on the item and drill down into editions sorted by year, or whatever.
But when you're doing search, it's terrible UX to be flooding it with tens of editions mixed in with other things with similar titles.
The author misunderstands 'work', as far as I know: A work is "intellectual or artistic content of a distinct creation. It refers to a very abstract idea of a creation e.g. Shakespeare's Romeo and Juliet and not a specific expression."[0]
In contrast, an "expression" is an "intellectual or artistic realization of a work. The realization may take the form of text, sound, image, object, movement, etc., or any combination of such forms."[0]
The Last Unicorn story is the work, "the book The Last Unicorn" is an expression as would be the film version or the computer game, etc.
[0] https://www.ifla.org/references/best-practice-for-national-b... (as of a few years ago)
tl;dr; - The ISBN is intended to be a physical Part Number, within the book business. Where "hardcover, or paperback, or trade paperback, or large print, or revised edition, or ..." very much matters.
This isn't even the half of it. On some digital books, I'll find a dozen ISBNs in the front matter. Of course there's the hardback, the clothbound (not always the same as the hardback), the alk. paper variant, paperback, trade paperback, epub, pdf, "Adobe digital", and "master digital e-book" (no idea what that even is myself). And that's all just issued together. If they reprint, it won't get a new ISBN, but if the rights convey to another publisher, that one will get a whole 'nother set again. Some popular titles likely have low hundreds of ISBNs, and keep in mind that these have only been a thing since the late 1960s (9 digit ISBNs, technically just SBNs back then). Then with the now dead paperback trade, you could go through a dozen different covers for the most popular books (King, etc) but they'd all use the same ISBN.
Then, and this one bites me the most... if archive.org scans in a hardback with its ISBN, what do I use for the scanned pdf? I've decided that for lack of a better alternative I have to use it, but if the publisher made their own pdf (even just scanning the hardback), then it is supposed to issue a new ISBN to it.
Cataloging my own library, I've had to use a hodgepodge of unique ids. ASINs, ISBNs, Worldcat's OCLC numbers, Open Library's, and a few others besides. And it still comes up short. The number of oddball publishers and pamphlets and so forth that have never been cataloged anywhere is enormous.
The scanned pdf just doesn't have an ISBN. ISBNs are assigned by publishers to products for inventory management. That's it. If archive.org scans a book, it's not a product that needs inventory control.
Personally, I have never had all these indicators match in any book. It also allows you to find a very specific publication using a semantic search, specifying a combination of tags/publisher/formats.
Archive.org would recommend using the OpenLibrary IDs instead of ISBNs. (OpenLibrary is an Archive.org project.)
> The number of oddball publishers and pamphlets and so forth that have never been cataloged anywhere is enormous.
I think it's more the case that number of catalogs is too many. At least with LibraryThing it always seems like somebody has cataloged everything, but we have such a hodgepodge of ID systems and catalog numbers in part because so rarely have all the catalogs been connected or have tried to be connected. It's only a relatively recent library phenomenon that so many small library catalogs can talk to each other on the same protocol, much less coexist in the same broader search tool.
> Cataloging my own library, I've had to use a hodgepodge of unique ids. ASINs, ISBNs, Worldcat's OCLC numbers, Open Library's, and a few others besides.
In part because most of my personal catalog is in LibraryThing, I've been impressed with LibraryThing's Works ID as a generally trustworthy unique ID for a book. LibraryThing benefits from an interesting mix of volunteer and professional librarian work (especially the work of a lot of tiny and interesting niche libraries across the world) in deduping and merging editions together into the same Work ID. StoryGraph and OpenLibrary are also doing interesting things in this space, but LibraryThing has the momentum of time (it's as old as GoodReads and not an Amazon side project) and the benefit of extra (nerdy) labor.
I also like the LibraryThing IDs because they are generally short, opaque (which is a weird feature sometimes), and don't look anything like an ISBN because they aren't intended for that. StoryGraph's IDs are GUIDs, which I will forever find ugly in their normal - delimited hexadecimal rendering. Open Library's look like ISBNs for reasons that I don't understand, but I do appreciate that you can use the last letter of the ID to distinguish between an edition ID (ends in M for reasons I don't know why) and a work ID (ends in W), and the OL prefix does help them stand out next to other catalogs' IDs.
I built a voting website for my current favorite book club and I thought I could do everything with just the LibraryThing Works ID but then I keep adding other IDs to the "database" (YAML frontmatter) as time goes on. LibraryThing doesn't have a Covers API because most of their edition covers come from Amazon and Amazon is restrictive on that. If I add the OpenLibrary Edition ID, I can use the OpenLibrary Covers API as Archive.org has very nice terms on that today. (Not the OpenLibrary Works ID, because covers are associated at the Edition level, which does make some sense, but the website UI shows a default cover from a random edition so I'm not sure why the API couldn't return that cover from the Works ID, but it is nice to pick and choose Edition covers anyway and I can't complain too much having a working cover image API from someone.) I started adding StoryGraph IDs because members of the club love StoryGraph right now and also because while StoryGraph doesn't have an Official API yet (it is on the Roadmap), I discovered StoryGraph's CWs section was amenable to easy scraping. I figured since an API for it is on the Roadmap a bit of light scraping (with attribution!) was fair. (My club wanted CW information to help decide on book voting. LibraryThing intentionally doesn't track CWs as too hot button and subjective, but StoryGraph has a rather nice "voting" experience for CWs and before I started to scrape StoryGraph's CWs we were already starting to copy and paste them by hand into the Markdown documents. The scraping provides better attribution and a unified display.)
https://www.wikidata.org/wiki/Q106545884
Gives us the ISBN, Goodreads work id, LibraryThing id, the OpenLibrary id, and the Google Knowledge Graph id all in one query.
I shop by ISBN often because I want specifically a particular edition in a particular cover. So it’s not just title and author. It’s not even title, author, publisher, edition, and cover honestly. Sometimes there’s an Indian subcontinent English printing of a book that’s laid out differently and on different paper from the US/Canada market version.
One small drawback is sometimes I’ll order a book by ISBN, and the bookseller will locate it by ISBN, and it will be a completely different item on a different topic by a different author. Sometimes if a book is a small printing or is a very old title the publisher will recycle the ISBN.
There is. https://hardcover.app
I used Letterboxd a lot before kids. I used Goodreads until the Trump inauguration when I de-Amazon'd myself as much as possible (Amazon owns Goodreads). I switched to Hardcover, which is a much better interface. There are ways to improve, but overall I prefer it over Goodreads.
I'm not sure there is anything I'd improve. I recall when I started creating a list of Akutagawa Prize winners in translation, it was a bit painful because I had to input a lot of ISBNs for rarer books. Also I struggled for a couple books with picking the right version so the page count was correct.
But I haven't done either of those in a while since I haven't read more rare stuff in a year. Possibly they've gotten better, too.
Take To Kill A Mockingbird as an example... No matter what (English) edition of the book you read, you're likely reading the exact same content, even the exact same words, as any other English edition. There might be a different preface near the front or different blurbs on the back cover or a different number of words per page, but the actual story is word-for-word the same. A simple title lookup makes sense here in most cases.
Compare that to something like The Iliad, where the English versions are all translations and can vary greatly from translator to translator. While all telling ultimately the same story, a bad translation doesn't begin to compare to an elegantly beautiful translation, so you almost certainly don't want to treat all editions of The Iliad the same.
Translations aren't the only times that you wouldn't want to treat all editions of a title the same. Some books have undergone abridgments, revisions, or corrections, so the content won't be word-for-word the same between editions, but might or might not be close enough that it's worth considering them the same. Some books have heavily annotated editions, so while not changing the underlying content that all the editions are based on, the reading experience is quite different.
I could go on with differences, but I hope it's clear that there _are_ differences between books and movies when it comes to variations/releases. For books, I think the lookup issue is closer to how it is for board games. Board games, like books, have many editions and translations and often get updated/revised between editions. Sometimes the updates change the gameplay significantly, and other times they don't. Boardgamegeek.com is one of the best (if not _the_ best) catalogs of board games that there is, and it has regular discussions/arguments about whether a new edition of a game is different enough that it deserves its own page or if it should just be relegated to be an easy-to-ignore note in the Versions section of the previous version's page. I think a letterboxd-like lookup for books would have similar regularly-occurring debates, and, like with board games, ultimately have to be fairly hand-curated.
ISBN is a an attribute/key, but not primary key, in database terms :)
ISBNs are messy and in real world you’ll see crazy amount of broken/edge cases that shouldn’t happen by the letter of the standard, but happen all the time in reality.
* For example, isbn can be reused by publisher for completely different book.
* 2nd edition, while very different, may have same isbn.
* Reissue of the same book could have different isbn.
* Textbook of same author for 6th and 7th grade could have same isbn.
* As soon as you’ll get in translations all bets are off.
* I already mentioned textbooks. How anbout about college books where each year there was slightly revised edition of same book.
If you ask yourself - wtf? You’re not alone.
—-
In my youth I heard horror stories about people who suddenly found multiple duplicate guids (uuidv1) in their databases because cheap Chinese knockoff network cards were using same MAC addresses. Think that with isbn that could Happen to you any time.
Honestly, right now I probably wouldn’t even try to code complex algorithm of book matching but fed all of books metadata, including book covers etc to llm and it would do better than what we had.
Our algorithm had tons of special cases coded and in results ui there was a button “needs manual review”, that was launching review workflow (not a joke, business people has special support team in India, because we were matching not only books) for cases when confidence score was low.
For my language, I know I can only rely on just a couple of shops that make a photo of an actual book under regular lights for catalogues and announcements as a matter of principle. Others just add the same image file everyone else used. For the older books that had multiple revisions, it's still strictly manual interaction with second-hand book sites and old listings to figure out the exact version.
On the other hand, most services here do care about providing table of contents, even if its crowd-sourced phone snapshots of the page. International market is insane in that regard. Not only sellers on Amazon and the like ignore them (sure, they were never going to do so much manual work for all the uploaded items anyway), the publishers don't even mention what's inside on their official pages. “Selected stories. New translation by X”. Are those the same stories that were published 5 or 10 years ago under a different name or cover? Has X translated a new set of stories and (partially) refreshed the collection? Both of these things happen. It is ironic that I need to download the full scan from the pirate library to find out what is available in print. Unless, of course, some anonymous worker in some library hasn't diligently typed the whole list of works into the description, hooray for those.
Two great features are: season names and episode groups. The other day there was a thread about Babylon 5, where seasons have names and the watching order is different from the airing order. Perfect application of both
I'm mildly surprised at exactly how successful ISBNs are. I worked in a book wholesaler's warehouse 35 odd years ago and the ISBN was used as the product code by the "system". I'd get a series of picking lists for pallets on good old green "staved" fan fold. I'd whizz around the warehouse with my trolley and pick from paper packets of books. The product lines had the rack and bay, last four from the SBN, quantity, title and full SBN. The packets of books had the rack/bay/last four from SBN printed on a label in large and small other details. I got very good at optimising my course around the warehouse and could pick at a right old rate, whilst listening to my mini cassette player. Its pretty boring work so you might as well game it!
Sometimes an individual book might fall off my trolley and be dumped in the big cardboard "skip" for rejects. For some reason casualties around me generally involved subjects like maths, material sciences, geology, surveying, hydrology. Oh and fractals!
I graduated in civil engineering.
Anyway. Surely all of us here know that really getting to grips with defining what it is that you are cataloguing/indexing/numbering/whatever and why can be quite tricky.
Both Dewey and SBNs catalogue "books" but for very different reasons. Both systems are extremely successful. You might think that in our world of LLMs n that, that books, Dewey and SBNs will go the way of the dodo.
Perhaps, but I doubt it.
Right, bugger all this old school nonsense. I've got a C64 (it rocks a SD card interface and a HDMI out (via SCART - must sort that out)) blinking away on my telly in the sittingroom and some mutant camels need a bloody good kicking.
For example, here's a non-exhaustive but still pretty long list of movies with reused titles: https://www.imdb.com/list/ls083468410/
But there is at least one case where it was on purpose. There's a set of reading primers from the UK called the Biff, Chip and Kipper books. We acquired a whole set of them at a garage sale, and when I went to enter them into my catalogue, I discovered that the publisher had assigned just one ISBN to the whole series. Which quite annoyed me when I discovered it. (I ended up just not cataloguing those books, because I didn't want to type the titles, author, copyright date, etc. in by hand for 50+ tiny books).