HNHacker News
TopNewBestAskShowJobs

smoores

321 karma · joined February 3, 2023

submissionscomments
smoores··on Automating Immersive Reading
Viterbi and Needleman-Wunsch are essentially the same algorithm, developed in parallel for two different domains! The Viterbi formulation of the algorithm is the one usually applied to signal decoding, since that's what it was originally designed for.
smoores··on Automating Immersive Reading
I'll post again when we release the v3 apps — even in alpha, they're really awesome, I think they blow the Kindle app out of the water!
smoores··on Automating Immersive Reading
Oh, cool! Yeah that seems like a good application.

The current Storyteller alignment algorithm actually does do just that! We use Whisper to transcribe the audio to text, and then use error-align[1] to align on the text.

There are a few disadvantages to this approach:

1. Whisper only supports ~25 languages, and only about 10 of those very well. We want to support more languages, and Massively Multilingual Speech supports "1000+" 2. Whisper's timing outputs are not very good. We want to do word-level highlighting, like in the demo at the top of the post, but in order for that to be a good user experience, those timings need to be very precise. Much easier to do that with CTC!

CTC Viterbi is the tried and true forced alignment algorithm for good reason. It's not really that it's heavier duty than running Whisper and aligning on the output. Rather, it's like you stop Whisper early, before it does the final step of actually producing text, and step in and say: take the data you just calculated and use it to produce _this_ text, specifically. And then, since it produced _your_ text, you don't have to do anything else, you just use the timestamps directly.

The only reason Storyteller never used it in the past is because I couldn't come up with a good way to do the boundary search I describe in this post! This is super important for books in a way that it may not be for your oral reading transcript use case, because chapters can be (and often are) out of order between the ebook and audiobook. But once I worked out the n-gram RANSAC approach, it became much more tenable.

[1]: https://github.com/corticph/error-align

smoores··on Automating Immersive Reading
You gotta give me more than that haha. It's not readable in what way?
smoores··on Automating Immersive Reading
Well, first of all, lots of people specifically do use this because they have print disabilities or neurodivergence that makes it challenging to read long stretches of text.

Also, disability isn't binary? Some people have an easier time reading text than others — is your argument that books should remain less accessible to those people for whom it's harder because you don't personally feel like they should need it?

And then there are _lots_ of people that just enjoy reading this way. Audiobook production is its own art form, and many people like experiencing the text and the audio together. For some people it helps them focus in a noisy environment, for some people they find they can read faster with the narration, and some people just enjoy it.

Every single time I post about Storyteller, there are multiple people in the comments insisting that this shouldn't exist because no one should ever consume books differently from how they do. I don't understand it, honestly. If you don't want to read books this way, that's totally fine! You can even still use Storyteller — it works great with plain EPUBs. Why the need to make others feel bad for the way they engage with stories? The world is a better place if more people read more books — discouraging people from reading books in the way that feels pleasant and accessible to them makes the world worse.

I think if you feel compelled to instruct people you don't know to "stop their current habits" because you don't personally understand their needs or wants, perhaps you should reconsider your current habits, yourself.

smoores··on Automating Immersive Reading
My guess is, like many accessibility tools, it will vary by person! The Storyteller apps don't actually support this multi-level granularity demostrated in the demo here — until this iteration of the alignment algorithm, the timing wasn't good enough for word-level highlighting. So currently we only do sentence-level highlighting.

When we do roll out multi-level granularity, it will indeed be something that you can configure yourself, including how each level is indicated (e.g., you might want to set a background color on the sentence and underling the word) and whether a each level is indicated at all!

The "spread out" highlighting is a really neat idea, I don't think I've heard that suggestion before!

smoores··on Automating Immersive Reading
Sure! Some people have print disabilities like dyslexia and neurodivergence that makes reading text for an entire novel-length book challenging. Other people are perfectly competent print readers, but find that they enjoy having their book read to them. Audiobooks are their own art form, and it's nice to be able to enjoy them alongside text.

"Immersive reading" seems to be the industry term for this feature — I used it here because I though it was most likely to be recognized by a wide audience. Personally, and within the Storyteller ecosystem, I call it "readaloud," which I think is at least a little bit more useful of a phrase.

smoores··on Automating Immersive Reading
Hm, I don't know what you're referring to. Is the text too small for you? I wrote the blog post on FF for Linux and FF for Android, and I didn't need to zoom in to review it, though some of the graphics do definitely end up with pretty small text on mobile.
smoores··on Automating Immersive Reading
Almost everyone who uses the readaloud features in Storyteller speeds up the audio playback quite a bit, often 2-3x, so that it matches their reading speed better. The Storyteller apps let you set different playback speeds for pure audio playback vs readaloud, since being able to see the text usually makes it easier to listen faster!

I supposed I should add: it's also, obviously, fine if this just isn't for you! Lots of people get a lot of value out of audiobooks (e.g., I really like to listen to audiobooks in the car or while running), but like to switch to reading when it's an option. And lots of people find readaloud super valuable, whether because they have a print disability or neurodivergence that makes reading challenging, or just because they like the experience of being read to. But lots of people are in neither of those groups!

smoores··on Automating Immersive Reading
You can use a KOReader plugin, https://github.com/stradichenko/audiobook.koplugin, which has work-in-progress support for Media Overlays (the EPUB spec that Storyteller uses for readaloud)!
smoores··on Automating Immersive Reading
... Okay!
smoores··on Automating Immersive Reading
There are two koreader plugins:

1. StorytellerSync (https://github.com/Sirozha1337/storytellersync.koplugin), which syncs your KOReader progress directly to your Storyteller server

2. Audiobook (https://github.com/stradichenko/audiobook.koplugin), which has a WIP media overlay implementation that works with Storyteller readalouds

smoores··on Automating Immersive Reading
Thanks! We're improving it all the time — we're working on big new releases ("v3") for the web app and mobile apps. There's a Discord server linked in the docs if you ever need any help out want to chat!
smoores··on Automating Immersive Reading
That is actually the exact use case that I originally built Storyteller for! I wanted to listen to my books on long runs, and then switch to reading when I got back.

The Storyteller mobile apps have great support for this, I think. They have fully fledged audiobook players and ebook readers, and you can switch between the two with one tap. And then of course you can also double tap on a sentence and start playing from there, with the app highlighting the currently read sentence.

smoores··on Automating Immersive Reading
I took a week of from work recently to reimplement Storyteller's forced alignment algorithm. Storyteller[1] is an open source, self hosted platform for creating, managing, and reading/listening to "readaloud" books — books that have audiobook narration built in and can highlight each sentence (and/or word, with this new algorithm!) as it's read aloud. Forced alignment is the process of determining where each piece of text starts and ends in the audiobook.

Anyway, I am really pleased with how the new algorithm turned out! Hopefully someone else finds it interesting as well.

[1] https://storyteller-platform.dev

smoores··on Making React ProseMirror Fast
Just finished a new blog post about React ProseMirror. Happy to chat if anyone has questions, hope you enjoy!
smoores··on We're building a better rich text editing toolkit
Hey folks!

Handle with Care is a software collective that builds and maintains open source rich text editing libraries, including React ProseMirror [1]. We all came from The New York Times’ content management system team, and we spend a lot of time thinking about rich text and collaborative editing.

Now we’re working on something new: Pitter Patter will be a fully featured collaborative rich text editing toolkit, with all of the bells and whistles you need for your own text editor.

The space we’re entering is not devoid of solutions — Lexical, Slate, ProseMirror, and Tiptap are all viable options for building modern, browser-based rich text editors. But we feel pretty confident that we’re going to be able to bring some value, nonetheless.

First of all, Lexical, Slate, and ProseMirror (especially ProseMirror, in our opinion!) are all excellent rich text libraries, but they are also quite low level. You can build nearly anything atop them, but you will have to do quite a lot of the building yourself. Sometimes that’s exactly what you’re looking for — in that case, Pitter Patter can still provide you some value, because we’re going to be releasing individual libraries (like a CodeBlock node view, advanced markdown serialization, and suggest changes) that interop with the existing ProseMirror ecosystem.

But if you want something that’s more batteries-included, you’re mostly left with Tiptap. Tiptap has been dominant in the space for a while, but we think we can do better!

- We’re building on top of React ProseMirror, a truly React-native ProseMirror view, that doesn’t have to make any of the compromises that Tiptap’s React integration currently makes [2]

- We have a deep understanding of ProseMirror’s internals (and we’re not afraid to use it!)

- Pitter Patter will be completely open source

- We’re building on top of prosemirror-collab-commit, the best (only?) rich text collaboration protocol that is both correct and fast [3]

Anyway, we’re posting here for two reasons:

1. Maybe there are some more collaborative rich text editing nerds here that will be exciting (or not!) to hear about this. Sign up for our newsletter if you want updates!

2. Maybe there are some companies that are looking for alternative solutions to what’s out there. Consider sponsoring us on GitHub [4], or reaching out if you want to be more involved!

[1]: https://github.com/handlewithcarecollective/react-prosemirro...

[2]: https://smoores.dev/post/why_i_rebuilt_prosemirror_view/

[3]: https://www.moment.dev/blog/lies-i-was-told-pt-2

[4]: https://github.com/sponsors/handlewithcarecollective

smoores··on Using React Transitions for low priority text editor updates
Howdy! React ProseMirror maintainer here. Our collective has been helping out a client with migrating their existing text editor to use React ProseMirror from @tiptap/react. They had a very complex system for deferring updates to their miniature editor preview, which involved queuing ProseMirror transactions and applying them to a second Tiptap Editor during idle time.

While migrating to React ProseMirror, initially I tried out just passing the primary editor's EditorState directly to the preview editor's <ProseMirror /> component, but the top level node view components turned out to be just slow enough to render that rendering them twice on every keypress introduced a noticeable lag. So I added a useDeferredValue to render the preview editor in a Transition! Here's a post about how that works and the tradeoffs involved. I added some interactive demos to illustrate how the Transition changes the render flow.

smoores··on [dead]
I found myself recently needing to implement the HTTP range request protocol in order to support video elements in Safari. It took some effort, so I figured I would document what I learned in case it’s useful for anyone else!
smoores··on Why I rebuilt ProseMirror's renderer in React
Thank you for the kind words, I really appreciate it! I'm glad you enjoyed it
smoores··on Why I rebuilt ProseMirror's renderer in React
Howdy folks. This is a somewhat long (… sorry!) deep dive into several years’ worth of work that started when I was a staff engineer at The New York Times and has followed me into open source development in the years since I left to do my own thing! In it, I break down the issues we’ve faced while attempting to integrate React and ProseMirror — there are loads of code snippets and live demos in there. I hope that at least one other person finds this interesting!
smoores··on Show HN: @smoores/epub, a JavaScript library for working with EPUB publications
Thanks! Looking forward to your feedback!
smoores··on Show HN: @smoores/epub, a JavaScript library for working with EPUB publications
Readium is fantastic, and Storyteller uses it basically whenever possible. But Readium is exclusively for reading EPUB contents, and doesn’t have any support for modifying or creating them, which is the primary purpose of this library!
smoores··on Show HN: @smoores/epub, a JavaScript library for working with EPUB publications
Thanks, yeah I agree. The vast majority of the code in this library lives in a single class definition. Is it possible to move the implementations into separate files? Totally. Would that make the codebase more legible? I think at the moment, I would argue no, it would actually hurt legibility. If the class needs to grow dramatically, then maybe we’ll need a different approach, but I think this is actually the right thing to do for now!
smoores··on Show HN: @smoores/epub, a JavaScript library for working with EPUB publications
This is like... comically hard to express succinctly haha. We've been having several conversations about how to explain it quickly to new folks. I like the phrase that the Readium folks use, "guided narration", but I don't know how useful that is for folks that aren't already familiar with what Storyteller does
smoores··on Show HN: @smoores/epub, a JavaScript library for working with EPUB publications
I think that the thing we need to account for (which, number of words per chapter would capture this, I think) is different publications of the same book, which would need different overlays if they have different chapter filepaths, etc.
smoores··on Show HN: @smoores/epub, a JavaScript library for working with EPUB publications
Yeah I was toying around with that, too… but folks often mess around with metadata in tools like Calibre and Audiobookshelf in ways that wouldn’t have an impact on Storyteller’s sync, but would change their hash. On the other hand, I don’t know how various publishers handle EPUB dc:identifiers and that may not be robust enough, either. We could try doing something like hashing only the contents of spine items (including their file names, since that’s how media overlays refer to content)
smoores··on Show HN: @smoores/epub, a JavaScript library for working with EPUB publications
Is the idea that you have some devices that you want to download just the text to, but have it sync with your other devices? I think we could support that natively, honestly! Storyteller already has the input files, and it uses a text-based position system that doesn’t require the audio to exist. If you’re already doing work on this, maybe we could add it to Storyteller?
smoores··on Show HN: @smoores/epub, a JavaScript library for working with EPUB publications
That’s a really interesting idea! The more I think about it, the more I like it.

A challenge I foresee is that the media overlays are only reusable if you have the exact same input EPUB file, and have processed it with Storyteller to mark up the sentence boundaries. EPUBs have unique identifiers, though, so maybe this would be fine! We’d need to add a new processing flow to Storyteller, but it should be doable.

Feel free to hit me up in the Storyteller chat if you want to discuss more! Thanks for sharing this idea!

smoores··on Show HN: @smoores/epub, a JavaScript library for working with EPUB publications
You can’t do much better than that; that’s the size of the audiobook! For what it’s worth, I also used Storyteller on Wind and Truth, and got it down to 1.2GB by using the OPUS codec with a 32 kb/s bitrate.
Page 1 of 3Next →