238 karma · joined September 18, 2009
Readability (Mozilla) is used when you open regular web pages, but PDF.JS (also Mozilla!!) is used to render each page, run YOLO region detection and then generate a clean HTML. Math formulas are rendered as Latex, and tables are rendered into HTML by detecting their structure using Texo/FormulaNet.
Finally, you can read out loud the readable articles/PDFs using another local TTS (I prefer PocketTTS personally, but wanted to try the newly released and tiny Inflect TTS as well, so it's an option).
Let me know if you find this useful.
I should probably publish them to the Chrome and Firefox stores.
Converted to ONNX of course, and also CoreML for the iOS app. I haven't yet tried on Android. Does it work for you if you have an Android phone?
Hope you find this useful. The original version used heuristics, but frankly it was breaking on many more PDF that I'd want to admint.
This new AI-based version that uses a YOLO detector trained specifically on Doc Layout dataset seems to do a very very good job.
It’s using a local YOLO detector trained specifically on detection pdf page regions.
https://www.appblit.com/pdfreflow
The old version works and has many users who love it to read scientific papers, but its heuristics based and was in my opinion failing on edge cases that this new AI approach solves.
(It's free up to 20 articles because there are real costs: I use Gemini to summarize the pages you open)
AI voices run locally on your iPhone/iPad (web extension version coming soon).
When you find something useful, you can share the overviews online (free hosting), e.g. https://voiceview.app/a/2J49UnwK
Hope this helps cut the noise and help folks save time.
Laurent
So you're telling the caller that it is an AI, and yet you can have a pleasant background audio experience.
OpenAI realtime voices are really bad though, so you can also configure your session to accept AUDIO and output TEXT, and then use any TTS provider (like ElevenLabs or InWord.ai, my favorite for cost) so generate the audio.
Or PipeCat Cloud / LiveKit cloud (I think they charge 1 cent per minute?)
Same with TTS: some like Deepgram and ElevenLabs let you stream the LLM text (or chunks per sentence) over their websocket API, making your Voice AI bot really really low latency.
That’s why I created a stack entirely in Cloudflare workers and durable objects in JavaScript.
Providers like AssemblyAI and Deepgram now integrate VAD in their realtime API so our voice AI only need networking (no CPU anymore).
Runs at around 50 cents per hour using AssemblyAI or Deepgram as the STT, Gemini Flash as LLM and InWorld.ai as the TTS (for me it’s on par with ElevenLabs and super fast)
I wrote this webapp that uses this method: it calls Gemini in the background to polish the raw transcript and produce a much better version with punctuation and paragraphs.
https://www.appblit.com/scribe
Open source with code to see how to fetch from YouTube servers from the browser https://ldenoue.github.io/readabletranscripts/
To solve the server IP sometimes being blocked by YouTube, the app fetches the transcripts in the browser.
200 free minutes on signup so you can try for free.
LLM corrected transcripts are really good, and you can highlight text which I find super useful to study and share quotes.