HNHacker News
TopNewBestAskShowJobs

ldenoue

238 karma · joined September 18, 2009

Building web, iOS and macOS apps at https://www.appblit.com
submissionscomments
ldenoue··on Show HN: Readability for web and PDFs with TTS and local AI summary
extension for Chrome will be here: https://chromewebstore.google.com/detail/gpgkihaonhnhfabcmgn...
ldenoue··on Show HN: Readability for web and PDFs with TTS and local AI summary
weekend project that uses the latest Chrome Summarizer API, but also YOLO to process PDF pages and find images, tables, and math formulas.

Readability (Mozilla) is used when you open regular web pages, but PDF.JS (also Mozilla!!) is used to render each page, run YOLO region detection and then generate a clean HTML. Math formulas are rendered as Latex, and tables are rendered into HTML by detecting their structure using Texo/FormulaNet.

Finally, you can read out loud the readable articles/PDFs using another local TTS (I prefer PocketTTS personally, but wanted to try the newly released and tiny Inflect TTS as well, so it's an option).

Let me know if you find this useful.

I should probably publish them to the Chrome and Firefox stores.

ldenoue··on Show HN: PDF reflow in the browser with local AI model
do you have sample PDFs?
ldenoue··on Show HN: PDF reflow in the browser with local AI model
Not yet. for now this one from Alessandro works well: https://huggingface.co/Armaggheddon/yolo26-document-layout

Converted to ONNX of course, and also CoreML for the iOS app. I haven't yet tried on Android. Does it work for you if you have an Android phone?

ldenoue··on Show HN: PDF reflow in the browser with local AI model
this app runs entirely in your browser. thanks to Codex, it now uses a YOLO detector using ONNX to detect regions in each rendered PDF page image (text, pictures, formulas, tables). based on this analysis, PDF Reflow then cuts out TEXT regions into tiny little word images and adds back the original text from the PDF (when available) so you can still select and find text in the reflowed HTML render.

Hope you find this useful. The original version used heuristics, but frankly it was breaking on many more PDF that I'd want to admint.

This new AI-based version that uses a YOLO detector trained specifically on Doc Layout dataset seems to do a very very good job.

ldenoue··on Ask HN: What are you working on? (June 2026)
Cooking a local AI version of PDF Reflow to show PDFs on mobile and keep all the look and feel (font, formulas, pictures, tables) yet formatted for a smaller screen.

It’s using a local YOLO detector trained specifically on detection pdf page regions.

https://www.appblit.com/pdfreflow

The old version works and has many users who love it to read scientific papers, but its heuristics based and was in my opinion failing on edge cases that this new AI approach solves.

ldenoue··on OpenScreen is an open-source alternative to Screen Studio
I created QuickScre because I wanted a no editing way of recording polished screen recordings for Slack etc. Free to try https://www.appblit.com/quickscreen
ldenoue··on OpenScreen is an open-source alternative to Screen Studio
I recently released QuickScreen give it a try for free https://www.appblit.com/quickscreen and one time purchase for $7.99 lifetime
ldenoue··on Show HN: VoiceView – Instant Audio Overviews (web, YouTube, pdf, X articles)
I grew tired of endless YouTube videos, X articles or web articles. So this app lets you open any link and you instantly get an AI summary + brief about the content.

(It's free up to 20 articles because there are real costs: I use Gemini to summarize the pages you open)

AI voices run locally on your iPhone/iPad (web extension version coming soon).

When you find something useful, you can share the overviews online (free hosting), e.g. https://voiceview.app/a/2J49UnwK

Hope this helps cut the noise and help folks save time.

Laurent

ldenoue··on Asterisk AI Voice Agent
It doesn't have to be. You can configure your bot to great the user. E.g. "Aleksandra is not available at the moment, but I'm her AI assistant to help you book a table. How may I help you?"

So you're telling the caller that it is an AI, and yet you can have a pleasant background audio experience.

ldenoue··on Asterisk AI Voice Agent
I don't but I should open source this code. I was trying to sell to OEM though, that's why. Are you interested in licensing it?
ldenoue··on Asterisk AI Voice Agent
I am not using speech to speech APIs like OpenAI, but it would be easy to swap the STT + LLM + TTS to using Realtime (or Gemini Live API for that matter).

OpenAI realtime voices are really bad though, so you can also configure your session to accept AUDIO and output TEXT, and then use any TTS provider (like ElevenLabs or InWord.ai, my favorite for cost) so generate the audio.

ldenoue··on Asterisk AI Voice Agent
Check out something like LayerCode (Cloudflare based).

Or PipeCat Cloud / LiveKit cloud (I think they charge 1 cent per minute?)

ldenoue··on Asterisk AI Voice Agent
Yes DO let you handle long lived websocket connections. I think this is unique to Cloudflare. AWS or Google Cloud don't seem to offer these things (statefulness basically).

Same with TTS: some like Deepgram and ElevenLabs let you stream the LLM text (or chunks per sentence) over their websocket API, making your Voice AI bot really really low latency.

ldenoue··on Asterisk AI Voice Agent
I built a voice AI stack and background noise can be really helpful to a restaurant AI for example. Italian background music or cafe background is part of the brand. It’s not meant to make the caller believe this is not a bot but only to make the AI call on brand.
ldenoue··on Asterisk AI Voice Agent
The problem with PipeCat and LiveKit (the 2 major stacks for building voice ai) is the deployment at scale.

That’s why I created a stack entirely in Cloudflare workers and durable objects in JavaScript.

Providers like AssemblyAI and Deepgram now integrate VAD in their realtime API so our voice AI only need networking (no CPU anymore).

ldenoue··on Asterisk AI Voice Agent
I developed a stack on Cloudflare workers where latency is super low and it is cheap to run at scale thanks to Cloudflare pricing.

Runs at around 50 cents per hour using AssemblyAI or Deepgram as the STT, Gemini Flash as LLM and InWorld.ai as the TTS (for me it’s on par with ElevenLabs and super fast)

ldenoue··on Show HN: Building a full-stack Cloudflare starter kit (Hono and D1 and Stripe)
Would be useful to get a preview of the code
ldenoue··on ShowHN: YouReadTube
Which browser and computer ?
ldenoue··on ShowHN: YouReadTube
YouReadTube is the new name because it’s easier to remember and also insert “read” on any YouTube link
ldenoue··on ShowHN: YouReadTube
In browser transcript beautification using a mix of small models (Bert, all-MiniLM-L6-v2 and T5) for restoring punctuation, finding chapter splits and generating the headers.
ldenoue··on Yt-transcriber – Give a YouTube URL and get a transcription
Check out https://ldenoue.github.io/readabletranscripts/ and the website https://www.appblit.com/scribe that use Gemini to post correct the raw transcripts
ldenoue··on Yt-transcriber – Give a YouTube URL and get a transcription
Unless you fetch directly from your browser. It works by getting the YouTube json including the captions track. And then you get the baseUrl to download the xml.

I wrote this webapp that uses this method: it calls Gemini in the background to polish the raw transcript and produce a much better version with punctuation and paragraphs.

https://www.appblit.com/scribe

Open source with code to see how to fetch from YouTube servers from the browser https://ldenoue.github.io/readabletranscripts/

ldenoue··on Fei-Fei Li: Spatial intelligence is the next frontier in AI [video]
Full playable transcript https://www.appblit.com/scribe?v=_PioN-CpOP0
ldenoue··on Andrej Karpathy: Software in the era of AI [video]
Full playable transcript https://www.appblit.com/scribe?v=LCEmiRjPEtQ
ldenoue··on Elevenlabs Conversational AI 2.0
How is the turn detection working? LLM prompting or a special AI audio plus text model?
ldenoue··on Veritasium: How Will AI Change Education? [video]
Readable video https://www.appblit.com/scribe?v=0xS68sl2D70
ldenoue··on TL;DW: Too Long; Didn't Watch Distill YouTube Videos to the Relevant Information
Here’s one video above 1 hour and it works with Scribe https://www.appblit.com/scribe?v=FQUo2r-ow-k
ldenoue··on Show HN: TubePen – My attempt to get more out of YouTube learning
looks great. I made a similar app called Scribe where you can highlight passages of the transcript. It's working on the web but also as an iOS app. https://www.appblit.com/scribe

To solve the server IP sometimes being blocked by YouTube, the app fetches the transcripts in the browser.

ldenoue··on Show HN: Free minutes for LLM-powered YouTube transcripts
Same method as my open source lib https://github.com/ldenoue/readabletranscripts but several folks asked for a hosted version so here you go.

200 free minutes on signup so you can try for free.

LLM corrected transcripts are really good, and you can highlight text which I find super useful to study and share quotes.

Page 1 of 6Next →