HNHacker News
TopNewBestAskShowJobs

mrkn1

262 karma · joined July 30, 2020

@kouhxp on twitter and bluesky, follow for more
submissionscomments
mrkn1··on What even is an OS now?
While reading your post, I kept thinking of K&R, and the beards of the 70s. Isn't this a return to early Unix? Users writing their own small tools in C, the system shipping building blocks... walls between users rather than programs?
mrkn1··on Yes, Claude can do Nine Loops
I'm sure there's a simple answer to this, and I'm probably just missing it, but what happened with the supergravity one?

And how many runs did it take before this one? They say "in one shot" with nothing more than "keep going." But we only see the successful run, reported by the people who ran it.

mrkn1··on Anyone using a Wispr Flow alternative that is non-cloud?
https://github.com/kouhxp/yapsnap free cpu first and open source
mrkn1··on What Do We Know About the Microplastics Inside Us?
fwiw kimchi-derived probiotic bacterium (Leuconostoc mesenteroides CBA3656) was shown to bind nanoplastics and help mice excrete more of them. But it’s not yet proven that eating kimchi removes microplastics from humans
mrkn1··on Wayfinder Router: deterministic routing of queries between local and hosted LLM
Has anyone tried the others listed? Any feedback?
mrkn1··on Mistral OCR 4
This runs for free on CPU https://github.com/kouhxp/textsnap
mrkn1··on Show HN: Trace – Offline Mac meeting transcripts you can flag mid-call
The key moments feat is neat. Been working on a free opensource offline transcriber that runs fast on CPU and does diarization too

https://github.com/kouhxp/yapsnap

mrkn1··on Show HN: Local-first fast CPU image to text for screenshots, PDFs, webpages
No rigorous eval, and I love Tesseract. Here's the example that motivated me to build textsnap (which is in the github's README), parsed with Tesseract:

https://imgur.com/a/i2eQra8

mrkn1··on Show HN: Local-first fast CPU image to text for screenshots, PDFs, webpages
No reason other than their Q4 model working reasonably well and fast on my CPU laptop. Should work with any ONNX VLM model
mrkn1··on Show HN: Local-first fast CPU image to text for screenshots, PDFs, webpages
109 languages, including other alphabets.
mrkn1··on Show HN: Local-first fast CPU image to text for screenshots, PDFs, webpages
Just ran

  textsnap "https://i.ytimg.com/vi/LBNDfxjEYlA/maxresdefault.jpg"
and got this

  $('.count').each(function () {
  $('this').prop('Counter', 0).animate({
    Counter: $('this').text()
  }, {
      duration: 4000,
      easing: 'swing',
      step: 'function (now) {
          $('this").text(Math.ceil(now));
      }
    }); 
  });
mrkn1··on Show HN: Local-first fast CPU image to text for screenshots, PDFs, webpages
thank you!
mrkn1··on Show HN: Local-first fast CPU image to text for screenshots, PDFs, webpages
thank you! what is it about?
mrkn1··on Show HN: CPU-only fast OCR for screenshots, images, PDFs, webpages
I have not. My experience has been a few seconds for an 1024x1024 with medium density of text, FWIW. Feel free to try it on a few test images, model is pretty small and fast, but yeah no formal evals on CPU.
mrkn1··on Show HN: CPU-only fast OCR for screenshots, images, PDFs, webpages
It should support 109 languages. More info here: https://huggingface.co/PaddlePaddle/PaddleOCR-VL
mrkn1··on Show HN: CPU-only fast OCR for screenshots, images, PDFs, webpages
No, simply, my laptop only has a CPU.
mrkn1··on Show HN: CPU-only fast OCR for screenshots, images, PDFs, webpages
I haven't. But the evals of the underlying model are published here, including on Omnibench. https://huggingface.co/PaddlePaddle/PaddleOCR-VL-1.5
mrkn1··on Show HN: Claude Code's $200 plan is a 17× subsidy on the raw API
https://github.com/kouhxp/fftext
mrkn1··on Disagreement among frontier LLMs on real-world fact-checks
For 100% local CPU fact checking, I made this: https://news.ycombinator.com/item?id=48301003
mrkn1··on Show HN: Claude Code's $200 plan is a 17× subsidy on the raw API
A lot of my queries are summarize/explain/fact check, and these are covered 100% on my CPU locally [0], reducing frontier model reliance

[0] https://news.ycombinator.com/item?id=48301003

mrkn1··on Show HN: CPU-only fast OCR for screenshots, images, PDFs, webpages
just released new version that implements your ideas
mrkn1··on Show HN: CPU-only transcription for YouTube, TikTok, X, Instagram videos
new version released is up to 3x faster on CPU. Let me know!
mrkn1··on Show HN: CPU-only fast OCR for screenshots, images, PDFs, webpages
thank you being thorough

clipboard: rn input is treated like any other source, so text gets written to ./textsnaps/clipboard_ocr.txt, and stdout just prints that path. Nothing goes back to the clipboard in this version (stay tuned)

portability: agreed, and it's a small change. textsnap already looks for the checksum manifest next to the script before falling back to the cache, so extending it should be easy. I make a note for next version.

mrkn1··on Show HN: CPU-only fast OCR for screenshots, images, PDFs, webpages
Great question. I'm not familiar with docling-serv but pretty different beasts from what I gathered. Docling is a heavier pipeline (actually uses GPU).textsnap is the opposite: single-file CLI, small VLM running on plain CPU cores, one command, no server. Tradeoff is CPU decode is sequential so it's slower on dense pages, and it OCRs one image rather than doing full layout.

If docling-serve is already meeting your needs it's probably not an upgrade. But it installs in one command, so would love to hear how it stacks up on your images, if you end up trying it.

mrkn1··on Show HN: CPU-only fast OCR for screenshots, images, PDFs, webpages
thanks! yapsnap is audio to text, and textsnap is image to text. Both have been daily use cases for me for a while. And yes, the feedback on yapsnap encouraged me to also release textsnap on github
mrkn1··on Show HN: CPU-only transcription for YouTube, TikTok, X, Instagram videos
Reading this means a lot, thank you! Even faster version in the coming days, stay tuned and PRs welcome!
mrkn1··on Show HN: CPU-only transcription for YouTube, TikTok, X, Instagram videos
Added Diarization / Speaker Separation that is fast and CPU only. Thank you all for the great feedback and support. PRs welcome!

  yapsnap "https://www.youtube.com/watch?v=NzKJ-xO-VhE" --diarize

  SPEAKER_00 [00:00]: Welcome to the show.
  SPEAKER_01 [00:03]: Glad to be here, thanks for having me.
  SPEAKER_00 [00:08]: Let's get started.
mrkn1··on Show HN: CPU-only transcription for YouTube, TikTok, X, Instagram videos
done! new version separates speakers on CPU fast
mrkn1··on Show HN: CPU-only transcription for YouTube, TikTok, X, Instagram videos
done! just pushed a new version with CPU diarization
mrkn1··on Show HN: CPU-only transcription for YouTube, TikTok, X, Instagram videos
Kroko's website says benchmarks aren't formalized yet. FWIW, this url says 5% WER for English [0]. though it doesn't specify the dataset, so not directly comparable to Parakeet's 6.32 on the Open ASR Leaderboard

Best way to judge is to try it on your own audio

[0] https://huggingface.co/hudaiapa88/sherpa-stt-onnx

Page 1 of 3Next →