HNHacker News
TopNewBestAskShowJobs

yujonglee

355 karma · joined May 14, 2022

prev: fastrepl.com (YC S25)
submissionscomments
yujonglee··on Litelm: LiteLLM Without the Bloat
- We recently migrated to shadcn, which also comes with dark mode support. - You can subscribe to the issue or email me at yujong at berri.ai

I’m happy to provide any support if you’re willing to try out the initial version.

yujonglee··on Litelm: LiteLLM Without the Bloat
thanks for the feedback. genuinely curious what you think we could be doing better, especially around our dev practices. Would love to hear specifics.
yujonglee··on Litelm: LiteLLM Without the Bloat
also if you had resource issues with your production deployment, keen to hear more details
yujonglee··on Litelm: LiteLLM Without the Bloat
hey - litellm will migrate its core to Rust soon. please follow this issue if you're interested! https://github.com/BerriAI/litellm/issues/31263
yujonglee··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
cool! we have something similar implemented: https://github.com/fastrepl/anarlog/tree/main/crates/owhispe...
yujonglee··on Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone
amazing work! impressed with https://deepgrove.ai/chat
yujonglee··on The Conductor Rewrite: What They Changed to Make It Fast
there is some overlap on how I built relay plugin: https://yujonglee.com/blog/hacking-tauri-for-designer/

Source code: https://github.com/fastrepl/anarlog/tree/d38413681a3c02f5ba3...

yujonglee··on Show HN: Compile-time model-id validation with declared capability
fyi this use vendored openrouter models info:

https://github.com/yujonglee/openrouter-toolkit/blob/main/cr...

yujonglee··on Shipping 100 hardware units in under eight weeks
Super impressed and inspired!
yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
kind of self-plug, but you might find https://github.com/fastrepl/hyprnote/blob/main/README.md interesting.

EDIT: typo

yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
It can run Whisper and Moonshine models locally, while also allowing the use of other API providers. Read the docs - or at least this post.
yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
what do you mean? this use-case is not llm. it is realtime stt.

also fyi - https://docs.hyprnote.com/owhisper/configuration/providers/o...

yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
sure. `owhisper pull --help`
yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
yes. metal is on
yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
Thank you!
yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
Its lot more than that.

- It supports other models like moonshine.

- It also works as proxy for cloud model providers.

- It can expose local models as Deepgram compatible api server

yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
probably end of this month or early next month. not 100% sure.
yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
we store data in R2 and range query sometime glitch... It might work if you retry it
yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
yeah we use whisper.cpp for whisper inference. this is more like a community-focused project, not a commercial product!
yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
got it. fyi if you run `owhisper pull --help`, this info is printed
yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
It's open-source. Happy to review & merge if you can send us PR!

https://github.com/fastrepl/hyprnote/blob/8bc7a5eeae0fe58625...

yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
Are you thinking about the realtime use-case or batch use-case?

For just transcribing file/audio,

`owhisper run <MODEL> --file a.wav` or

`curl httpsL//something.com/audio.wav | owhisper run <MODEL>`

might makes sense.

yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
This is CLI entry point:

https://github.com/fastrepl/hyprnote/blob/8bc7a5eeae0fe58625...

yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
> I'm on linux

I didn't tested on Linux yet, but we have linux build: http://owhisper.hyprnote.com/download/latest/linux-x86_64

> also, it looks like the `owhisper run` command gives it's output as a tui. Is there an option for a plain tex

`owhisper run` is more like way to quickly trying it out. But I think piping is definitely something that should work.

> Same question for streaming, is there a way to get a streaming text output from owhisper?

You can use Deepgram client to talk to `owhisper serve`. (https://docs.hyprnote.com/owhisper/deepgram-compatibility) So best resource might be Deepgram client SDK docs.

> diarisation

yeah on the roadmap

yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
I use VAD to chunk audio.

Whisper and Moonshine both works in a chunk, but for moonshine:

> Moonshine's compute requirements scale with the length of input audio. This means that shorter input audio is processed faster, unlike existing Whisper models that process everything as 30-second chunks. To give you an idea of the benefits: Moonshine processes 10-second audio segments 5x faster than Whisper while maintaining the same (or better!) WER.

Also for kyutai, we can input continuous audio in and get continuous text out.

- https://github.com/moonshine-ai/moonshine - https://docs.hyprnote.com/owhisper/configuration/providers/k...

yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
Correct. About the deepgram-compatibility: https://docs.hyprnote.com/owhisper/deepgram-compatibility

Let me know how it goes!

yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
yeah that is on the roadmap!
yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
It separate mic/speaker as 2 channel. So you can reliably get "what you said" vs "what you heard".

For splitting speaker within channel, we need AI model to do that. It is not implemented yet, but I think we'll be in good shape somewhere in September.

Also we have transcript editor that you can easily split segment, assign speakers.

yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
If your use-case is meeting, https://github.com/fastrepl/hyprnote is for you. OWhisper is more like a headless version of it.
yujonglee··on Show HN: OWhisper – Ollama for realtime speech-to-text
Happy to answer any questions!

These are list of local models it supports:

- whisper-cpp-base-q8

- whisper-cpp-base-q8-en

- whisper-cpp-tiny-q8

- whisper-cpp-tiny-q8-en

- whisper-cpp-small-q8

- whisper-cpp-small-q8-en

- whisper-cpp-large-turbo-q8

- moonshine-onnx-tiny

- moonshine-onnx-tiny-q4

- moonshine-onnx-tiny-q8

- moonshine-onnx-base

- moonshine-onnx-base-q4

- moonshine-onnx-base-q8

Page 1 of 4Next →