HNHacker News
TopNewBestAskShowJobs

karimf

2,384 karma · joined December 24, 2016

fikrikarim.com
submissionscomments
karimf··on Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
This is awesome. Thanks for pushing the audio pareto frontier forward.

Probably far fetched for now, but I think the next big evolution is building the pareto/much cheaper alternative to GPT-Live-1.

The STT/TTS market is quite saturated, while today, there's almost no cheap/open source alternative to GPT-Live-1.

karimf··on DeepSeek v4.1 Flash
While this is very impressive benchmark-wise, GPT-6 Astra showed us that benchmarks don't always correlate 1:1 to intelligence of a model.

When Astra launched, I think Artifical Analysis showed that it was on par with GPT-5.6 Sol and lower than Opus or something like that? Then, they updated the scoring.

I hope that more open source models, including this model, to be "as good to use" as Astra.

karimf··on Show HN: Free Inference Engineer and Model Training Roadmap
I think a curriculum like this is neat and might help with interviews since you go wide and have a checklist of things that you need to learn.

I'm on a totally different path for learning inference engineering. I self-host a voice AI app that has ~2000 monthly active users on my own GPU box.

This forces me to learn about production serving, KV cache, quantization, inference engine, observability and economics, prefill optimization since I'm optimizing for TTFT instead of decode speed, and many more.

It's fun since every optimization you do directly translate to a better user experience or allow you to serve more users using the same hardware.

karimf··on GLM-5.3 Artificial Analysis Benchmarks
Yes. Please seriously try other models. See relevant thread here: https://news.ycombinator.com/item?id=49296740
karimf··on Why does Opus 5 feel worse to work with?
This 100%. I was Anthropic-pilled. I had a $200/mo subscription and I only used Anthropic models. I was frustrated by the verbose output and the writing style. I tried ASD-STE-100, it helped a bit, but it's still too verbose for my taste.

Then I tried GPT 5.6 Sol. It's night and day.

I think Anthropic just RL too hard on coding capabilities and never calibrated or benchmarked the writing styles.

karimf··on llama.cpp
Not sure why it's on the front page now, but I highly recommend using llama.cpp for running AI model locally vs using other inference framework, unless you have a very specific requirement.

ggerganov and the team have done a stellar job maintaining the quality while still being fast to implement new models/improvements.

karimf··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
Yes, and also waiting for the next iteration of Gemma. Muse or Qwen are optimized for coding, while IMO Gemma is still better for non-coding tasks.

https://x.com/osanseviero/status/2086107547535122767

karimf··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
Practically ~20GB with KV cache

> We quantize weights to ~4-bit, bringing the LM under 20 GB. We validated minimal to no degradation on agentic tasks under compression.

https://www.reddit.com/r/LocalLLaMA/comments/1vkgsum/introdu...

karimf··on Turn And Face The Strange
Most people are going under identity crisis right now because of recent LLM advancements. This post is a good example that shows that it's not only happening at the individual level, but also on the company/organization level.

Is it still worth building products or companies that can be one-shotted by AI? Probably not.

One interesting consequence is that this force everyone to be more ambitious and do something bigger that's impossible before.

I hope more people are working on something that can always bring net positive to humanity even if there are hundreds of people working on the same thing, like clean energy.

karimf··on Local, CPU-Friendly, High-Quality TTS (Text-to-Speech) with Kokoro
This repo is a good starting point for comparing TTS models https://github.com/5uck1ess/tts-bench

Kokoro is a really good model, considered it’s released 1.5 years ago. It’s punching above its weight https://5uck1ess.github.io/tts-bench/scores.html

karimf··on 1-Click GitHub Token Stealing via a VSCode Bug
I've been using Zed for a few weeks now and these two are also my main complaints as well.
karimf··on Cloudflare Email Service
Oh yeah for sure. At that point, using SES is probably a better option compared to running a VPS just for SMTP. I posted that to let them know that SMTP support is a requirement for some developers.
karimf··on Cloudflare Email Service
Ok I just tried the service since I want to migrate from Resend.

Seems like you can only send email via the worker or REST API for now?

Can I send via SMTP? I'm using Supabase and it needs the SMTP credentials.

I can't find anything on the dashboard or on the docs, even though last year they said it supports SMTP [0]

[0] https://blog.cloudflare.com/email-service/

karimf··on Google Gemma 4 Runs Natively on iPhone with Full Offline AI Inference
Related: Gemma 4 on iPhone (254 comments) - https://news.ycombinator.com/item?id=47652561
karimf··on TanStack Start Now Support React Server Components
This is an interesting approach.

> How does this compare to Next.js App Router?

> Next.js App Router is server-first: your component tree lives on the server by default, and you opt into client interactivity with 'use client'.

> TanStack Start is isomorphic-first: your tree lives wherever makes sense. At the base level, RSC output can be fetched, cached, and rendered where it makes sense instead of owning the whole tree. When you want to go further, Composite Components let the client assemble the final tree instead of just accepting a server-owned one.

The sudden server-first change on Next.js App Router definitely trips some people, especially since React started as client-only library.

karimf··on Show HN: Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B
Well, on my demo it's around 2.5s and I already consider it as a "real-time". One way to improve it is to disable the image input.
karimf··on Show HN: Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B
What's your average response time with M1 max and what's the target?
karimf··on Show HN: Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B
Is this the problem? https://news.ycombinator.com/item?id=47669954
karimf··on Show HN: Ghost Pepper – Local hold-to-talk speech-to-text for macOS
In the /r/macapps subreddit, they have huge influx of new apps posts, and the "whisper dictation" is one of the most saturated category. [0]

>“Compare” - This is the most important part. Apps in the most saturated categories (whisper dictation, clipboard managers, wallpaper apps, etc.) must clearly explain their differentiation from existing solutions.

https://www.reddit.com/r/macapps/comments/1r6d06r/new_post_r...

karimf··on Show HN: Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B
During my limited testing, it works better than I expected at handling multiple languages in a single session. Perhaps I just had a low expectation since I've mostly worked with English-only STT models.
karimf··on Show HN: Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B
Nothing unique, it's just taking a snapshot when it's processing the input. Even processing a single image will increase the TTFT by ~0.5s on my machine, so for now, it seems to be impossible for feeding a live video and expecting a real-time response.

In regards to the video capability, I haven't tested it myself, but here's a benchmark/comparison from Google [0]

[0] https://huggingface.co/blog/gemma4#video-understanding

karimf··on Show HN: Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B
Huh that's weird. I just tried it and it works on my machine. Could you perhaps create a GitHub issue and share the reproduction steps and any relevant logs?
karimf··on I won't download your app. The web version is a-ok
This. I posted this on my other comment, but there's a meme that "Gen Z Kids Don't Understand How File Systems Work" [0].

There seems to be a disconnect between some developers and the younger folks.

[0] https://news.ycombinator.com/item?id=30253526

karimf··on I won't download your app. The web version is a-ok
This is my stance as well, but keep in mind that a lot of people have the opposite preference.

They didn't grow up with the world wide web. They only started using technology when Android and iPhone was popular. They only know Whatsapp, Youtube, TikTok. They're not used to using the browser.

There's a meme that "Gen Z Kids Don't Understand How File Systems Work" [0]

So, it'll depend on your target audiences.

[0] https://news.ycombinator.com/item?id=30253526

karimf··on Gemma 4 on iPhone
Oh wow, that's awesome. Thanks a lot, dang!
karimf··on Show HN: Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B
Thank you. This reminds me of a paragraph from the LatentSpace newsletter [0]

> The excellent on device capabilities makes one wonder if these are the basis for the models that will be deployed in New Siri under the deal with Apple….

https://www.latent.space/p/ainews-gemma-4-the-best-small-mul...

karimf··on Gemma 4 on iPhone
Thanks for sharing! I'm still torn about it. Sure it'll feel more natural if you have the AI head animation, but I don't want people to get attached to it. I don't want to make the loneliness epidemic even worse.
karimf··on Gemma 4 on iPhone
Thanks! Although, I can't claim any credit for it. I just spent a day gluing what other people have built. Huge props to the Gemma team for building an amazing model and also an inference engine that's focused for edge devices [0]

[0] https://github.com/google-ai-edge/LiteRT-LM

karimf··on Gemma 4 on iPhone
This app is cool and it showcases some use cases, but it still undersells what the E2B model can do.

I just made a real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B. I posted it on /r/LocalLLaMA a few hours ago and it's gaining some traction [0]. Here's the repo [1]

I'm running it on a Macbook instead of an iPhone, but based on the benchmark here [2], you should be able to run the same thing on an iPhone 17 Pro.

[0] https://www.reddit.com/r/LocalLLaMA/comments/1sda3r6/realtim...

[1] https://github.com/fikrikarim/parlor

[2] https://huggingface.co/litert-community/gemma-4-E2B-it-liter...

karimf··on Google releases Gemma 4 open models
Update: Just made one that runs on Macbook M3 Pro https://github.com/fikrikarim/parlor
Page 1 of 5Next →