HNHacker News
TopNewBestAskShowJobs

primaprashant

544 karma · joined June 8, 2021

Writing Agentic Coding Weekly (https://www.agenticcodingweekly.com/) | ML Engineer
submissionscomments
primaprashant··on Ask HN: What are you working on? (September 2026)
Haven't added a demo to the repo yet (will do soon) but building this TUI [1] in Go to manage my agent skills.

I don't like putting 20-30 agent skills in the .claude/skills/ dir and letting the agent figure it out. I keep all the skills I've created and adapted over time in a separate git repo and then from there I copy them to the project skill dir when i need them for a session and then remove them when I'm done.

This TUI just replaces all the manual work with ls, cp -r, and rm -rf commands with a few keystrokes.

[1] https://github.com/primaprashant/sei

primaprashant··on Show HN: Yap – OSS on-device voice dictation for macOS with no model to download
Pretty cool. Added to this awesome-style GitHub repo I maintain containing all the best open-source voice typing tools:

https://github.com/primaprashant/awesome-voice-typing

primaprashant··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
A couple tidbits:

> Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready.

> We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.

primaprashant··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Pricing per million input/output tokens:

2.5 Flash: $0.3 / $2.5

3.0 Flash: $0.5 / $3

3.5 Flash: $1.5 / $9

3.6 Flash: $1.5 / $7.5

---

2.5 Flash-Lite: $0.1 / $0.4

3.1 Flash-Lite: $0.25 / $1.5

3.5 Flash-Lite: $0.3 / $2.5

primaprashant··on Show HN: Hns – Local open source speech-to-text CLI
Hi all, demo is at the bottom of the readme. Or you can watch here:

https://github.com/user-attachments/assets/2aa3752e-bd16-453...

primaprashant··on Transcribe.cpp
If you're talking about translated text, then that should be super easy. Most of these dictation tool support post-processing with LLM to remove filler words, fix punctuation, etc. I'd imagine you can change the system prompt for the post-processing step to do the translation instead, and you'd get translated text.
primaprashant··on Transcribe.cpp
Handy is an amazing cross-platform app for dictation from the author. There are other awesome open-source dictation tools as well like native macOS ones. You do not need SaaS subscription in this day and age for transcription.

I maintain this list of all the best open-source ones in this awesome-style GitHub repo. People looking for open-source dictation tools, hope you find something that works for you here:

https://github.com/primaprashant/awesome-voice-typing

primaprashant··on Transcribe.cpp
Totally understandable, but I’ve found that software that transcribes everything after I finish recording actually works better for me. I’ve tried both kinds, and systems that continuously type what I’m saying distract me from completing my thought. I end up reading what’s being typed and noticing transcription mistakes instead of focusing on what I’m trying to say.

I often prefer to dictate everything in my head about a particular thing for 5–10 minutes and then go through it afterward. I find that much more useful because it doesn’t break my thought process the way continuous transcription does.

primaprashant··on Ask HN: What Are You Working On? (July 2026)
Continuing my newsletter about agentic coding:

https://www.agenticcodingweekly.com/

primaprashant··on Text art tools
I built this [1] for myself so that I can comment lgtm in PR comments with an ASCII art. Pretty silly but fun.

[1]: LGTM ASCII Art as a Service - https://lgtms.app/

primaprashant··on How to ask for help from people who don't know you
sending nohello website [1] back to people who just say hello without additional context is a classic solution

[1] https://nohello.net/en/

primaprashant··on Claude Sonnet 5
Based on both performance vs price charts, it seems using Opus 4.8 with med effort is almost a better choice than using Sonnet 5 at xhigh effort
primaprashant··on Apple raises prices of MacBooks, iPads
These are the price changes mentioned in the article:

Macs

  MacBook Neo: $699 (up from $599)
  13-inch MacBook Air: $1,299 (up from $1,099)
  15-inch MacBook Air: $1,499 (up from $1,299)
  M5 MacBook Pro: $1,999 (up from $1,699)
  M5 Pro MacBook Pro: $2,499 (up from $2,199)
  M5 Max MacBook Pro: $4,099 (up from $3,599)
  iMac: $1,499 (up from $1,299)
  M4 Max Mac Studio: $2,499 (up from $1,999)
  M3 Ultra Mac Studio: $5,299 (up from $3,999)
iPads

  iPad: $449 (up from $349)
  11-inch iPad Air: $749 (up from $599)
  13-inch iPad Air: $949 (up from $749)
  11-inch iPad Pro: $1,199 (up from $999)
  13-inch iPad Pro: $1,499 (up from $1,299)
  iPad mini: $599 (up from $499)
More products:

  Apple TV 4K: $199 (up from $129)
  HomePod: $349 (up from $299)
  HomePod mini: $129 (up from $99)
  Vision Pro: $3,699 (up from $3,499)
primaprashant··on Apple rejected my dictation app for using the accessibility API
I am big fan of VoiceInk which is also local and open-source. I also maintain this list of all the best open-source ones in this awesome-style GitHub repo. People looking for open-source dictation tools, hope you find something that works for you here: https://github.com/primaprashant/awesome-voice-typing
primaprashant··on Gemini CLI will stop working from June 18, 2026
Well, there is no IDE in antigravity 2.0
primaprashant··on Apple unveils new accessibility features
Open-source STT apps are plenty and just as good. Pick one from this list:

https://github.com/primaprashant/awesome-voice-typing

primaprashant··on Apple unveils new accessibility features
I’d say STT is pretty much a solved problem. Everyday there is a new product and can be one-shotted by any current top of the line LLMs. Take a look at this [1]. Apple is just stuck in the past.

https://github.com/primaprashant/awesome-voice-typing

primaprashant··on Google I/O
Looks like it. Found this:

> On June 18, 2026, Gemini CLI and Gemini Code Assist IDE extensions will stop serving requests for Google AI Pro and Ultra, as well as those using it free of charge using Gemini Code Assist for individuals.

https://developers.googleblog.com/an-important-update-transi...

primaprashant··on Google I/O
So Spark is cloud hosted openclaw?
primaprashant··on DeepSeek v4
While SWE-bench Verified is not a perfect benchmark for coding, AFAIK, this is the first open-weights model that has crossed the threshold of 80% score on this by scoring 80.6%.

Back in Nov 2025, Opus 4.5 (80.9%) was the first proprietary model to do so.

primaprashant··on Ask HN: What Are You Working On? (April 2026)
I've been using speech-to-text tools every day now especially for dictating detailed prompts to LLMs and coding agents. I personally use VoiceInk which is open-source.

I tried to look for what other solutions are available and I've collected all the best open-source ones in this awesome-style GitHub repo. Hope you find something that works for you!

https://github.com/primaprashant/awesome-voice-typing

primaprashant··on Native Instant Space Switching on macOS
I love using (tiling) window managers, and one of the most important requirements for me is having a key binding for switching to the last active workspace. The proposed solution in the blog doesn't achieve this. I use Aerospace on macOS right now and think it's the best solution available.

I generally have fixed workspaces for different things: first for a browser, second for a code editor, third for a terminal, and so on. If I want to switch between the browser and code editor, I can do that with a single key binding, usually Alt+Tab. The same binding lets me switch between the code editor and terminal just as easily.

When you have something like 10 different workspaces, not having this key binding becomes annoying. If you need to alternate between windows on workspace one and workspace eight, you're stuck using both hands to press Control+1 and then Control+8. But with a last-active-workspace key binding, you can just Alt+Tab between them. This is the killer feature I always need.

primaprashant··on Show HN: Ghost Pepper – Local hold-to-talk speech-to-text for macOS
For me personally, it's not really about typing speed. While I can type pretty fast and most likely I speak faster than typing, but typing and dictating are just different way of doing things for me. While the end result of both is same, but for me it's just like different way of doing things and it's not a competition between the two.

I regularly just sit down and often just describe whatever I'm trying to do in detail and I speak out loud my entire thought process and what kind of trade-offs I'm thinking, all the concerns and any other edge cases and patterns I have in my mind. I just prefer to speak out loud all of those. I regularly speak out loud for 5 to 10 minutes while sometimes taking some breaks in between as well to think through things.

I am not doing it just for vibe coding, I'm using it for everything. So obviously for driving coding agents, but also for in general, describing my thoughts for brainstorming or having some kind of like a critique session with LLMs for my ideas and thoughts. So for everything, I'm just using dictation.

One other benefit I think for me personally is that since I'm interacting with coding agents and in general LLMs a lot again and again every day, I end up giving much more context and details if I'm speaking out loud compared to typing. Sometimes I might feel a little bit lazy to type one or two extra sentences. But while speaking, I don't really have that kind of friction.

primaprashant··on Show HN: Ghost Pepper – Local hold-to-talk speech-to-text for macOS
Speech-to-text has become integral part of my dev flow especially for dictating detailed prompts to LLMs and coding agents.

I have collected the best open-source voice typing tools categorized by platform in this awesome-style GitHub repo. Hope you all find this useful!

https://github.com/primaprashant/awesome-voice-typing

primaprashant··on Stop Typing Prompts to Your Coding Agent
Author here. My argument is: we give instructions to coding agents dozens of times a day. Over time, speaking those instructions naturally tends to produce more detailed context than typing them out, because the friction of typing makes you abbreviate.

I've been using VoiceInk on macOS for a few months now. The workflow is just: hold shortcut, speak, release, text appears at cursor and works in terminal, editor, chat, wherever.

The post covers Handy, Whispering, VoiceInk, OpenWhispr, and FluidVoice. All open-source, all do local transcription, all paste directly into the active window. The differences are mostly platform support, model selection, and how much extra stuff (AI post-processing, voice-activated mode, etc.) they add.

Happy to answer questions about any of these or about the voice-typing-for-agents workflow in general.

primaprashant··on Prompt-caching – auto-injects Anthropic cache breakpoints (90% token savings)
I've found RTK CLI proxy [1] quite useful for reducing token usage

[1]: https://github.com/rtk-ai/rtk/

primaprashant··on Voice Typing – Curated list of open-source speech-to-text tools
I encourage everyone to use speech-to-text tools to give detailed context to coding agents. As a developer, I love my keyboard and I can understand if you're skeptical. I was too. But using speech-to-text is one of the high-leverage things you can do as a developer.

We all know LLMs work better when given more context and clear instructions. When you're working with coding agents, you're giving instructions to them multiple times a day, every day. Over time, you end up giving much better instructions and detailed context if you use speech-to-text compared to manually typing all those instructions all the time.

There are tons of open source and proprietary products for speech-to-text which offer inference on local machine or on cloud. So I put together a curated list of 30+ open-source tools across Linux, macOS, Windows, Android, and iOS. Most support offline recognition. Pick whatever you find suitable, but I definitely recommend giving speech-to-text a try for your LLM workflows. And if you're skeptical, give it a week and then re-evaluate.

https://github.com/primaprashant/awesome-voice-typing

primaprashant··on Show HN: What was the world listening to? Music charts, 20 countries (1940–2025)
I picked India and a random year, 1985 [1]. The number 3 song caught my eye cause it had the thumbnail of a famous movie that came out in 2004, although the correct song played. When I went to the linked Spotify playlist for that year, the included song at number 3 was wrong and linked to the song from the 2004 movie.

Not sure what the data source is, but needs a little bit of cleaning and validation. Not critiquing, this project is awesome, just giving a heads up.

[1]: https://88mph.fm/in/1985

primaprashant··on Ask HN: What Are You Working On? (March 2026)
Continuing my weekly newsletter about agentic coding updates:

https://www.agenticcodingweekly.com/

primaprashant··on Show HN: A Self-Paced Exercise to Build a CLI Coding Agent from Scratch
I ran this live in Tokyo with ~50 engineers. The biggest "aha" moment was Phase 4, when the agent loop closes and the LLM starts chaining tool calls autonomously. People go from "I'm building a chatbot" to "oh, this is an agent".

Also, there have been plenty of "build a coding agent in 200 lines" posts on HN in the past year, and they're great for seeing the final picture. I created this simple structured exercise so we start from an empty loop and build each piece ourself, phase by phase. So instead of just reading the implementation, I hope more people try the implementation themselves.

These are the 7 phases of implementation:

  1. LLM in the loop: replace the canned response with an actual LLM call
  2. Read file tool: implement the tool + pass its schema to the LLM + detect tool use in the response
  3. Tool execution: execute the tool the LLM requested and display the result
  4. Agent loop: the inner loop where tool results go back to the LLM until no more tool calls
  5. Edit file tool: create and edit files
  6. Bash tool: execute shell commands with user confirmation
  7. Memory: use the agent to build the agent, add AGENTS.md support for persistent memory across sessions

Feedback and PRs welcome. Happy to answer any questions.
Page 1 of 3Next →