93 karma · joined November 1, 2021
Context7 is a good MCP. If one were to point an agent to a docs website, the amount of tokens consumed would use up too much of the context window to be able to do a meaningfully complex task with it.
Figma MCP translates Figma to a language an agent understands.
Not everything can be a cli.
Apply Here:
- https://job-boards.greenhouse.io/carta/jobs/7544237003
- https://job-boards.greenhouse.io/carta/jobs/7504452003
⎯
Carta connects founders, investors, and limited partners through world-class software, purpose-built for everyone in venture capital, private equity and private credit.
Frontend Platform: You will own the core infrastructure and monorepo that underpins our entire frontend. Our mantra: "Create a development experience that people will miss when they leave."
Design Systems (Ink): You will build the foundational React components, tooling, and accessibility standards used by hundreds of Carta engineers to ship consistent UI. Check out our docs site at https://ink.carta.com
⎯
Tech: React, TypeScript, Rspack/Vite/Rush, Node.js.
We are looking for engineers who care deeply about developer experience, accessibility, and enabling great design at scale.
Setup:
Terminal:
- Ghostty + Starship for modern terminal experience
- Homebrew to install system packages
IDE:
- Zed (can connect to local models via LM-Studio server)
- also experimenting with warp.dev
LLMs:
- LM-studio as open-source model playground
- GPT-OSS 20B
- QWEN3-Coder-30B-AEB-quantized-4bit
- Gemma3-12B
Other utilities:
- Rectangle.app (window tile manager)
- Wispr.flow - create voice notes
- Obsidian - track markdown notes
the jazz metaphors do not help provide additional context.
You haven’t shared any architectural details. What model? What size? How can anyone be sure that what you’re building is truly offline?
It can process a set of 3-hour audio files in ~20 mins.
I recorded a demo video of how it works here: https://www.youtube.com/watch?v=v0KZGyJARts&t=300s
[1] https://github.com/naveedn/audio-transcriber
I alluded to building this tool on a previous HN thread: https://news.ycombinator.com/item?id=45338694
- https://huggingface.co/pyannote/speaker-diarization-3.1 - https://github.com/narcotic-sh/senko
I personally love senko since it can run in seconds, whereas py-annote took hours, but there is a 10% WER (word error rate) that is tough to get around.
For preprocessing, I found it best to convert files to a 16kHz WAV format for optimal processing. I also add low-pass and high-pass filters to remove non-speech sounds. To avoid hallucinations, I run Silero VAD on the entire audio file to find timestamps where there's a speaker. A side note on this: Silero requires careful tuning to prevent audio segments from being chopped up and clipped. I also use a post-processing step to merge adjacent VAD chunks, which helps ensure cohesive Whisper recordings.
For the Whisper task, I run Whisper in small audio chunks that correspond to the VAD timestamps. Otherwise, it will hallucinate during silences and regurgitate the passed-in prompt. If you're on a Mac, use the whisper-mlx models from Hugging Face to speed up transcription. I ran a performance benchmark, and it made a 22x difference to use a model designed for the Apple Neural Engine.
For post-processing, I've found that running the generated SRT files through ChatGPT to identify and remove hallucination chunks has a better yield.
You're probably right in that most of those employees are itching to leave already, but financially it would make more sense just quit than take a 2 month sabbatical and then lose a larger payout.
https://www.naveed.dev/posts/senior-engineer-interviews-brok...
https://www.naveed.dev/posts/leetcode-alternatives-compared
https://www.naveed.dev/posts/alternative-senior-engineer-int...