HNHacker News
TopNewBestAskShowJobs

eevmanu

28 karma · joined May 2, 2018

submissionscomments
eevmanu··on Testing is better than data structures and algorithms
deterministic simulation testing[1] (DST)

[1] https://notes.eatonphil.com/2024-08-20-deterministic-simulat...

eevmanu··on Show HN: Python Audio Transcription: Convert Speech to Text Locally
If I understood correctly, VAD has superior results than using ffmpeg silencedetect + silentremove, right?

I think latest version of ffmpeg could use whisper with VAD[1], but I still need to explore how with a simple PoC script

I'd love to know more about the post-processing prompt, my guess is that looks like an improved version of `semantic correction` prompt[2], but I may be wrong ¯\_(ツ)_/¯ .

[1] https://ffmpeg.org/ffmpeg-filters.html#toc-whisper-1

[2] https://gist.github.com/eevmanu/0de2d449144e9cd40a563170b459...

eevmanu··on What fact do you wish everyone understood?
understand digital advertisement incentives and modern marketing tactics and techniques

https://news.ycombinator.com/item?id=40675527

https://news.ycombinator.com/item?id=43187603

maybe this could help people to be more thoughtful on how they invest their attention

eevmanu··on Ask HN: Claude Code–style agent, but Aider-like and model-agnostic?
claude code router is a great alternative

another one could be https://github.com/opencode-ai/opencode

eevmanu··on Coding with LLMs in the summer of 2025 – an update
Open-weight and open-source LLMs are improving as well. While there will likely always be a gap between closed, proprietary models and open models, at the current pace the capabilities of open models could match today’s closed models within months.
eevmanu··on Building Rq: A Fast Parallel File Search Tool for Windows in Modern C
This tool looks great. Congrats for the making of this tool. I would like to know if there is any internal file that could be used to create a benchmark from it in order to compare how fast it gets results in comparison with Void Everything, which is the most common app if a Windows user would like to improve file search.
eevmanu··on Ask HN: What makes you keep coming back to Hacker News?
I get nerd sniped every day.
eevmanu··on Show HN: Spegel, a Terminal Browser That Uses LLMs to Rewrite Webpages
great POC

looks very similar to a chrome extension i use for a similar goal: reader view - https://chromewebstore.google.com/detail/ecabifbgmdmgdllomnf...

eevmanu··on Jepsen: TigerBeetle 0.16.11
Thanks for your answer aphyr and for this amazing analysis
eevmanu··on Jepsen: TigerBeetle 0.16.11
I have a question that I hope is not misinterpreted, as I'm asking purely out of a desire to learn. I am new to distributed systems and fascinated by deterministic simulation testing.

After reading the Jepsen report on TigerBeetle, the related blog post, and briefly reviewing the Antithesis integration code on GitHub workflow, I'm trying to better understand the testing scope.

My core question is: could these bugs detected by the Jepsen test suite have also been found by the Antithesis integration?

This question comes from a few assumptions I made, which may be incorrect:

- I thought TigerBeetle was already comprehensively tested by its internal test suite and the Antithesis product.

- I had the impression that the Antithesis test suite was more robust than Jepsen's, so I was surprised that Jepsen found an issue that Antithesis apparently did not.

I'm wondering if my understanding is flawed. For instance:

1. Was the Antithesis test suite not fully capable of detecting this specific class of bug?

2. Was this particular part of the system not yet covered by the Antithesis tests?

3. Am I fundamentally comparing apples and oranges, misunderstanding the different strengths and goals of the Jepsen and Antithesis testing suites?

I would greatly appreciate any insights that could help me understand this better. I want to be clear that my goal is to educate myself on these topics, not to make incorrect assumptions or assign responsibility.

eevmanu··on O4-mini vs. Claude 3.7 vs. Gemini 2.5 Pro on code generation
great article!

is it possible to mention in the article the cost per model to run that benchmark?

eevmanu··on Ask HN: How to unit test AI responses?
+1 to evals

https://github.com/anthropics/courses/tree/master/prompt_eva...

https://cookbook.openai.com/examples/evaluation/getting_star...

eevmanu··on Write Prompts Like a Pro: Checkout These Prompt Engineering Tools
what do you think about:

- openai meta prompt[1] or

- anthropic prompt generator[2] or

- deepseek model prompt word generation[3]?

you don't see these as user-friendly enough to consider prompt engineering tools? I mean, I feel like they're way more than just guidelines. Seems like they're a step beyond.

[1]: https://platform.openai.com/docs/guides/prompt-generation#me...

[2]: https://docs.anthropic.com/en/docs/build-with-claude/prompt-...

[3]: https://api-docs.deepseek.com/prompt-library - go to 模型提示词生成 card

eevmanu··on Show HN: I'm learning Go and built a scraper. How can I improve it?
as an idea I'd suggest to evaluate your codebase with the multiple llms available now using the free tier (deepseek V3, llama 3.3 versatile on groq, gemini-exp-.. on google ai studio and so on)

on each llm provider you can find prompting best practices or techniques to improve your eval prompt and eval your code correctly

at least the feedback loop would be faster

if you already thought about it or already did that, well ... great

eevmanu··on Researchers Revert Cancer Cells into Normal Cells
https://onlinelibrary.wiley.com/doi/10.1002/advs.202402132

article here

eevmanu··on Ask HN: What nascent tech are you interested in?
artificial life
eevmanu··on Ask HN: How do you find and consume content online?
Whenever I explore content on different platforms (Hacker News via Algolia, X, Discord, etc.), I always check if the source supports search operators. If it does, I use them to filter out noise and discover interesting or relevant content tailored to my preferences. For platforms without built-in search functionality, I rely on Google’s search operators to surface the content I’m looking for, as long as it’s accessible on the web.
eevmanu··on Ask HN: What programming languages are you learning currently or in 2025
erlang? awesome threading model
eevmanu··on Show HN: I launched a super cheap and simple to use OCR tool for macOS
Any decent alternative for Linux or Ubuntu-based OS? Thanks.
eevmanu··on Ask HN: What's an interesting software development niche?
Artificial Life - https://en.wikipedia.org/wiki/Artificial_life
eevmanu··on Ask HN: How to transcribe a couple thousand calls per day?
Consider "Whisper Large V3" on console.groq.com, imo is fast reliable and cheap ($0.03/hour transcribed).
eevmanu··on Show HN: Kaption AI – WhatsApp Web Extension for Audio Transcription
Do you think will take time for Whatsapp team to activate this feature in Whatsapp web? Because Whatsapp android app in beta version (2.24.17.7) already has a transcribe native feature for voice messages, I recently noticed that around one week ago.

edit: btw I don't know which model Whatsapp team is using behind scenes but definitely doesn't looks like is whisper because is notoriously slow in comparison with using whisper inside chatgpt android app or groq web interface.

eevmanu··on Ask HN: How do you manage articles from different sources on your phone?
Telegram's Saved Messages[1].

Previously, I used to send URLs to myself using WhatsApp, but searching through messages on <web.whatsapp.com> is painfully slow compared to Telegram's search capabilities. The speed difference is likely due to Telegram leveraging the benefits of a desktop-native app, though I'm not totally sure why it's so much faster.

Notion supposedly has beefed up its search engine recently, but my last experience with it was about three years ago, and it was frustratingly slow. I've been hesitant to revisit and see if they've actually improved.

[1]: https://telegram.org/blog/albums-saved-messages#saved-messag...

eevmanu··on Anyone knows a free extension to send personalized WhatsApp messages daily 100
just to complement parent:

https://faq.whatsapp.com/5957850900902049

Unauthorized use of automated or bulk messaging on WhatsApp

https://faq.whatsapp.com/508673377786103

Seeing the message "Your phone number is banned from using WhatsApp. Contact support for help."

https://www.whatsapp.com/legal/terms-of-service

Legal And Acceptable Use

(e) involve sending illegal or impermissible communications such as bulk messaging, auto-messaging, auto-dialing

https://www.whatsapp.com/legal/business-terms/

Restrictions

(h) create software or APIs that function substantially the same as our Business Services and offer them for use by third parties in an unauthorized manner on or utilizing the WhatsApp platform or our Business Services.

eevmanu··on Pg_vectorize: Vector search and RAG on Postgres
Might sound like a rookie question, but curious how you'd tackle semantic chunking for a hefty text, like a 100k-word book, especially with phi-2's 2048 token limit [0]. Found some hints about stretching this to 8k tokens [1] but still scratching my head on handling the whole book. And even if we get the 100k words in, how do we smartly chunk the output into manageable 250-350 word bits? Is there a cap on how much the output can handle? From what I've picked up, a neat summary ratio for a large text without missing the good parts is about 10%, which translates to around 7.5K words or over 20 chunks for the output. Appreciate any insights here, and apologies if this comes off as basic.

[0]: https://huggingface.co/microsoft/phi-2

[1]: https://old.reddit.com/r/LocalLLaMA/comments/197kweu/experie...

eevmanu··on Show HN: I made an app to use local AI as daily driver
Sure thing, your RAG approach sounds intriguing, especially since you're sidestepping vector databases. But doesn't the input context length cap affect it? (chatgpt plus at 32K [0] or gpt4 via open ai at 128K [1]) Seems like those cases would be pretty rare though.

[0]: https://openai.com/chatgpt/pricing#:~:text=8K-,32K,-32K

[1]: https://platform.openai.com/docs/models/gpt-4-and-gpt-4-turb...

eevmanu··on We were not accepted into Google Summer of Code. So, we started our own
Just dropping my $0.02 cents on the Jepsen testing for Qdrant's distributed guarantees. While Jepsen is a solid choice, have you considered throwing Antithesis [1] into the mix as well? If it's already on your radar, then no worries, just figured I'd mention it.

[1]: https://antithesis.com/

eevmanu··on Show HN: I create a free website for download YouTube transcript, subtitle
to whom it may be useful

- `.` - precision flag [1]

- `150` - precision amount [2] - bytes to count after casting to binary representation

- `B` - special conversion type [3] - bytes

[1]: https://docs.python.org/3/library/stdtypes.html#printf-style...

[2]: https://docs.python.org/3/library/stdtypes.html#printf-style...

[3]: https://github.com/yt-dlp/yt-dlp?tab=readme-ov-file#output-t...

eevmanu··on How to copy a file between devices? (2022)
If you want to copy files:

- from a linux computer/laptop

- to an android phone and

- not follow rule 2,

I'd like to suggest another alternative: `adb` (just platform tools) which has the `push` [1] command

    push [--sync] [-z ALGORITHM] [-Z] LOCAL... REMOTE
         copy local files/directories to device
         --sync: only push files that are newer on the host than the device
         -n: dry run: push files to device without storing to the filesystem
         -z: enable compression with a specified algorithm (any/none/brotli/lz4/zstd)
         -Z: disable compression
[1]: https://developer.android.com/tools/adb#copyfiles
eevmanu··on Is something bugging you?
Just a heads-up for anyone diving deeper into this thread - I dug into the original tweet and managed to track down the parent tweet right here: [1]. Moreover, there's a snapshot on archive.org [2] capturing the reply along with the quote in question. Interestingly, there's also a snapshot from foundationdb.com [3] that discusses the outcomes of running Jepsen tests on FDB. Worth checking out for those interested in the technical nitty-gritty.

[1]: https://twitter.com/obfuscurity/status/405016890306985984

[2]: https://web.archive.org/web/20220805112242/https://twitter.c...

[3]: https://web.archive.org/web/20150325003526/http://blog.found...

← PreviousPage 2 of 3Next →