HNHacker News
TopNewBestAskShowJobs

underlines

786 karma · joined August 22, 2012

submissionscomments
underlines··on Experimental web browser optimized for rabbit-holing
I specifically just switched back from chrome to Firefox after about 10 years, because there were still no native hierarchical tab solutions in chrome.

FF + sideberry for every day use and rabbit holing.

underlines··on Show HN: A real time AI video agent with under 1 second of latency
That's a cool tech demo, I really like it. I thought about something similar with only open sourced components:

1. Audio Generation: styletts2 xttsv2 or similar for and fine tuning 5min of audio for voice cloning

2. Voice Recognition: Voice Activity Detection with Silero-VAD + Speech to Text with Faster-Whisper, to let users interrupt

3. Talking head animation: some flavor of wav2lip, diff2lip or LivePortrait

4. Text inference: Any grok hosted model that is fast enough to do near real time responses (llama3.1 70b or even 8b) or local inference of a quantized SML like a 3B model on a 4090 via vLLM

5. Visual understanding of users webcam: either gpt-4o with vision (expensive) or a cheap and fast Vision Language Model like Phi3-vision, LLaVA-NeXT, etc. on a second 4090

6. Prompt:

You are in a video conference with a user. You will get the user's message tagged with #Message: <message> and the user's webcam scene described within #Scene: <scene>. Only reply to what is described in <scene> when the user asks what you see. Reply casual and natural. Your name is xxx, employed at yyy, currently in zzz, I'm wearing ... Never state pricing, respond in another language etc...

underlines··on Brazilian Electric "Suicide" Shower Heads [video]
Why do people only discuss flow heating vs gas heating? In Switzerland I never saw either of them and only know warm water boilers my whole life. I freaked out seeing electricity going into the bathroom in Thailand, where ungrounded homes are common. Are boilers that uncommon in the rest of the world?
underlines··on Show HN: Put ful.co/ in front of any URL to easily copy its SVGs and images
It 404ed with the two biggest swiss websites: digitec.ch and 20min.ch
underlines··on Contextual Retrieval
We build a corporate RAG for a government entity. What I've learned so far by applying an experimental A/B testing approach to RAG using RAGAS metrics:

- Hybrid Retrieval (semantic + vector) and then LLM based Reranking made no significant change using synthetic eva-questions

- HyDE decreased answer quality and retrieval quality severly when measured with RAGAS using synthetic eval-questions

(we still have to do a RAGAS eval using expert and real user questions)

So yes, hybrid retrieval is always good - that's no news to anyone building production ready or enterprise RAG solutions. But one method doesn't always win. We found semantic search of Azure AI Search being sufficient as a second method, next to vector similarity. Others might find BM25 great, or a fine tuned query post processing SLM. Depends on the use case. Test, test, test.

Next things we're going to try:

- RAPTOR

- SelfRAG

- Agentic RAG

- Query Refinement (expansion and sub-queries)

- GraphRAG

Learning so far:

- Always use a baseline and an experiment to try to refute your null hypothesis using measures like RAGAS or others.

- Use three types of evaluation questions/answers: 1. Expert written q&a, 2. Real user questions (from logs), 3. Synthetic q&a generated from your source documents

underlines··on Objective Bayesian Hypothesis Testing
A great question that I came across in Hypothesis Driven Development a long time ago: Should you use Frequentist Statistics or Bayesian Statistics? It's relevant when you do A/B or Multivariate Testing.

As it was very difficult for someone like me without higher stats or math education, I can highly recommend the following additional sources:

- https://www.redjournal.org/article/S0360-3016(21)03256-9/ful...

- https://amplitude.com/blog/frequentist-vs-bayesian-statistic...

- https://indico.cern.ch/event/568904/contributions/2651065/at...

underlines··on Postgres.new: In-browser Postgres with an AI interface
Benchmarks suggest otherwise. Toqan's sql benchmark shows other models way up in the ranking compared to gpt-4o [1]

Open Weight models specifically fine-tuned on sql generation and modification also rank pretty well compared to SOTA proprietary models. If you want to eval alternative models, check out sqleval [2]

1 https://prollm.toqan.ai/leaderboard/stack-unseen?type=concep...

2 https://github.com/defog-ai/sql-eval

underlines··on GitHub Copilot – Lessons
I think what many people miss, is how long context LLMs get better (needle in a needle stack) and how Context is very important. With github copilot or with continue for VS Code the main issue lies in how they decide which context to give: Ideally it's a graph of all function calls and instantiations of classes so whenever your cursor is in a particular spot you could hop through the graph and get all important code pieces as context.

Currently continue uses vector search of chunks which is just a crutch. I am not sure what copilot or Aider does, but the right context is key.

Another way to improve a coding assistant is to move away from simple paradigms to agent based Workflow that can deploy sub agents and break down tasks. All in the back autonomously while you code and then it surfaces suggestions or changes. This will get possible with increasing inference speeds (groq) and better agent frameworks.

Most devs at our company say LLMs are useless for coding and github copilot is a glorified autocomolete costing 20USD. I think the tech and ecosystem will improve a lot over time and they underestimate LLM abilities due to their bias.

underlines··on Show HN: Local voice assistant using Ollama, transformers and Coqui TTS toolkit
Honestly, there are so many Project on Github doing STT - LLM - TTS that I lost count. The only revolutionary thing that feels like magic is if the STT supports Voice Activity Detection and low latency LLM inference on Groq, so conversations feel natural.
underlines··on Regular expression functions in Excel
Coincidentally using the same function names as Google sheets did for years. :D
underlines··on Falling in love again with disposable cameras
Just get a refurbished film camera like the excellent Canon AE-1 Program and enjoy film photography.

The chemicals used for developing color film are bad enough already, no need to add pollution by single use plastic to the problem again.

underlines··on Gifski: Optimized GIF Encoder
I use it for Documentation (when interactive), for Teams with integrated short clips, for Powerpoint Presentations when I'm doing a live demo (having a second slide with the gif as a backup, when the live demo doesn't work).

The editor is minimal and has a non existent learning curve.

AFAIK screentogif even has the gifsky encoder plus a few others.

underlines··on NoTunes is a macOS application that will prevent Apple Music from launching
Ditto. And my Spotify library shows cracks more and more. I am a playlist curator by heart, did that with my 100GB MP3/FLAC Music collection for 20 years before switching to Spotify, but now and then you get an empty artist/track name entry and can't even see what song it was, due to the Song rights being lost.
underlines··on NoTunes is a macOS application that will prevent Apple Music from launching
Making iMessage work on Android is a non issue for most parts of the world: mobile messaging is fragmented and the US uses SMS/iMessage, while most of the rest of the world uses either whatsapp, LINE or WeChat.
underlines··on URLhaus: A database of malicious URLs used for malware distribution
abuse.ch is a non-profit, initially private. Working on cyber security issues for 15 years. Mainly focused on botnets and malware. Since 2021, abuse.ch is under the Institute for Cybersecurity and Engineering ICE at Bern University of Applied Sciences. To date the project has been funded entirely from private-sector donations.

They have mainly two goals:

1 Research: Research into malware and botnets

2 Open source threat intelligence: indicator of compromise – IOC for the public to prevent threats

underlines··on GPT-4o's Memory Breakthrough – Needle in a Needlestack
Such tasks don't need a large context window. Just good RAG.
underlines··on Mixtral 8x22B
tons and tons of papers, most of them had some disadvantages. Can't have the cake and eat it too:

https://arxiv.org/html/2404.08801v1 Meta Megalodon

https://arxiv.org/html/2404.07143v1 Google Infini-Attention

https://arxiv.org/html/2402.13753v1 LongRoPE

and a ton more

underlines··on Fast and secure translation on your local machine with a GUI
The models used, without really trying them yet, seem to be much older and much worse compared to seamless-m4t-v2 [1] which is multi-modal and support the tasks of:

Speech-to-speech translation (S2ST) Speech-to-text translation (S2TT) Text-to-speech translation (T2ST) Text-to-text translation (T2TT) Automatic speech recognition (ASR).

across

101 languages for speech input. 96 Languages for text input/output. 35 languages for speech output.

I tried it for low resource languages like Thai to German for text and audio, and it works quite well.

1 https://huggingface.co/facebook/seamless-m4t-v2-large

underlines··on Jpegli: A new JPEG coding library
JPEGLI = A small JPEG

The suffix -li is used in Swiss German dialects. It forms a diminutive of the root word, by adding -li to the end of the root word to convey the smallness of the object and to convey a sense of intimacy or endearment.

This obviously comes out of Google Zürich.

Other notable Google projects using Swiss German:

https://github.com/google/gipfeli high-speed compression

Gipfeli = Croissant

https://github.com/google/guetzli perceptual JPEG encoder

Guetzli = Cookie

https://github.com/weggli-rs/weggli semantic search tool

Weggli = Bread roll

https://github.com/google/brotli lossless compression

Brötli = Small bread

underlines··on DBRX: A new open LLM
This paper partially finds disagreeing evidence: https://arxiv.org/abs/2403.17887
underlines··on DBRX: A new open LLM
Waiting for Mixed Quantization with MQQ and MoE Offloading [1]. With that I was able to run Mistral 8x7B on my 10 GB VRAM rtx3080... This should work for DBRX and should shave off a ton of VRAM requirement.

1. https://github.com/dvmazur/mixtral-offloading?tab=readme-ov-...

underlines··on Mutt on Windows Without WSL
I can't say WSL2 eats host RAM or vRAM:

I moved all my ML training and finetuning stuff (OSS LLMs, TTS, STT, Text2Img) to WSL2 and have only minimal overhead. A clean Win11 host eats away less than 2GB and I love how you can just use your cuda devices on WSL2, use something like micromamba and cmake your wheels while still being able to switch to Win tools whenever necessary. Idle Cuda devices use around 0.5GB vRAM.

Especially the two new experimental features in the last update of WSL2 added a nice QoL improvements:

- autoMemoryReclaim – Makes the WSL VM shrink in memory as you use it by reclaiming cached memory

- Sparse VHD – Automatically shrinks the WSL virtual hard disk (VHD) as you use it

underlines··on Show HN: Not sure you're talking to a human? Create a human check
In certain scenarios: Why would it matter if written communication is done with a human or not?

A) If it's a support chat and the AI helps you as well (or better) than a human support assistant (especially if they implemented function calling and allow it to access and modify DB entries in their core application, it could be better than a human in helping you.

B) If it's a scam/spam call, not so much.

underlines··on 2600.network Dialup Service
For reference, 2600 is a famous phreaking and hacking e-zine from the earlier days. The name 2600 itself is a reference to the 2600hz frequency used in Phone phreaking.

https://en.m.wikipedia.org/wiki/2600:_The_Hacker_Quarterly

underlines··on AI Infrastructure Landscape
I do something like that for open source:

https://github.com/underlines/awesome-ml

But it lost a bit of traction lately.

It needs re-work for the categories, or better, a tagging system, because these products and libraries can sit in more than one space.

Plus it either needs massive collaboration, or some form of automation (with an LLM and indexer), as I can't keep up with it.

underlines··on Analyzing Spotify Stream History
I collected 10 years worth of my complete listening habits via a lastfm plugin, in the past for Winamp, then for iPhone/Android Media players and finally within Spotify. Now lastfm has shut down their free API access and visualization tools stopped working. I printed a 2x1m large poster of my over 10 years of data as a beautiful LastGraph visualization.
underlines··on Better Call GPT: Comparing large language models against lawyers [pdf]
Conflict of interest in the paper, as this mainly is a PR piece from Onit:

"Onit Announces Generative AI-Powered Virtual Legal Operations Assistant for In-house Counsel"

underlines··on Show HN: Tool to calculate how much milk is needed to make an amount of cheese
Swiss here who worked in the national dairy laboratory before switching into IT.

What does this calculator mean by the cheese type "Swiss"?

underlines··on OpenVoice: Versatile Instant Voice Cloning
True. But the better way forward IMHO is to give access to technology in equal ways, instead of keeping it in the hands of a few who promise they are the good ones. Because then an chilling effect can happen, and the playing field is level.
underlines··on OpenVoice: Versatile Instant Voice Cloning
This aera is barely new. Look at how old some of the projects are:

https://github.com/underlines/awesome-ml/blob/master/audio-a...

The thing that changes is the complexity to run it. I was training my wife's voice and my voice for fun and needed 15min of audio and trained on my 3080 for 40 minutes.

Now it's 2 Minutes.

← PreviousPage 3 of 9Next →