235 karma · joined February 3, 2024
DM: properbrew@blazingbanana.com
Also great video, it hooked me.
One of the hardest parts I've found is the diarisation (who said what) side of things. Trying to tune this and have it working in a way that doesn't absolutely grind the laptop to a halt or take forever to complete has been _hard_ but also extremely rewarding.
Another part has been the fine tuning side of the Phi-4 model, I'm on version 10 now, getting that pipeline down was a journey in itself, but I've got some great results. I wrote a bit about it in a comment here - https://news.ycombinator.com/item?id=48385906#48389625
I absolutely love working on this, I still wake up and the first thing I think about is voice transcription pipelines (sad I know), but I'm excited to see how much further performance and utility I can squeeze out.
Do you have any old documentation that it's picking up and referencing? If you set all claude settings back to default do you see the same issue?
I'm from a hardware / networking / infrastructure background. I've had extensive exposure to (web) application development as I'm working closely with development teams and I do have the bash/powershell scripting knowledge.
But honestly, if I tried this "the old fashioned way" it probably would have taken me about 6 to 7 years to develop that application, that's an optimistic estimate. You really do have to have a passion for what you're building, I didn't know that voice transcription and local LLMs would be such a driving force for me, but it's all I think about, so much that I find it hard to go to sleep sometimes.
It's taken _a lot_ of time and effort, but this is an example of what can be developed using LLMs alone.
You have to have dedication and a goal to reach, but you can absolutely build anything if you're building with the right foundations in mind.
I'm not sure if it's because I've iterated through so many sites that LLMs have produced that "slop" is instantly recognisable and it just feels soulless.
Not like web pages ever had a soul, but it's not there on the generic LLM generated sites.
Sure an app can be built and spun up in an afternoon, but are you willing to spend another 6 months ironing out all those little bugs, tuning it a bit, testing, tweaking, testing etc.
2 hours later he's got a fully working piece of local software that does exactly what he wants, yet yours is not able to even sort dates correctly. Feel free to download it if you want to see for yourself, I didn't even do any UI tweaks as this was just a tool for him to use:
Linux - https://downloads.blazingbanana.com/whistle-subtitles/unstab...
Windows - https://downloads.blazingbanana.com/whistle-subtitles/unstab...
Mac - https://downloads.blazingbanana.com/whistle-subtitles/unstab...
How can there be such a massive gap in what can be produced?
This seems so wide reaching if it's catching simple things like explaining a paper. Does this also refuse to help with any already developed training pipelines?
I can kind of understand the generation of synthetic data, but nerfing the assistance of training pipelines just seems like a really shitty thing to do.
Has to be a personal debit (not credit) card - https://www.gov.uk/pay-corporation-tax/debit-or-credit-card
Yea absolutely, but man, where to even start, it is very specific.
Fundementally I didn't use any wrappers like unsloth or axolotl, although I have used the latter before a year or two back and it was good, but I needed something very very custom. I also wanted the whole fine tuning pipeline to exported OpenVino model to be seamless.
I heavily leaned on codex, claude and some manual sleuthing around the internet to understand what I needed. I'd played about with QLoRA finetuning with axolotl before and felt most comfortable with that. So I needed to keep everything as stripped down as possible and figured I can just utilise the 3 main huggingface libraries (transformers, peft and datasets) and also bitsandbytes (as suggested by claude to quantize the model to keep this working on my GPU) along with some custom scripts generated by claude/codex (each cross referencing each other) that will do the different stages of the training run.
The next part was the data. Obviously didn't have access to thousands of meetings and associated output documents but I did have a 3090ti sitting there and a codex subscription. So I set about working out what format I needed the data in (many thanks again, to claude/codex) and started generating hundreds of different transcripts, different amounts of speakers, content, tones, subjects, spelling mistakes - like all the different things you could think a meeting would have. Then it's a case of actually generating a good meeting document off the back of the transcripts and creating the "gold standard" that we'd use.
I'm going to gloss over a lot here as I'd rather not detail it as it relates to some propriatary stuff that I had to work through, but you basically pair the transcripts together and run the training.
At the verification stage, there was pretty much 3 things:
1. "just" do some regex string matching to see if there's any of the source transcript key facts in the output to ensure fact preservation. Same with owner fabrication (who said what), I don't want something attributed to someone when it wasn't them that said it and then finally markdown validation.
2. Using codex/claude to validate the transcript and output from the model - I used the latest frontier models, probably overkill for my task, but they were good at the job
3. Finally me going through some actual recordings of myself, groups, meetings and manually verifiying the output
So a fair bit of work, and for context I'm on version 10 now, so it's been a journey!
If you have a very specific idea for local model use you can find a way to make it work very well, you don't even need to have a graphics card or NPU chip. You just have to be extremely constrained in how it's used. I think as a generic chatbot they're not great, I'd use a hosted SOTA model and I'm a big fan of local LLMs myself.
I have worked retail before, and to add onto the things you put it was the lack of problem solving for me that was absolutely mind numbing. Sure there were the little "problems" to solve of shelving, stock order, tidiness etc but it doesn't push the brain (and maybe they're done with that part, which is fair), but until you've experienced it I would be very surprised if this person finds retail better than tech.
I have a few live websites built using LLMs and they will just go for default generic templates and colours if there's no vision.
Scripts in this repo will mirror the site, do the required replacements and OS list update and it'll work on Linux (Only tested on Ubuntu).
Also means that if they retire the page you can still pair your older devices with the unifying receiver. You'll need to use .deb version of Chrome.
And how much with Opus 4.7? 5x?
Some websites can run only on ads. Is it such a bad thing that they would die off?
I say this as someone that likes the old web and has fun hitting the "surprise me" button on https://wiby.me/ (not affiliated) and browsing the random sites. Just giving an alternative view.
The whole point should be that the people in the meeting can actually focus on the conversation rather than half listening while trying to write down enough notes to remember it later. If you're paying people good money to be in a meeting, having them spend half of it doing low quality note taking is a bit mad.
This is basically the reason I've been building Whistle Enterprise (https://whistle-enterprise.com). I'd much rather have something where I choose to record a meeting, process it locally, generate the document and then decide what to keep or delete. I mean yea, it still creates a record so it doesn't solve the legal / discovery side, but at least you're not also adding a random third party into the middle of every conversation.
It looks like generic AI slop, the site doesn't even render the headings for their SEO spam "Curated AI Tool Collections by Use Case" section properly and they're half cut off. The images all have the very distinct generic AI hue to them without any attempt of bringing it into a specific style or brand.
Who is upvoting this stuff? Do people not care? Is it just bots gaming the system? Am I an old man shouting at a cloud?
> 23 March 2026
This is something I need to do, have no knowledge about and is definitely going to be harder than building the thing.
> it would be cool to record the calls and feed them to AI for some simple/crude auto-summary which automatically pulls out rejection reason, concerns, interesting points etc
This is very close to what the software I've been building over the last few months is, offline note transcription with summary file generation in a nicely formatted PDF/Docx using local models. Codex is available if you don't want to do the inference on your laptop (I will sort claude out soon as well).
If it sounds like something of interest then feel free to send me a message. More than happy to send you a 30 day trial version.
Why is there this massive disparity in experience? Is it the automatic routing that ChatGPT auto is doing? Does it just so happen that I've hit all the "common" issues (one was flashing an ESP32 to play around with WiFi motion detection - https://github.com/francescopace/espectre) but even then, I just don't get this "ChatGPT is shit" output that even the author is seeing.