HNHacker News
TopNewBestAskShowJobs

userhacker

44 karma · joined February 1, 2016

submissionscomments
userhacker··on Show HN: An open source framework for voice assistants
Just made https://feycher.com thats similar, but has realtime lip syncing as well. Let me know if you are interested and we can chat
userhacker··on Ask HN: What apps would you like to see on Vision Pro
A new age of empires game or any top down real-time strategy game.
userhacker··on Insanely Fast Whisper
I'm the creator of Revoldiv.com, We do speaker diarization and transcription at the same time. Give it a try.
userhacker··on MacWhisper: Transcribe audio files on your Mac
Try to upload it on https://revoldiv.com/ we pre-process the file to make it a little Intelligible and you can supply your context when uploading.
userhacker··on MacWhisper: Transcribe audio files on your Mac
Good point but the problem with local hosting is that if you want to use the larger models it will take a long time to transcribe a file. We use multiple gpus and we do speaker detection, sound detection and it is has a rich audio editor.
userhacker··on MacWhisper: Transcribe audio files on your Mac
If you want a quick and free web transcription and editor tool, We've built https://revoldiv.com/ with speaker detection and timestamps. Takes less than a minute to transcribe 1 hour long video/audio
userhacker··on Whispers of A.I.’s Modular Future
On the end user side Revoldiv.com lets you pick any podcast you want and transcribe it
userhacker··on Launch HN: Buildt (YC W23) – Conversational semantic code search
Nice product! Any integration planned for jetbrain ides?
userhacker··on Introducing ChatGPT and Whisper APIs
contact us at team@revoldiv.com and we are offering an API on a case by case basis
userhacker··on Introducing ChatGPT and Whisper APIs
I suggest you give revoldiv.com a try, We use whisper and other models together. You can upload very large files and get an hour long file transcription in less than 30 seconds. We use intelligent chunking so that the model doesn't lose context. We are looking to increase the limit even more in the coming weeks. It's also free to transcribe any video/audio with word level timestamps.
userhacker··on OpenAI quietly launched Whisper V2 in a GitHub commit
Can you send me the audio that caused it, you can email me at team AT revoldiv .com. If there is going to be a lot of interest, yes we can provide it as an api service. Our service has some niceties like word level timestamp, paragraph separation, sound detection etc... for now it is a free service you can use as much as you want
userhacker··on OpenAI quietly launched Whisper V2 in a GitHub commit
For revoldiv.com we have profiled, many gpus, the best one is 4090. We do a lot of intelligent chunking and detect word boundaries and run the model in parallel in multiple gpus and we get about 40 to 50 seconds for an hour long audio but without expect 7 minutes for an hour long audio on tesla t4

  on tesla-t4-30gb-memory-8vcpu google cloud
   on tiny and tiny.en
    for 10 minute = 30 seconds
   on medium
    for 10 minute = 1m 30s
    for 60 minute = 7m
   on large
    for 60 miutes = 13m
  on NVIDIA GeForce RTX 4090
   on tiny
    for 10-minute = 5.5 seconds
    for 60-minute = 35 seconds
   on base
    for 10-minute = 7 seconds
    for 60-minute = 50 seconds
   on small
    for 10-minute = 14 seconds
    for 60-minute = 1 min 35 sec
   on medium
    for 10-minute = 26 seconds
    for 60-minute = 3 mins
   on large
    for 10-minute = 40 seconds
    for 60-minute = 3 min 54 sec
userhacker··on OpenAI quietly launched Whisper V2 in a GitHub commit
It's not a model you can run on your own server but a free service on revoldiv.com. You can expect 40 to 50 second wait time to transcribe an hour long video/audio. We combine whisper with our model to get word level timestamps, paragraph separation and sound detections like laughter, music etc... We recently added very basic podcast search and transcription.
userhacker··on OpenAI quietly launched Whisper V2 in a GitHub commit
> https://modal-labs-whisper-pod-transcriber-fastapi-app.modal...

Interesting, which model are you using? We use the medium model which is the sweet spot between time/performance ratio. We also chunk, We try to detect words and silences to do better chunking at word boundaries but if you do more chunking and you don't get the word boundaries right it seems like whisper loses some context and the accuracy suffers. We will soon support longer hours. We just want to make sure the wait time for transcription doesn't suffer for most users. But great demo, reach out to me if you want to collaborate

userhacker··on OpenAI quietly launched Whisper V2 in a GitHub commit
I recently swapped out the AI model for voice transcription on revoldiv.com and replaced it with Whisper. The results have been truly impressive - even the smaller models outperform and generalize better than any other options on the market. If you want to give it a try, our model is capable of faster transcription by utilizing multiple GPUs and some other enhancements, and it is all free
userhacker··on Tell HN: Otter.ai bot recording meetings without consent
I created revoldiv.com. It's privacy focused and login is not required to transcribe. You can record your meeting and upload the video or audio to transcribe it
userhacker··on Astrofox – Turn Audio into Videos
You can use revoldiv.com. If you go to export and choose audiogram, it will convert the audio/video you uploaded to text and create an audiogram.
userhacker··on Show HN: Recut automatically removes silence from videos – built with Tauri
Thanks yes it is, we implemented all the ai models in house, that cuts our cost.
userhacker··on Show HN: Recut automatically removes silence from videos – built with Tauri
You can use revoldiv.com to cut out filler words or any words of your choosing, after you upload your file and it finishes sound detection, you can click on the search box to bring up the toolbar to delete sounds
userhacker··on Show HN: Convert audio or video files to text
Hey IndySun, We are in the process of writing one, but we store the file to transcribe and it gets deleted automatically after a certain amount of time has elapsed (this varies on the server configuration, but in a short amount of time). No human accesses the audio or text transcription you have uploaded
userhacker··on Show HN: Convert audio or video files to text
The reason for capping it at an hour is, since the service is free we want to make the experience fast for everybody. We are gauging how our service is going to be used and allocating resources accordingly. We will allow longer transcriptions in the near future.
userhacker··on Show HN: Convert audio or video files to text
I'll hangout here to answer any questions
userhacker··on Show HN: Podcast Audiograms – Make clips with captions, title and audio waves
Pretty cool application, we are also working on a similar tool at https://revoldiv.com/
userhacker··on Ask HN: What browser extensions are a must-have in 2021?
https://chrome.google.com/webstore/detail/text-to-speech/bkj... has a nice text to speech software, something like what the edge browser has
userhacker··on Show HN: Cleanvoice – Automated Podcast Editing
revoldiv.com has a similar feature set
userhacker··on Source: Google Hangouts for consumers will be shutting down sometime in 2020
If you are on a mac, I created a simple web wrapper for the google voice website and puts it on the menubar, you can SMS on the desktop that way. Link to the app voicenotifies.com
userhacker··on Coinbase Index Fund
unless you change your holdings into fiat thats not the case. You can rebalance with exchanging among your coins
userhacker··on Google Trips is a killer travel app for the modern tourist
as a hack, you can use a regular gmail address and forward all your emails to that address. you can also respond from a regular gmail address to your abc@yourname.com by changing the from field (you have to verify you own the email address). currently using a similar setup, custom email address but responding from a regular gmail address
userhacker··on I am a fast webpage
yup Atlanta is slow for me too
userhacker··on Instapaper is joining Pinterest
I second Readability, it works great for article heavy webpages. I used it to build a reading time estimator for chrome https://chrome.google.com/webstore/detail/read-time/nccohhim... and its open source https://github.com/usergit/read-time bonus, you can click on the extension to show only the main content of the page
Page 1 of 2Next →