HNHacker News
TopNewBestAskShowJobs

yagudaev

36 karma · joined February 11, 2015

submissionscomments
yagudaev··on Politician reads AI prompt during assembly
* AI Response. Not prompt.

I was expecting him to read a proper prompt. The prompt would be what he would tell the AI system.

“You are a speech writer, an expert on XYZ… your task is to write ABC…”

It’s not too bad, at least he is using ai. And agree we can replace many of them with AI. At least we will have the opportunity to talk to the ai directly about concerns.

yagudaev··on All of Paul Graham's essays as Audio
Hi all ,

I converted all of Paul Graham's essays to audio to make it easier to consume.

https://www.audiowaveai.com/playlists/paul-graham

Listening to his essays from 30 years ago up to the present day changed the way I think of startups.

I tried summarizing it with AI and reading summaries, but honestly, it doesn't do it justice.

Rather than binge-watching a Netflix series, just add this to your podcast player and listen to it.

Hope it helps .

P.S.

If you have any thoughts on how to make it better, please reply below.

I'd love to organize it more and make it even more accessible.

yagudaev··on I've built my first successful side project, and I hate it
Thank you for sharing that honest experience you had.

Here is a link to an audio for those of us who like to listen to it instead of reading : https://www.audiowaveai.com/p/2310-ive-built-my-first-succes...

yagudaev··on React.useState for 3 years, before finding the limits
Hi everyone ,

Wrote a short article about my state management journey in React. The tl;dr; is 10 years of react, I decided to keep things simple and only `useState` and basic react hooks.

For the last 3 years, I've been using basic react hooks and it was more than enough. Until, I tried to build a more complex UI and finally needed a bit more.

Endedup using Zustand to help manage the state and avoid stale inter-connected state.

Hope it helps someone out there

yagudaev··on Show HN: Wordsnapper – Rank your ideas by search volume
This is a great idea and I’m glad you built it .

Looking forward to trying it out over the next few weeks. I’ve been building 52 startups in 52 weeks and documenting it under 52shipped.com

yagudaev··on OpenAI Email about Unsupported API Traffic on Vercel and Cloudflare
A few makers like myself got a message this morning from OpenAI warning us about access from unsupported regions. It was cryptic and confusing.

We suspect it has to do with Edge function being used on Vercel which in turn uses Cloudflare. Thus, Vercel and cloudflare customers might be affected.

yagudaev··on Show HN: Affordable text-to-speech for long-form content
Yeah great point, I will simplify the pricing and just say "Listen up to 10hrs of audio" instead.

It is a lot more clear and avoids any misunderstandings. Just need to make the changes to the app to do that.

Thanks for the great feedback

yagudaev··on Show HN: Affordable text-to-speech for long-form content
Thanks so much yeah a few people asked for an API to make it easy too. Added it to my list of TODOs :)
yagudaev··on Show HN: Affordable text-to-speech for long-form content
Totally fair, and I was really excited about the iOS safari built-in version. It didn't work for the book I bought and then after using AI bases TSS, I just couldn't go back.

I would love to be able to run the models on the device, and it will come in the future. The OSS models are not quite there from a quality standpoint yet. But they will surly get there

yagudaev··on Show HN: Affordable text-to-speech for long-form content
Thanks a lot Nayam :). Let me know how it goes when you get a chance to try it, added chatbox to make it easier to leave feedback now
yagudaev··on Show HN: Affordable text-to-speech for long-form content
Thanks so much :). The pricing is towards the bottom, should I just add a link to the footer/header to make it more visible scrollable you think?
yagudaev··on Show HN: Affordable text-to-speech for long-form content
Thank you so much! Fixed it now :)
yagudaev··on Show HN: Affordable text-to-speech for long-form content
Just OpenAI TTS for now, tried a bunch of others and they were no where as good yet sadly.

There is some promise from Myshell models on huggingface and I hope to see them keep evolving.

yagudaev··on Show HN: Affordable text-to-speech for long-form content
Yeah I tried a bunch of them and OpenAI's TTS was by far the best.

Outside of that standard tech stack Next.js, Postgres, TailwindCSS.

It is still early days for ML TTS, and it will be exciting to see the compute requirements drop and for it to run on the device. OSS models have some promise, but still not there from quality perspective.

yagudaev··on Show HN: Affordable text-to-speech for long-form content
Hi, it does split chapters with AudiowaveAI using markdown. It converts other formats to markdown and uses that.

So you can either copy and paste a markdown file there or upload it. PDF and epub would work too, but it is a bit more finky.

I'm happy to help; there are a few authors who contacted me, and I'm releasing an HD quality for authors this week, too (less static noise).

Ping me at michael@audiowaveai.com

yagudaev··on Show HN: Affordable text-to-speech for long-form content
Correct, it's OpenAI TTS.

Costs still need to come down and they will in the future, especially with OSS models improving daily.

Will share a comparison table I made in the near future (just need to clean it up).

yagudaev··on Show HN: Affordable text-to-speech for long-form content
Using OpenAI TTS, it costs ~$10 for ~10hrs. So at $15, margin is 50%.

The costs will go down in the future, I hope, and there are promising open-source projects coming up. Their quality is still pretty subpar and I have a comparison table I'll share soon on HN.

Otherwise it is Next.js on Vercel with Postgres DB.

Hope it helps

yagudaev··on Show HN: Affordable text-to-speech for long-form content
Nice, well you can download each audio file you convert to your device at any time.

I found a single file didn't make sense as you lose your place in it easily.

Instead splitting each chapter into a file seemed to do well.

But, you really do want a listening app. Otherwise it get harder to share and listen to on your phone.

So far I created a listening app as a PWA for the time being.

yagudaev··on Show HN: Affordable text-to-speech for long-form content
Super cool! Thanks for sharing that . Just clean up the UI a little to make it more attractive and hide certain options.

I also experimented with a desktop app and tried to run open-source models locally.

Being a "real programmer" actually hurts you, I had a lot of things to unlearn to just ship fast and keep iterating. I was too stuck looking for the "best practices" or for it to be "just right" (code words for perfectionism).

So keep iterating and writing many projects . This is project 12 in 16 weeks. I've been doing this challenge of 52 startups in 52 weeks. It's been tremendously helpful. (more about it: http://52shipped.com)

yagudaev··on Show HN: Affordable text-to-speech for long-form content
Super cool! A lot of what you are describing I want to do in the future too.

The issue I personally found with traditional TTS is the lack of emotional range and lack of thoughtful pauses. ML models are better at this and picking up on small queues that are hard to program into a TTS otherwise.

I love the iPhone on Safari has a built-in TTS now and was excited to use it. It actually didn't work on Make by Pieter levels after I bought it. So I went to explorer other options. After I started listening to AI generated TTS, I just couldn't go back. It's like 270p vs 2160p (4K).

yagudaev··on Show HN: Affordable text-to-speech for long-form content
Thank you so much ; you've given me a lot to think about.

I'll admit I don't listen to a lot of fiction content, I prefer to watch it.

I do plan on adding additional voices and one brand of fiction I do like listening to is stories for adults. I used it on the Calm app to fall asleep before, and it is great for calming the mind.

For the time being the product is best for non-fiction, but I'll keep an eye on opportunities to make it work better for fiction.

Thank you also for opening my mind on how copy is perceived by blind users. It will be an honour to help them more. Empowering others through tech is why a lot of us work so hard. This is a great reminder.

Any golden classics I should try for fiction?

yagudaev··on Show HN: Affordable text-to-speech for long-form content
Hi Andrew , these are fantastic questions. Let me answer them one at a time:

> How do I delete projects?

Three dots on the side of the project, you can delete it

> I must have tapped three times after submitting a Wikipedia article and it created three projects that apparently cannot be deleted.

> How do I delete my account?

Just email me support@audiowaveai.com with form that email and I'll delete it of you. Still MVP no functionality for that yet.

> And for $15 I get credits. How many credits do I get foe $15? Is each credit a word translate? 1 credit == 1 word translated to audio?

1 credit = 1 character. You are right I need to be more clear on it. $15 would give your about 10hrs of audio or 100 articles (~5-6mins). ElevenLabs will cost you $99 for the same audio.

yagudaev··on Show HN: The best Podcast Details + Podcast Search API
Great job on explaining the different podcasting APIs available and how Taddy became a need with your previous product .

Looking forward to playing around with the API soon

yagudaev··on Show HN: Use GIFs to market your product
Great job! Nice simple and clear idea and reasonable pricing.

Bookmarked for future use

yagudaev··on Show HN: Meme creator with agile/scrum based templates
Nice job creating fun little project .

Made me laugh, so mission accomplished!

yagudaev··on Lottie – Use after effects animations in web and native apps
It is pretty great . I used it on a RN project a few months ago.

There is an amazing collection of animations to choose from here: https://lottiefiles.com/.

You can also buy after effects templates on many websites and use them for motion graphics.

yagudaev··on Show HN: LED Website Indicator – an LED lights up when someone visits your site
That is a cool idea! I have a spare LED panel from a fun IOT project I did one X-MAS, I should try this .

It would be cool to see a video and get the github repo for it

yagudaev··on Ruby Together and Ruby Central, coming together
Great news! Hopefully it can raise more money to help advance the language and tooling. Typescript as an example has been amazing for the JS community. Having something like that for Ruby would be a game changer. (Yes, I know Ruby 3 types)
yagudaev··on Ruby Together and Ruby Central, coming together
For those of us not familiar, can you link to an article or expand on it a little?
yagudaev··on How Phoenix LiveView Works
I’ve tried .Net Blazor, Stimulus Reflex and Phoenix Livewire. Blazor is by far the closest as it allows you to run Native client-side code and server-side code.

The conceptual model didn’t work for me. Using client-side frameworks with json over the wire is a simple model. Ideally the server decides where things are rendered.

These are important tech innovations, however the missing piece is still WASM. Being able to write the full-stack in one language. A step further WASM enables is interoperability between languages. Being able to call a Python ML lib client-side from your JS/Ruby/C# would be a game changer.

Sadly, WASM for ruby and Python is not quite there yet.

Page 1 of 2Next →