Show HN: Unreal Speech – Text-to-Speech API
unrealspeech.com
unrealspeech.com
Shady, really shady. It's a shame, I would not mind a good competitor to Polly/Google.
But yeah sometimes I'd have to regenerate before I got non-glitchy audio. Still though, very cool.
I get lots of weird random artifacts in Unreal. Sometimes it's a weird "vibrato" on some words, some words are unrealistically raspy. Plus it's a bit too sibilant in general compared to AWS. Clicking the "Redo" button fixes it but then other artifacts crop up in other places, which leads me to believe this is completely fixable.
And what do you mean by cherry-picked? I don't think the Amazon example is cherry picked, you can type your own custom text there! They're just using the AWS API themselves. Maybe I misunderstood you.
Anyway, it's a bit unfair because Amazon probably spent millions in their product, but I wouldn't exactly call it better...
Your heading makes it seem like you want the direct comparison to AWS polly so maybe add a table to directly compare different aspects of your product vs aws that make it better. Sound quality is just one attribute to compare. What about SDKs, limits, code samples, use cases, more nitty gritty sound comparison details, etc.
Pricing should also be more transparent if you want to compare aws to yours - what maths did you do to get 8x cheaper because at first glance that is misleading.
I agree on the table idea to more candidly and clearly compare other aspects. This launch/experiment was that the quality/cost would be the main factors, and we’d go from there, iterating and customizing it to work for early customers.
The 8x math is per 1M characters. I do see that since we’re charging a subscription, it may not be a fair comparison. But the minimum commitment is is so small that I thought it wouldn’t matter for the customers I’m targeting right now. I do think it can be misleading because people might expect pay-as-you-go.
We aren’t able to provide pay-as-you-go right now, so I’ll look into updating the copy or how we communicate the subscription model!
A. Impressive display of capability
B. A very clever choice as it's recognizable but not as universal as say Obama's voice or Joe Rogan's would be
C. Brilliant marketing
I probably wouldn't offer it as a "real" voice for use in bulk through the API due to the legal concerns, but on the marketing page it's really cool and I would hang on to it. Plus if you get sued it would be great publicity :-)
Frankly, trying to fight against the Goliath, with 0 marketing budget, and I’m desperately hoping to create noise. Breaking a rule or two is something AWS can’t do at their scale.
On the bulk offering, I do have a clear path forward. In short, I can create new synthetic voices. It’s like those this-person-does-not-exist images but for voice. “unreal speech”
This is blunt and clear immoral behavior and you are fully aware of it. You just think that you have the right to reach success not by your own skills, but by stepping on other people, and until they complain you'll keep doing it.
PS: Even if Amazon does the same thing it's no excuse for you. Look up tu quoque fallacy.
Bringing others down will never be an antidote to not building or achieving anything in your life. Your time will be better spent if you focus on bringing yourself up and others around you.
I think you want to get a rise out of someone over the Internet, with your identity hidden, to get attention. Hence, this will be my last comment-- but if you want to turn it around, I'd be down to help.
Been thinking a lot about how to accomplish this myself for a similar product I'm building, glad to hear someone else is thinking about it too!
https://www.theverge.com/2020/4/28/21240488/jay-z-deepfakes-...
has more details
I personally think it's fine, but I'd expect some people finding it searching for speech synthesis for the unreal engine and then getting outraged about getting "tricked".
Not sure what to do about it though. It just jumped to the top of my mind while looking at the landing page.
I spend a lot of time listening to AWS's text-to-speech. It can be distracting at times. But with Unreal Speech's text-to-speech, I'd lose focus incredibly quickly and focus on the issues with cadence or general weirdness (Professor's pitch change is too abrupt, and weirdly gets caught on the word used which throws off the cadence of natural speech).
I might argue, though, you might get used to our voices quickly once you start using it frequently. I've had this feedback before from someone who was very used to AWS's monotonous voice, and he actually changed his opinion after listening to a couple of articles. Previously, I built https://audioread.com which got me to talk to a bunch of users.
kudos to you on having a wonderful idea :-D Guess I'll move on to my next idea
> Yo! Let’s at least chat and brainstorm about it. Would you want to ping me on twitter? @automationism
But honestly, THIS is why I keep coming back to HN, rather than take an adversarial route and needlessly bicker entrepreneurs should be collaborating and utilizing each others strengths.
Well done, I know want to see your progress as your product matures! Who knows, I (a AI/ML student) might have a need for your services yet.
> What the fk did you just fking say about me, you little btch? I'll have you know I graduated top of my class in the Navy Seals, and I've been involved in numerous secret raids on Al-Qaeda, and I have over 300 confirmed kills.
Having "Better" on the title of your landing page probably isn't the best idea. This is such a subjective territory that it's very difficult to say that one voice sounds better than another unless you have a blind poll with at least a few hundred people claiming they preferred your product over AWS.
It doesn't matter if you can argue its better, it only matters if you can show that the vast majority of your target demographic prefers your product over AWS, which I couldn't find any evidence of.
The "8x cheaper than AWS" feels deceptive though. The pricing is not apples to apples so it's only 8x cheaper at the most favorable point in the graph (when spending $1,000 a month and using exactly the number of characters offered). For me that's outrageously more expensive than AWS. It's more like 8x more expensive than AWS rather than 8x cheaper, which is such a dramatic swing I wondered if I was misreading the pricing somehow. My usage on Polly last month was about $75 but the month before was $5 and this month will be closer to $20.
It's a real shame because after hearing the output I was ready to move everything over. I should have looked at pricing before getting excited, but I took the 8x claim at face value.
I use a lot of AWS Polly neural and have listened to a hundred or so hours of Polly output (building an MVP/prototype). The cost of AWS is high (for me as an independent dev) but I've tried several other services and none have been as good at making natural speech (the kind one could tolerate for an audiobook for example).
If this was cheaper or if I could buy a small time license and run it on my own hardware (which takes the heavy costs away from Unreal) I would totally do that. Alternatively, I'd be willing to pay close to AWS pricing for a pay as you go, with a one-at-a-time rate-limit in place (to avoid the scaling/provisioning challenges on the ops side). I know it's not likely to happen, but just wanted to throw it out there.
> Sarah Connor?
and
> Come with me if you want to live.
But the professor edges out on top for:
> Never harm a human or through inaction allow a human to come to harm.
Ed: according to Wikipedia I got it pretty close:
> First Law
> A robot may not injure a human being or, through inaction, allow a human being to come to harm.
> Second Law
> A robot must obey the orders given it by human beings except where such orders would conflict with the First Law.
> Third Law
> A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.
https://en.m.wikipedia.org/wiki/Three_Laws_of_Robotics
Just remember, Asimov himself pointed out through his works that the three laws are not enough! :)
Any good links or HN comment threads that you’d recommend?
Using this slower generative approach it will allow to produce large high-quality enriched audio datasets with parametric text, timings and emotion.
Then you use these datasets to bootstrap in a supervised fashion the existing traditional architectures to make the generation faster.
The usual problem of text-to-speech is that you have to go from a low-information space (aka text) to a high-information space (aka sound). And therefore training is ill-defined because one input text can have several correct sound. But once you have an enriched text with inflections and parameters, speaker embedding, the mapping then become one enriched text to one exact audio and the training become well-defined and easy.
"$99/mo (non-commercial use) too rich for my blood."
Don't get me wrong it's pretty cool tech if you trained your own model.
I'm also having a hard time coming up a use case. Outside of spam and games I'm having a hard time coming up with generative AI is useful at all anyway.
So one would need to consume the maximum every month, for several months, for the price comparison to be true.
And I'm not talking about the quality of your product. I mean to use Jordan Peterson without consent or reference is not bold as other are saying, but is verging on criminal.
How can you even calmly post this on HN?
But that "professor" voice is extremely recognizable as him.
The entrepreneur voice is familar as well but I can't place it. It might be a better blend of more people.
Found this from comments on ProductHunt
> thanks for a great question. Candidly speaking, I’d say we’re taking a “move fast” and “seek forgiveness later” approach. The plan is a) to try to apologize and get a license if there are demands and if that doesn’t work b) create new synthetic voices that are not of real people.
This infuriates me to no end. People with that kind of mentality fuck up the entire startup ecosystem for everyone.
So that makes this intentional and the excuse of "I didn't know any better" no longer plays.
I really don't see the outrage. If I do an impression of Jordan Peterson am I also in violation of a social code?
It's like someone selling stolen credit cards and then claiming they're doing nothing wrong because people who are buying them are the ones committing the crimes.
I was having fun making Jordan Peterson say silly things when I clicked random. First Russian came up and failed, then x-rated sir-mix-a-lot struck!
Entrepreneur = Gary Vaynerchuk
Honestly not a good look.
Otherwise cool demo!
A very good look for many, many people, so I'll assume you are talking about the other guy who I've never heard of.
I guess people get upset when someone on the internet (that they could easily ignore, just as you have) tells them to work harder.
He's also loves promoting arbitrage plays... so I kind of blame him for the insane secondhand markets on a ton of normal stuff. Everyone is buying up all the stock and trying to flip it for a profit so they can be like GaryVee at garage sales. This does bother me. I went to go buy a new pair of shoes I bought 3 years ago for $90, but no one has them. I looked on eBay and they are going for $600. That's madness. It's not just Gary's fault though, it's the companies that opt for "drops" and hype over scale and actually meeting demand. But that's a whole differnet rabbit hole.
Both people have a lot of fans, but also get a lot of hate from a particular demogrpahic. I assume the majority is indiffernet.
Personally, AWS' and Google's pay-as-you-go plan win every time regardless of the "better" and "cheaper" claims IMO. You must introduce pay-as-you-go plans if you really want some good traction. I really like the Entrepreneur voice (the professor sounds like Jordan Peterson)
On more thing, I think in the comparison you have sampled the AWS non-neural voice while using the price of neural voice for the price comparison - this does not sound like a fair comparison to me (but please correct me if I'm wrong and I'll edit my comment).
But for those who are already spending, let's say, $250+ a month on TTS, this is a sweet deal. They are my initial target customers.
We're sampling the Neural AWS Polly Matthew voice. FYI: Neural Matthew ($16/1M chars): https://unreal-tts-live-demo.s3.us-west-1.amazonaws.com/aws/... Standard Matthew ($4/1M chars): https://unreal-tts-live-demo.s3.us-west-1.amazonaws.com/aws/...
Well, our $250/mo is at $2.6 per 1M chars which is even cheaper than Standard Matthew, and I think ours clearly sounds better?
Its like you have a Tupac hologram advertising cheese sicks on TV, no permission, you use his voice and body, but he is alive.
You are using someone's 'likeness' for profit.
Voice might be under protected biometric data as well.
IDK what you mean about biometric data being protected. I'm pretty sure there's no law stopping me from pulling your fingerprint off of a coffee mug you left at Starbucks, creating a high resolution scan of it, and posting it at godmode2019-fingerprint.com
EDIT: after some quick Googling it looks like there are some biometric privacy laws on the books in certain U.S. states that would prevent something like godmode2019-fingerprint.com, but it does not appear to be comprehensive across the US. Not sure about other countries.
The voices sound fine (some are worse than AWS, but some are indeed better). However, as a queer person, putting in sample texts like these[1] kinda put me off from your product no matter how good they are. IMO, that's absolutely uncalled for.
On a professional note, it's very immature( and silly even?) to use a product page to voice hostile idiosyncratic political opinions in general (regardless of whether I agree with them or not).
You're of course entitled to your opinions, and welcome to market your product however you want, though; I'm not trying to encroach on that.
EDIT: they have responded it was text people tried. sorry for the link with Peterson in that case.
Also - if you are storing previous inputs (for any purpose) I’d let people know upfront!
I would maybe have a list of paragraphs from a few select public domain works in there, instead of using unfiltered user input.
You need to make it pull from Wikipedia or some other semi-moderated source, else your Random button is going to turn into Microsoft's Tay real quick. I also notice your Professor voice is pretty clearly Jordan Peterson, which some people may have a problem with.
If you're looking to position yourself as an inclusive company, don't regurgitate text put in by previous users. Because idiots on the internet will idiot. And that idiocy now is cosigned with your company name and logo.
People have been putting stupid stuff into text-to-speech since the early 80s after it was popularised by the SP0256 chip.