HNHacker News
TopNewBestAskShowJobs

narrationbox

192 karma · joined January 30, 2020

We do generative ML, specializing in speech technology.

https://narrationbox.com

Cost effective voiceovers and narrations at scale. Voices for agentic AI and audiobooks.

Check out our blog here: https://narrationbox.com/blog

Follow us at: https://twitter.com/narrationbox

Subscribe to our newsletter for generative text to speech, voice cloning, and accent conversion research.

submissionscomments
narrationbox··on ElevenLabs, TwelveLabs, ThirteenLabs
Are you still working on it?
narrationbox··on How we made a text-to-speech model respond in sub-50 ms
Haven't read the full report yet, just a quick question. Are your numbers for cold start without pre fill or is it after warmed cache?
narrationbox··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
> Typical production voice agent is an ensemble of three models: STT, an LLM, and TTS.

To use a claudism, I would like to push back on this. The industry is very much moving towards one-model-does-all end to end trained similar to LLMs and VLMs. Mostly for latency reasons and partially because the results for the end to end trained models are just so much better than those using three pieces architectures.

I think most of the value prop is in automatic evals, not routing specifically. A better pitch for you would be "the LM Arena of voice models" rather than comparing yourself to openrouter because the value add is rather questionable. For TTS specifically, the current SOTA for production systems are all using prompt based voice gen i.e. instead of having 10 different Tacotron models trained on 10 different models, these days it's all a single large model and the "style" is a prompt in the system prompt. The input is usually something like

  <System prompt>
  Speak in a deep smooth voice similar to a documentary narrator
  </System prompt>
  <Text to Narrate>
  Speko is the ultimate evaluation platform for voice agents. We do automatic  evals.
  </Text to Narrate>
It's the same for voice cloning too, you just pass the reference speech as an input file for all generations. A lot of systems don't have any separate style vector extraction step or model-specific fine-tuning anymore.

So something like OpenRouter for voices offer questionable value given that stakeholders usually make this sort of decisions once at the start of the project. On the other hand if you can offer automatic evals and figure out which prompts give the most similar results across different voice providers, that would offer a lot more value. It would be nice to be able to switch from e.g. Grok voice agents to ChatGPT voice agents knowing that the output style won't change too much. There are many companies now with evals as a core business model: LM Arena, Artificial Analysis, Prompt foo (before they got acquired and pivoted to security only) so many take a look at them.

Source: we have been building TTS systems for over a decade too https://narrationbox.com

narrationbox··on VibeVoice: Open-source frontier voice AI
Yes, the SOTA is currently much more advanced.
narrationbox··on Eleven v3
Give us a try, I think we are what you are looking for

https://narrationbox.com

narrationbox··on Google Illuminate: Books and papers turned into audio
A lot of our customers use us [0] for that, it works pretty well if executed properly. The voiceovers work best as inserts into an existing podcast. If you see the articles of major news orgs like NYT, they often have a (usually) machine narrated voiceover.

[0] https://narrationbox.com

narrationbox··on Another Text to Speech API
Yeah, neural codecs are pretty amazing. The most incredible part is that they can do compression well across the temporal domain, something which has been non-trivial.
narrationbox··on Project Gutenberg releases 5k free audiobooks
Their sentence segmentation heuristics were not configured correctly. It's not an inherent limitation of the technology itself.

The newer transformer based generators are a bit better in this regard (since they can maintain a longer context window, not just in short tiny snippets).

narrationbox··on PlayHT2.0: State-of-the-Art Generative Voice AI Model for Conversational Speech
Mel + multispeaker vocoder is very much a classic (tacotron era) TTS approach
narrationbox··on DeepFilterNet: Noise supression using deep filtering
Since it does the signal processing in the Fourier domain, does this suffer from audio artefacts e.g. hissing in the output? Torch's inverse STFT uses Griffin-Lim which is probabilistic and if you don't train it sufficiently, you may sometimes get noise in the output.

https://pytorch.org/docs/stable/generated/torch.istft.html#t...

An alternative would be to use a vocoder network (or just target a neural speech codec like SoundStream).

narrationbox··on US Marines defeat DARPA robot by hiding under a cardboard box
> Recognizes "human" and recognizes "desk". I sit on desk. Does AI mark it as a desk or as a chair?

Not an issue if the image segmentation is advanced enough. You can train the model to understand "human sitting". It may not generalize to other animals sitting but human action recognition is perfectly possible right now.

narrationbox··on NaturalSpeech: End-to-end text to speech synthesis with human-level quality
Your average mobile processor doesn't have anywhere near enough processing power to run a state of the art text to speech network in real-time. Most text to speech on mobile hardware are stream from the cloud.
narrationbox··on Ask HN: Non-tech professionals on HN?
For the high end stuff no, but many of the lower tier jobs are under threat.
narrationbox··on Show HN: Automated Binance Trading Bot – Buy Low/Sell High
We used to be in this field too (https://kloudtrader.com/narwhal). It is a very crowded market and monetisation is tricky.
narrationbox··on Quickdraw with Google AI
You can throw this together pretty quickly using one of the AutoML APIs.
narrationbox··on YouTube can now warn creators about copyright issues before videos are posted
So if I am launching a Spotify/Audible-style music/audiobook streaming platform, how will the pricing work out? Do we pay you instead of the original author for any user uploaded content? If the original author chooses to upload their content onto our platform, do we scan it with your API and explicitly whitelist it and pay them directly?

Your company looks very cool btw. What's your email?

narrationbox··on Launch HN: Pry (YC W21) – Finance for Founders
What about TransferWise ewallets?
narrationbox··on YouTube can now warn creators about copyright issues before videos are posted
Do you have any public pricing or startup plans?
narrationbox··on Launch HN: Pry (YC W21) – Finance for Founders
Does the accounting system support Canada and other commonwealth countries? Or is this US only?
narrationbox··on Launch HN: Enombic (YC S20) – Create your own stock indexes
Some ETFs also allow easy investment in stocks of foreign countries without requiring additional brokerages on behalf of the end user. I presume this does not offer that functionality?
narrationbox··on Show HN: Turn scripts into fine-tuned voices via Wiki markups
I will look into it, Wiki2SSML looks very handy.
narrationbox··on Show HN: Turn scripts into fine-tuned voices via Wiki markups
It looks great, have you considered adding a visual editor?

We have one for our systems: https://narrationbox.com

narrationbox··on Launch HN: Feroot (YC W21) – security scanner for front-end JavaScript code
Are there any free plans or discounts for HN users?
narrationbox··on Clubhouse for X – The emerging trend of audio-only applications
We never expected audio to be this popular either when we built Narration Box :)
narrationbox··on Lua and Python (2020)
Lua is a great language, but the fragmentation between versions makes it difficult to be adopted outside of niche embedded areas. On a side note, we maintain a Lua newsletter but haven't had time to update it. This article looks perfect for a new edition.

https://luadigest.netlify.app

narrationbox··on WhatsApp gives users an ultimatum: Share data with Facebook or stop using app
You can try locking down the app, it's not ideal but it is better than nothing:

https://news.ycombinator.com/item?id=25664130

narrationbox··on WhatsApp gives users an ultimatum: Share data with Facebook or stop using app
A while back I wrote about this

https://medium.com/@kloudtrader/reducing-whatsapp-digital-fo...

Not sure if it still applies to the latest version of Android and WhatsApp but it might help. However it only mitigates certain real-time tracking and contact discovery, not to mention switching profiles is somewhat of a hassle.

narrationbox··on I Created a Trading Simulator
This looks great! Congrats on launching. We used to be in this space too. The most tricky part is performance, especially if you are backtesting against an algorithm instead of manual trading. Having to wait for the results on 30 years of trading data can get rather annoying at times. If you do implement support for algorithmic trading, it might be helpful to rewrite the core in WebAssembly.
narrationbox··on Stripe Treasury
Interesting, it seems their products are geo-locked. The Canadian site still shows "Request invite".
narrationbox··on Stripe Treasury
Is this like Stripe Issuing where it is only available to a small number of companies? A lot of newer Stripe products seems to be available to large enterprises only. We applied for that waitlist multiple times but never heard back. What are the revenue or scale requirements for your targeted customers? Is it suitable for fintech e-wallets/banks?
Page 1 of 4Next →