HNHacker News
TopNewBestAskShowJobs

dylanbfox

442 karma · joined September 9, 2014

dylan at assemblyai dot com
submissionscomments
dylanbfox··on Launch HN: Kable (YC W22) – All-in-one platform for API products
Looks interesting! Do you guys offer any sort of visibility tools/reports into customer usage of different endpoints, tracking of actual API requests (including payloads/etc) per request, etc? Is it possible to use you guys just for billing if we already have our own auth?
dylanbfox··on Show HN: Top headlines of the world to start the day
Agree. The places I go on the web have become more and more centralized/limited. I think projects like this that help to surface and aggregate interesting content from the web (which is really what I come to HN to find) are great.
dylanbfox··on Show HN: Podcast Audiograms – Make clips with captions, title and audio waves
Nice work! The UI is really simple - and love not having to log in to use it. Have you thought about leveraging the ListenNotes API (https://www.listennotes.com/api/) to automatically pull in the podcast episodes via search vs having to upload them?
dylanbfox··on Greening of the Earth Mitigates Surface Warming (2020)
I agree with all your points - but one thing I think about is: how do we fix what we have today? How do you fix the concrete jungles that most cities are today in the US. Or is it inevitable that more concrete will just be poured over time until some major natural disaster allows for a reset?
dylanbfox··on How to train large deep learning models as a startup
Hi there - OP here - thanks for reading!

This blog is more of an intro to a few high level concepts (multi-GPU and multi-node training, fp32 vs fp16, buying hardware and dedicated machines vs AWS/GCP, etc) for startups that are early into their deep learning journey, and that might need a nudge in the right direction.

If you're looking for a deep dive into the best GPUs to buy (cost/perf, etc), the link in the below comment gives a pretty good overview.

PS - I can send you some benchmarks we did that show (at least for us) Horovod is ~10% faster than DDP for multi-node training FWIW. Email is in my profile!

dylanbfox··on How to train large deep learning models as a startup
Author here. Thanks for your comments!

In general - this is expensive stuff. Training big, accurate models just requires a lot of compute, and there is a "barrier to entry" wrt costs, even if you're able to get those costs down. I think it's similar to startups not really being able to get into the aerospace industry unless they raise lots of funding (ie, Boom Supersonic).

Practically speaking though, for startups without funding, or access to cloud credits, my advice would be to just train the best model you can, with the compute resources you have available. Try to close your first customer with an "MVP" model. Even if your model is not good enough for most customers - you can close one, get some incremental revenue, and keep iterating.

When we first started (2017), I trained models that were ~1/10 the size of our current models on a few K80s in AWS. These models were much worse compared to our models today, but they helped us make incremental progress to get to where we are now.

dylanbfox··on How to train large deep learning models as a startup
Dylan from Assembly here. If you want to send me one of your audio files (my email is in my profile) I'd be happy to send you back the diarized results from our API.

You can also signup for a free account and test from the dashboard without having to write any code if that's easier.

Other than lots of crosstalk in your group conversations - is there anything else challenging about your audio (eg, distance from microphones, background noise, etc?)

dylanbfox··on How to train large deep learning models as a startup
Great question. This is technically referred to as "Wake Word Detection". You run a really small model locally that is just processing 500ms (for example) of audio at a time through a light weight CNN or RNN. The idea here is that it's just binary classification (vs actual speech recognition).

There are some open source libraries that make this relatively easy:

- https://github.com/Kitt-AI/snowboy (looks to be shutdown now) - https://github.com/cmusphinx/pocketsphinx

This avoids having to stream audio 24x7 to a cloud model which would be super expensive. This being said, I'm pretty sure what the Alexa does, for example, is send any positive wake word to a cloud model (that is bigger and more accurate) to verify the prediction of the local wake word detection model AFAIK.

Once you are positive you have a positive wake word detected - that's when you start streaming to an accurate cloud based transcription model like Assembly to minimize costs!

dylanbfox··on How to train large deep learning models as a startup
Interesting. How do you guys manage spot interruptions when training on spot instances?
dylanbfox··on How to train large deep learning models as a startup
This is tricky. The de facto metric to evaluate an ASR model is Word Error Rate (WER). But results can vary widely depending on the pre-processing that's done (or not done) to transcription text before calculating a WER.

For example if you take the WER of "I live in New York" and "i live in new york" the WER would be 60% because you're comparing a capitalized version vs an uncapitalized version.

This is why public WER results vary so widely.

We publish our own WER results and normalize the human and automatic transcription text as much as possible to get as close to "true" numbers as possible. But in reality, we see a lot of people comparing ASR services simply by doing diffs of transcripts.

dylanbfox··on How to train large deep learning models as a startup
> Salary costs are probably even higher than compute costs.

Yes exactly. Managing that much compute requires many humans!

dylanbfox··on How to train large deep learning models as a startup
Dylan from Assembly here. Most of our customers have actually switched over to us from Google - this Launch HN from a YC startup that uses our API goes into a bit more detail if you're interested:

https://news.ycombinator.com/item?id=26251322

My email is in my profile if you want to reach out to chat more!

dylanbfox··on Is Word Error Rate a Good Metric for Speech Recognition Models?
Interesting. It seems like in the "real world" WER is not really the metric that matters, it's more about "is this ASR system performing well to solve my use case" - which is better measured through task-specific metrics like the one you outlined your paper.
dylanbfox··on Is Word Error Rate a Good Metric for Speech Recognition Models?
> since less common words tend to be more important for meaning.

Exactly. Errors with proper nouns are usually more problematic than errors with stop words, yet they're weighted equally in the WER calculation. Ie, deleting "Bob" and "but" both count as a deletion of the same degree according to WER, but we as humans know that deleting "Bob" is potentially a lot more problematic than deleting "but".

dylanbfox··on Is Word Error Rate a Good Metric for Speech Recognition Models?
Yes! Perplexity is a great idea. Although you could technically have a low perplexity prediction that is not similar to the ground truth transcription.

CER is definitely more granular. There are papers that basically count Deletions, for example, as 0.5(D) when calculating WER - since they consider Deletions "less bad", but if these weights aren't standardized then WER scores will be super hard to compare.

Personally I think some metric including some type of perplexity is the way to go.

dylanbfox··on Launch HN: Svix (YC W21) – Webhooks as a Service
Love this idea. Had this on my list of side projects to build for a while - definitely see the use for this. It's one less thing for a dev team to maintain in-house.

What does latency look like on delivering webhooks? From the time your service is hit, to the time when the webhook is sent?

dylanbfox··on Building an end-to-end Speech Recognition model in PyTorch
That's right - most literature does show that encoder-decoder architectures outperform CTC. I think one of the main reasons for this is that CTC assumes the label outputs are conditionally independent of each other, which is a pretty big flaw in that loss function.

The blog does mention Listen-Attend-Spell (which is an encoder-decoder architecture) as an alternative to the CTC model.

dylanbfox··on Building an end-to-end Speech Recognition model in PyTorch
One of the main factors for this is probably due to dataset size. Commercial STT models are trained on 10s of thousands of hours of data from real-world data. Even a decent model architecture is going to perform pretty well on that much data.

Most open source models are trained on Libri, SWB, etc. which are not really big or diverse enough for real-world scenarios.

But to max-out results the devil is in the details IMO (network architecture, optimizer, weight initialization, regularization, data augmentation, hyperparam tuning, etc) which requires a lot of experiments.

dylanbfox··on AssemblyAI: speech-to-text API
Those test sets are definitely more real-world than Libri Clean in our experience, but still are not as "dirty" (noise, muffling, cross talk, recording quality, etc.) as we see in production.
dylanbfox··on AssemblyAI: speech-to-text API
Thank you! Please let me know if you need any help. My email is in my profile.
dylanbfox··on AssemblyAI: speech-to-text API
Thanks for sharing this! It's awesome to see another independent company tackling STT well. If you guys ever want to try to collaborate, let me know! My email is in my profile.
dylanbfox··on AssemblyAI: speech-to-text API
There are some good open source STT solutions out there like Mozilla's DeepSpeech (https://github.com/mozilla/DeepSpeech) that wouldn't be too hard to put behind an API. We also plan to open source more of our stack in the future so you can run it yourself.

Privacy aside, the biggest problem with self hosting these kinds of models in our opinion is the compute required. The really accurate models are so large, they require GPUs for inference. You could run the models on fast CPUs if you don't care too much about latency, but the throughput would be pretty low. Either way, GPUs and fast CPUs get expensive fast, so our hope is that by us specializing in hosting these models, we can offer you a price point that would be cheaper than if you were to try to host it yourself.

dylanbfox··on AssemblyAI: speech-to-text API
Dylan from AssemblyAI here.

Right now we only support English, but are going to be launching more languages very soon.

We are also getting closer to launching a transfer learning API, so you can train your own acoustic models. This _could_ work for new languages if you had enough data, but we haven't done much testing around this yet. It's something that is definitely on our to do list!

dylanbfox··on AssemblyAI: speech-to-text API
Thank you for the feedback!
dylanbfox··on AssemblyAI: speech-to-text API
Right, we noticed similar findings. We do automatic punctuation now, and do diarization when there is more than one channel in the audio file. We're launching diarization on single channel audio with multiple speakers very soon. We're currently focused on improving some of our customization features, and then we plan to ship single channel diarization.

Thanks for sharing your results! We have more samples here if you want to do more comparisons: https://blog.assemblyai.com/2018/08/09/cutting-edge-phone-ca...

dylanbfox··on AssemblyAI: speech-to-text API
Hi there! Thanks for the support. We were very surprised to see ourselves on HN tonight!

We've been working towards a big launch in a few weeks, which is why we don't have the benchmarks ready just yet. But we are comparable to most of the big guys (IBM, Google, etc.) out of the box. And if you customize the API for your specific use case, you can get a lot better accuracy.

We do have a real-time endpoint but it's not production ready yet. Our primary use case to date has been async transcription for phone call recordings and podcast recordings, but we are definitely working on making the real-time endpoint production ready.

Real-time is just a little bit trickier, because it's more expensive to run since we deploy our models onto GPUs.

dylanbfox··on AssemblyAI: speech-to-text API
Dylan from AssemblyAI. We're working on this! We didn't plan to launch on HN tonight which is why the benchmarks aren't ready, but we know we definitely need these for the community.

We've found most public benchmarks, especially Libri, are not that representative of real world data we see in production. Most real world data we see is a lot noisier, and has worse recording quality like low bitrates and compression from mp3 encoding.

We do worse than state of the art benchmarks on Libri Clean today, for example (I think we are around 7% WER last time I checked), but are much more accurate on real world data than models reporting 3-5% WER on Libri. This is why we want to make sure we are thorough when we report our benchmarks on popular datasets like Libri.

dylanbfox··on AssemblyAI: speech-to-text API
This is Dylan from AssemblyAI. We're really surprised to see ourselves on HN tonight!

We had a big launch planned for 4-6 weeks from now, and have been working towards getting things ready for that. As a result, we're missing a lot of things we know we need like benchmarks comparing ourselves to other services. Please bear with us!

If you do end up trying the API, we'd love your feedback. We're trying to build a really simple speech-to-text API for developers, that you can get up and running with in just a few minutes, and that doesn't require you to implement <insert big tech co here> into more of your stack.

There's a lot more we offer too, like:

- customization via transfer learning for higher accuracy (language models now, acoustic models soon)

- supporting lossy audio like low bitrate mp3 files from phone calls

- transcribe any audio file without having to specify it's metadata like sampling rate or encoding

We're also constantly improving our models for higher accuracy. Every few weeks we ship accuracy improvements based on improvements to our DNN architectures, better data, better data augmentation, etc.

In terms of our STT stack, we're using CTC based models combined with RNN-LMs -- all built in TensorFlow and PyTorch -- for decoding. Happy to provide more info around our stack if you're interested!

For any questions off HN -- you can reach me at dylan at assemblyai dot com

Thank you!!

dylanbfox··on AssemblyAI: speech-to-text API
Dylan from AssemblyAI here.

We're surprised to see ourselves on HN tonight -- but if you do try out the API I would love to see what you think! Thanks for your interest.

DeepSpeech is an awesome open source project, and we absolutely support open source speech-to-text. We're going to be open sourcing parts of our stack in the future as well.

We're a lot more accurate on real-world data (like phone calls, podcasts, accents, noise, etc.) than the current DeepSpeech model. We're actually less accurate on LibriSpeech Clean than what DeepSpeech reports, but we've found Libri Clean isn't very representative to real-world data and as a result isn't a great benchmark.

We're planning to put out more thorough benchmarks comparing our API to other services (including DeepSpeech, Kaldi, and CMU Sphinx) in about 1-2 weeks. If you want to email me at dylan[at]assemblyai.com I can send you the benchmarks once we have them.

Another thing that's hard to do is host RNN based models like DeepSpeech in production at scale. It gets expensive fast, since you need to do inference on GPUs to keep your latency time down. We spend a lot of time optimizing our infrastructure to keep costs down, so there's a good chance we can host RNN based models cheaper than if you were to do it on your own since we specialize in this.

dylanbfox··on AssemblyAI: speech-to-text API
Dylan from AssemblyAI here. Thanks for trying the API!

Time to implementation is one thing we’re focusing a lot on. We’re really trying to make it fast to get up and running with Assembly — for example not requiring you to specify any meta info about your audio (like sample rate, bitrate, etc).

We do have a real time endpoint but it’s not production ready yet — our primary use case right now is for phone call recording transcription.

If you want to try the real time endpoint, you can email me at dylan at assemblyai dot com and I would love to get your feedback!

Page 1 of 3Next →