HNHacker News
TopNewBestAskShowJobs

kenarsa

448 karma · joined March 26, 2018

I like deep learning, machine learning, and scientific computing.

https://picovoice.ai/

submissionscomments
kenarsa··on Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
[flagged]
kenarsa··on Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
Try https://github.com/Picovoice/orca
kenarsa··on [dead]
GitHub: https://github.com/Picovoice/picollm Platform: https://picovoice.ai/blog/picollm-local-llm-platform/ LLM Quantization Algorithm: https://picovoice.ai/blog/picollm-towards-optimal-llm-quanti... LLM Inference Engine: https://picovoice.ai/blog/picollm-inference-engine-for-x-bit...
kenarsa··on The importance of initialization and momentum in deep learning [pdf]
Reflecting on Ilya Sutskever's early contribution to DL. This paper saved my job a decade ago
kenarsa··on Why is WebAssembly not supporting 256/512 SIMD registers?
100% of AVX/AVX2 have 256 SIMD registers, and Arm also has non-NEON registers, which are 256/512. of course, this requires runtime detection, but that is ok because even for native code, people need to do that.
kenarsa··on How Many Languages a Developer Should Know?
"Know" is a vague word. Everyone should know one language inside out as if they can code within a Notepad without IDE support. I think if you know three languages really well, that puts you in the top 1% of devs. I would go for C, Python, and JS.
kenarsa··on OpenAI Launches Voice Assistant Inspired by Hollywood Vision of AI
It's interesting to see the cycle. Alexa said everyone wanted to toggle the lights and get the weather forecast. Hence, they folded. Now, we are back with "stronger" tech and want people to change their behavior.
kenarsa··on Offline Voice Assistant on a Microcontroller with 192KB RAM
Have you noted that the board has no connectivity chip? If I had a way to connect to internet without the required chip I had a better story to tell. You snippet of FAQ is correct for all other platforms we support aside from microcontrollers ...

[1] https://www.st.com/en/evaluation-tools/stm32f4discovery.html

kenarsa··on Offline Voice Assistant on a Microcontroller with 192KB RAM
Picovoice runs on almost anything: web browsers, mobile, desktop, single board computers, and microcontrollers. For the platforms that have connectivity (i.e. almost anything aside from microcontrollers), we do call home for license management. This helps us keep the `Free Tier` free for personal users, hackers , and skunkworks projects, but make sure we get paid by enterprise customers with deployments at scale [1]. On a microcontroller like the one in this tutorial, there is NO connectivity option. Hence, in this specific case it is 100% offline with no license management. In other cases voice recognition is 100% offline but the call home for license management needs connectivity.

[1] https://picovoice.ai/pricing/

kenarsa··on Offline Voice Assistant on a Microcontroller with 192KB RAM
Yes, the board does not have a connectivity chip of any sort [1]. Even if one really wants to there is no way to connect to anything from this board.

[1] https://www.st.com/en/evaluation-tools/stm32f4discovery.html

kenarsa··on Offline Voice Assistant on a Microcontroller with 192KB RAM
context, context, and context!! Then deep learning tricks and efficient implementation.

[1] https://picovoice.ai/blog/end-to-end-intent-inference-from-s...

kenarsa··on Offline Voice Assistant on a Microcontroller with 192KB RAM
would love to check out what you build with it :) I've been constantly and pleasantly surprised but what people can build given this tech
kenarsa··on Offline Voice Assistant on a Microcontroller with 192KB RAM
Yeah, I post about my company that I founded and I am super into it. Which part of this is MISLEADING? The fact that I care or are you saying Picovoice's tech doesn't work? I made the latter easy cause you can now go and try it without me in your way. You comment is misleading.
kenarsa··on Offline Voice Assistant on a Microcontroller with 192KB RAM
I'm sure smart people can find a way to hack this! But check my other comment, the chip does NOT have connectivity module. Don't take it from me. Read the ST's product spec.

[1] https://www.st.com/en/evaluation-tools/stm32f4discovery.html

kenarsa··on Offline Voice Assistant on a Microcontroller with 192KB RAM
have you noted that there is `Free Tier` that cost you $0? You can train using that. For this tutorial the cost of board ($20) is all you need to pay. Same for personal projects and even small skunkworks projects within companies. Picovoice makes money from large-scale deployments done by device makers

[1] https://picovoice.ai/pricing/

kenarsa··on Offline Voice Assistant on a Microcontroller with 192KB RAM
That board doesn't have connectivity! If we had a way to connect to the internet without a connectivity chip I would have had a more exciting post!
kenarsa··on DeepSpeech 60x Smaller, 9x faster, and 2x accuracy
The benchmark is using the latest stable of DS >>> https://github.com/Picovoice/speech-to-text-benchmark/blob/m...

the data for cloud-based is from 2022.

kenarsa··on DeepSpeech 60x Smaller, 9x faster, and 2x accuracy
mycroft already integrated our wake word (porcupine). we remain neutral and anyone can use our tech :)
kenarsa··on DeepSpeech 60x Smaller, 9x faster, and 2x accuracy
Clarification. Google STT has to services. standard and enhanced. we are just better than enhanced and much better than standard. enhanced is much pricier than standard if you wonder what the diff is.

It runs real-time on NVIDIA Jetson Nano and RPI 3/4.

If you think we should consider other embedded platforms we love to hear what and why

kenarsa··on DeepSpeech 60x Smaller, 9x faster, and 2x accuracy
Thank you for the clarification. `[1]` and `[2]` hyperlinks in the intro section point the benchmark
kenarsa··on Offline voice AI within 512 KB of RAM [video]
This is a really good point. That is why we partially open-sourced our technology to enable unbiased third party evaluation. You can run the exact same demo on a Linux box or Raspberry Pi (any variant) using what's available on the project GitHub repository here: https://github.com/Picovoice/rhino

We are in process of open-sourcing a statistically-significant benchmark for this tech. But this will happen in 2019.

kenarsa··on Offline voice AI within 512 KB of RAM [video]
The device you are referring to is quite different. I am taking this is the board you are using?

https://www.arrow.com/en/reference-designs/imx6slevk-imx-6so...

It was an ARM Cortex-a9 with NEON extension instead of ARM Cortex-M7. It is basically a different family of i MX processors.

kenarsa··on Offline voice AI within 512 KB of RAM [video]
Yes, we will publish an article about our speech-to-intent engine and add a link to it on our website. This should happen before the new year.

You can find some information about the wake-word engine here https://medium.com/@alirezakenarsarianhari/yet-another-wake-...

kenarsa··on Offline voice AI within 512 KB of RAM [video]
Thank you. The engine is definitely domain-specific. This would apply to vocabulary and also inference. For example, if you want to use the tech for smart lighting then you would need a different model/context.

The demo is done with some noise and there is some reverberation as well. The speakers are also somewhat accented. That being said we would like to open-source a benchmark for this (similar to other products we have). The comment on accuracy is a bit tricky as it would depend on parameters you mentioned and specific task. I will provide more information when we open source the benchmark in Q1 2019.

kenarsa··on Offline voice AI within 512 KB of RAM [video]
Thanks a lot for the link. I'll be sure to take a look into it in more detail.

Keyword spotting is one of the modules we run in this demo. That's how we detect "Hey Barista". We also run an engine we can "Speech-to-Intent" that infers user request from follow up command.

One thing I wanted to mention is that there are two challenges when running DNNs on embedded platforms (1) limited compute power (CPU) (2) limited memory (RAM). RPi zero is definitely bound by (1) but not (2) as you get 100 MBs of RAM on it.

kenarsa··on Offline voice AI within 512 KB of RAM [video]
Pruning is a great idea to reduce memory usage. One thing to be careful with is that pruned matrices use irregular memory access and they might be slower as we don't have SIMD support for sparse matrix multiplication on generic CPU's (e.g. ARM) yet.
kenarsa··on Offline voice AI within 512 KB of RAM [video]
Hello. This is Alireza. I am the founder of Picovoice.

I totally understand the need to support makers community. We do have GitHub repositories for engines demoed here which allows you to use these technologies to some extent (not the full set of capabilities). I am working with our partners (both Soc and distribution) to come up with a maker-specific product for evaluation and personal use. It most probably will be a HW/SW product (i.e. a board that comes with our software). The product should allow you to use the full set of features on that specific board. I am expecting this to happen in 2019 and I will disclose the information as I am figuring things out.

kenarsa··on Offline voice AI within 512 KB of RAM [video]
You are absolutely right. Compressing (using this for lack of better terminology) is extremely important for power-efficient applications as one of the main power draws on a device is external RAM.

We had to come up with a bunch of ideas on how to fit our stack into the on-chip RAM (512 KB) and leave enough for OS and the actual application.

kenarsa··on Picovoice – Embed private voice AI into any product instantly
two reasons:

1- the business model 2- in some cases, it actually needs some engineering. for example a new brand name, etc.

kenarsa··on Picovoice – Embed private voice AI into any product instantly
I suggest doing a quick read on one of the demos (Python for example) that should clear things up. If not you can always open an issue...
Page 1 of 3Next →