Show HN: Wit – Natural language for your app
wit.ai
wit.ai
1. It doesn’t do any speech recognition (speech -> text), so not sure why they put Siri in the title. It is also not clear how they can ‘hijack’ the text from Siri to do this analysis. The ASR engines they talk about (CMU, OpenEars) have pretty horrible accuracy (compared to Siri or google voice).
2. Looks like they do some form of text normalization/correction, again not clear how they do it.
3. The actual service they provide is a form of named entity recognition (confusing named intent which clashes with the android intent mechanism in their examples).
4. Also they let you define your own entities to match. You can train them using a drop –down menu. Not sure how you can train hundreds of examples using point and click.
This different from alchemy (or many others) because this is open source(?) http://www.alchemyapi.com/products/features/entity-extractio...
Given this service was for developers with an interest in NLP, it would have been good if they didn’t hide behind a snow job title like “Siri as a service”.
Currently most Wit users use Google or Nuance with great success. You can even use Android's offline speech rec.
That being said, CMU and OpenEars work well, as long as you provide them with good language models (which you can't do if you hack a quick project). Our plan is for Wit to automatically generate the right language models from your instance configuration.
> 2. Looks like they do some form of text normalization/correction, again not clear how they do it. 3. The actual service they provide is a form of named entity recognition (confusing named intent which clashes with the android intent mechanism in their examples).
We abstract the full NLP stack for the developer. How we do it is not really what matters to our developers, as long as it works :) Actually we use a combination of many different NLP and machine learning techniques.
> 4. Also they let you define your own entities to match. You can train them using a drop –down menu. Not sure how you can train hundreds of examples using point and click.
You don't need to train hundreds of examples. Plus, our users are not NLP/ML experts and they prefer a graphical UI. But that's true it could be still more efficient, we have good features in the roadmap for that :)
> This different from alchemy (or many others) because this is open source(?) http://www.alchemyapi.com/products/features/entity-extractio....
Alchemy is great as a set of NLP tools, some of them quite academical, but it's not designed from scratch to solve the problem we're trying to solve: enable the masses of developers to easily add a natural language interface to their app.
How you do it is most certainly what matters to developers, as soon as it doesn't work as expected :)
Theory is when you know everything but nothing works.
Practice is everything works but no one knows why.
Here, theory and practice are combined: nothing works and no one knows why
:)
I'm running this at home and it works great for adding custom actions to Siri
But to give the authors their credit back, it's not what I guessed. It's much more a GUI around a complex toolset that would require you to dig deep into mudwater, bad docs etc. So, yes it makes life easier and sense to use this in your app. I have not evaluated the quality of their service yet, but it's a Startup, it's not going to stop improving (hopefully) :)
(Google/Apple are essentially powered by "Nuance", but with different qualities of training-data.)
More NLP/AI Startups and more colloboration on HN please!
2cents: I hope people don't sell to their first working protoype to the Google/Apple/Microsoft Empire, but try to get big on with friendly startup-colloboration and with the help of investors/angels.
I'll be honest, HN has a tendency (myself included) to have a first natural reaction of "how can I criticize this?" But just because something isn't faster than enterprise, or not-as-scalable, or not made in your framework of choice doesn't mean it's worthless. I think this project is amazing. Great job and I can't wait to see this mature.
I think there was a picture of a robot on the screen for a few seconds, but that's all I remember.
Would disabling javascript do the trick?
EDIT: All animations (except "What we do") should be disabled. Please, email me at willy@wit.ai if you still have issues.
I imagine as the developer you don't notice it. But as somebody trying to read a page, it's really jarring to have that happen. Enough so that I give up trying because I just want it to stop doing that to my eyes.
Any chance you could turn it off completely and just put some arrow icons on there?
You should make it clearer that you don't actually handle voice recognition. When I read: "Developers use Wit to easily build a voice interface for their app." I expect you to handle things from start to finish.
Also, let me try it! It's frustrating because the UI looks like you can experiment but it's only an animated demo (or am I missing something??) In particular the mic logo is used to record on Google and here it doesn't seem to do anything?
You're right, we'll make it more clear on the landing page. A full out-of-the-box integration with some voice recognition engines (we love CMU Sphinx, open source) is in our roadmap.
> Also, let me try it!
We purposely didn't provide a "end-user" demo (something that would look like chatting with Siri) because we want to focus first on the developer experience, when they configure Wit to understand their very own end-users intents. You can require an invite and try this in less than 5 minutes.
Fair enough, but then you should make it clear: "Want to try it out? require an invite and try this in less than 5 minutes!"
You usually have to wait several days when you apply for a beta like this.
Seeing this message, I bit the bullet and requested an invite anyway, and have seen no action in the couple of hours since... thus validating my initial reluctance.
But as a bootstrapped startup, we have to make tradeoffs as our budget is limited. We have to accept invitations gradually today to keep our servers alive. We should be able to accept everybody within a few days at most. Sorry for the inconvenience.
On this topic, I invite people to try out my non-prototype, non-project toy that uses Google's support for HTML5 speech recognition. It's pretty funny how wrong things go when you try to say something even a bit out of the ordinary:
http://arachnoid.com/speech_to_text
If I say, "Now is the time for all good men to come to the aid of their country," an old teletype test sentence, the Google recognizer always nails it. If I say, "I hit an uncharted rock and my boat is being repaired," things go hilariously wrong, and every time differently.
For instance: "In most cases, before beginning to listen, the browser will ask permission to monitor your microphone."
Came out as: "In Las Cruces, f listen, permission to monitor your microphone."
Just a heads up, but Get Started on the pricing page does nothing. It's natural progression for me to go home page->pricing->OK, looks good, let's get started.
Wit takes the output of the voice recognition engine as input. It's quite robust to voice recognition errors. Most devs use Google's engine or the open source CMU Sphinx engine.
Fixed the Get Started link, thanks!
https://github.com/dpaola2/jarvis (work in progress)
I absolutely would love a better NLP api. Please let me in!
You should be able to sign up now.
Bringing Natural Language Understanding to the masses of developers is hard and we still have a lot of work ahead of us. Please don't hesitate to reach out to us!
Here is a tutorial for quick Android integration: https://wit.ai/docs/android-tutorial
Only messing, it was taking a lot of CPU though.
Working on a fix now.
I don't know if people will remember it and be receptive to this touch but I like it.
Can you email me your GitHub username at willy@wit.ai to make sure we got your request? Thanks!
Wit is 100% open and flexible, you can create any intent you need for your app, you're not limited to a static set of domains/actions.
EDIT: @ragebol we are very interested in ROS and robotics, don't hesitate to get in touch with me arthur at wit dot ai. In the future we would like to provide an off-the-shelf human/robot communication module for developers.
With Maluuba, we can't make a command like "Introduce yourself" or "Grab that can for me" because of the limited set of categories. Wit should be able to handle those as well, from the looks of it.
I applied for alpha access using the github username "marks"
Openness is one of our core values. We're inspired by companies like GitHub.
- Regarding data, we encourage users to share their data, making “public” Wit instances free of charge (à la GitHub). We’re also working on standard formats for NLP/ML data (models, sets, etc).
- Regarding open-source, we plan to release our algorithms and infrastructure piecemeal (à la Prismatic). We’ll announce our open source plans in the near future.
Are there any test cases I can use, which utilize the full power of the API?
It seems like WIT will take the text that has already been translated from a user's voice to text and make it easily accessible to my application but how does WIT access the text generated from a Siri request in the first place for example? Does WIT have some other way of getting at this data that has already been converted from voice to text by Siri or Google or some other speech-to-text engine?
Actually there are ways. On Android devices, voice rec is available to devs (even offline if the user enabled it!). We have a simple tutorial about how to integrate on Android https://wit.ai/docs/android-tutorial
Right now on iOS you have two options (none of them involves Siri, which is kept closed by Apple):
1/ Do the voice rec server-side (Siri does that)
2/ Use OpenEars to do it client-side
Server-side, you have many voice rec options, including open source CMU Sphinx.
Providing a fully-integrated solution with speech rec out of the box is in our roadmap.
Except Stremor has a Query Language so you don't have to do anywhere near as much heavy lifting.
Looks like you focus on search, summary, entity and sentiment extraction with a rule-based approach.
Wit's focus is to power human/machine interfaces, and our priority is to provide developers with a 100% configurable solution, with no prior assumption on their domain. And we don't believe in rules, we chose a machine learning approach.
Unlike Wit it also offers the option to use the API's that are already integrated or Bring Your Own Backend so that you can have a mix of info/responses from your own system, or leverage what is already there.
Meanwhile you can decide to share your configuration data and get Wit for free (à la Github) :)
I would be weary of using the Github Octocat mascot though. I believe Octocat is protected under copyright.
After reading http://octodex.github.com/faq.html, we thought it was okay to put this image given that we advertise and reference GitHub a lot, we heavily integrate with it and we love Octocat!
Do you think we should remove it?
Thanks!
I don't know if Ask Ziggy is 100% self-service for the developers. That's a key requirement for us.
- Spanish (Mexican, Castellano, others?)
- Chinese (Mandarin, Cantonese)
- Hindi
These would be logical next steps with some important commonalities: broad base of native speakers, high importance in the US market (maybe less so for Hindi), and very important dialect differences. Mandarin, Cantonese, Hindi, and Russian could also force the issue of non-Romanized character sets.Having it online for configuration has at least two advantages:
1/ Easier to start (1 minute and you're playing with it)
2/ Leverage existing, "live" data to build your configuration
We'll release a Ruby tutorial soon!