New services expand IBM Watson capabilities to images, speech, and more
developer.ibm.com
developer.ibm.com
All the Watson services are still in beta but will start going GA very soon (first one next month). If you have any questions, please fire up, the Watson team is ready to answer.
> If you have any questions, please fire up, the Watson team is ready to answer.
So that's what you built Watson for :-)
Watson Jeopardy itself is built on top of Apache open source stack (Apache UIMA and Hadoop): http://en.wikipedia.org/wiki/UIMA
[1]: http://www.ibm.com/smarterplanet/us/en/ibmwatson/developercl...
I just want to access some services via API from my own servers. I think the documentation is not that good, there should be curl examples at least. For instance, for the STT or TTS include some curl examples.
Does the STT have speaker identification or does it output text in one stream?
I tried to access: https://gateway-s.watsonplatform.net:8443/speech-to-text-bet...
I used my bluemix l/p. It did not work. Are there other api credentials that are needed?
is in the watson-developer-cloud organization in github: https://github.com/watson-developer-cloud/speech-to-text-nod...
In fact, all the samples have the code there.
For example, I searched Google for "photo of girl", and found this image which seems very easy:
http://www.wagggsworld.org/shared/uploads/img/rachel-s-p-pho...
Watson says:
Color 71%
Human 67%
Photo 65%
Dog 59%
Person 57%
Placental_Mammal 56%
Animal 50%
Long_Jump 50%
Huh?This isn't me cherry picking bad results; aside from their demos I'm not finding any photos that are accurately classified. I even tried a headshot of a person isolated on a white background, and Watson told me I uploaded a photo of "shoes".
Seriously - how is this data useful? What could I build with this level of accuracy?
Watson team - do you agree? Is this product about to get a lot better, soon, or is this considered "pretty good"?
We are also believe that the first applications (e.g., classifying animals or plants or landmarks in dedicated apps) will have narrower use case that give better accuracy.
Also, the other results are very wrong. (i.e., Watson is more confident that this is a dog than a person. And I have no idea where it got "Long Jump" from). This makes it hard for me to trust Watson.
Is the recommendation that I incorporate a "confidence in Watson" metric, and ignore most of the results?
What confidence from Watson would you say indicates an answer that is probably accurate? And how confident are you that Watson's self-reported confidence is accurate?
Watson is a training API rather than say the more fanciful emergent AI type API. More data, the better it gets. It is like Google's voice recognition isn't good because someone coded the magic constants for various accents, rather it is good because Google fed it millions of samples of spoken words and corrects it when they get it wrong.
I'm not in this field, so I'm having trouble understanding what use cases / consumer facing features this API unlocks. Your comment is very helpful in that regard.
One example of a use is at Kiva we require borrowers to have a picture of themselves posted for their loan. But sometimes we get pictures of things like goats or cows instead (those are kind of nice to but gotta follow policy). Currently this is something we have to manually review for, but if we could automate that review piece it would save a lot of time (especially if at some point it could also count the number of humans in a photo).
Or Watson is not confident that humans are placental mammals and placental mammals are animals?
Watson gives high confidence to it being a color photo of a human (which is a Person, and an Animal). Which is right. But the only part that a human would ever really care about is that there's another human in the picture.
It gets things wrong with a reasonable confidence for Dog, Placental_Mammal and Long_Jump...importantly, these are wrong in ways that humans would never get wrong.
Just as important are the omissions. A human would probably describe this as a picture of a girl or young woman, laughing or smiling, with curly brown hair wearing a scarf -- and maybe some other incidental information.
Of that description, Watson only got the superclass of one part correct (Human, Person) and didn't provide any of the other parts.
AI fundamentally "thinks" differently than a human, and that makes it hard for humans to use AI as a cognitive enhancement tool in the same way humans use calculators, books, writing, etc. We don't trust what an AI is doing or the answers it provides because for the information it provides, AIs tend to provide right-and-irrelevant, weirdly wrong, or omits obvious and necessary information that a human might use for informational purposes.
If humans ever encounter aliens, it's likely that their mode of thinking will be just as different. So bridging that gap, and figuring out how to make AI like this useful could be a useful endeavor.
portrait
youth
fashion
facial expression
women
european
girl
model
female
actress
Photo 75% Shoes 69% Nature_Scene 69% Meat_Eater 63% Object 63% Mammal 63% Vertebrate 63% Cat 63% Indoors 62% Room 60% Person 58% Color 57% Judo 54% Person_View 53% Human 51% Leisure_Activity 50%
If you give the classifier a hint (animal) it gives: Meat_Eater 63% Mammal 63% Vertebrate 63% Cat 63%
So, clearly needs work as a general classifier, but still potentially useful.
Watson http://text-to-speech-demo.mybluemix.net/
Nuance http://www.nuance.com/for-business/text-to-speech/vocalizer/...
I prefer the Watson version voicing a sample paragraph. Both are good enough for an application that selects on price. For a voice-first application, maybe Watson is better for TTS.
For speech to text, Nuance has been the leader, e.g. Apple's Siri. Has anyone compared IBM speech recognition to Nuance, Microsoft & Google?
But - sometime in 2014, and I can't really place it - but right around June/August, Siri all of a sudden turned a corner, and her dictation ability got markedly better - so much now, that I don't even bother typing into my iPhone if I'm in a place where I can talk to it - dictation is 99% flawless. much better than my typing, and unquestionably faster.
For whatever reason, Apple hasn't been making a big deal of this - perhaps because they don't want to admit how crappy it was before - but it really is a big deal. Siri is, 3 years later, what she should have been in 20111.
Can't wait to see what the next step in this evolution will be...
The idea was that you'll also host your application in bluemix, although I think the services are actually accessible elsewhere once you create the instance in bluemix.
Vocalware https://www.vocalware.com/index/demo CereProc https://www.cereproc.com/
It is getting increasingly difficult to pick one as the clear leader for "natural sounding". The results are good enough for voicing canned text, and certainly better enunciated than many thick-accented English speakers. Improvements through training can still be made in parsing the text.
For example, IBM Watson interprets "IT" as "it", in the following sentence.
Thank you for calling the IT department.
Vocalware and CereProc correctly parse that.
Who I would really like to hear opinions from are professional voice actors, though they would tend to be understandably leery to lend a hand to improve TTS. Is there a standardized form of writing text that communicates the kind of emphasis, placement of silence and warping of phonemes these actors use in their delivery to concisely convey emotion, that TTS products can adopt?
Edit: And to add insult to injury, the English voices do pronounce "Español" correctly!
When this was first announced I remember reading about their pricing model where they would take a percentage of app revenue. I'm glad to see they offer flat pay-as-you-go pricing now. Some of the Watson services are intriguing.
For visual recognition, I used a picture of a snowmobile from http://www.1888goodwin.com/2013/11/14/what-do-you-need-to-do..., which it identified with 73% confidence as "Invertebrate".
Speech to text is a parody twitter account waiting to happen. Here's me asking it how it does with technical transcription:
How do you doing technical words.
If you were going to have to talk about get an jute cushion pull.
And you wanted to discuss the impact on a file server memory.
Issues that cross processes talk about home forks rivers slowed difficult.
It's not possible to train their service with your data, unlike wit.ai for example. Seems obvious to me that people would want to train with their own data.
I decided to test it a little. I copied phonem challenges and non-sensical phrasing from the web. Then I added some stuff that I know has problems from past experience.
----- Let's explore some complicated conversions, shall we? The old corn cost the blood. The wrong shot led the farm. The short arm sent the cow. How can I intimate this to my most intimate friend? Don't desert me here in the desert!. They were too close to the door to close it. The buck does funny things when does are present. Today is 1/1/2015. Today is Jan 5th, 1992. It's currently half past 12. Or 12:30PM. Twenty thousand dollars. 20,000 dollars. 20 thousand dollars. 2^5 = 32. NASA is an acronym. This ... is a pause. EmailAddress@somedomain.com.
1. No "special characters" allowed in passwords when creating an account. 2. ...where's the REST API? I've "added a service" (TTS), but I have to write a webapp to expose it over HTTP? It sure is a different experience than your typical API documentation.
As far as I can tell, most of the documentation essentially begins, "First, deploy a web app on our platform". Which is fine I guess, but isn't nearly as simple as the HTTP APIs you see from many other recent SaaS providers. As least for me, I'm pretty unlikely to jump through those hoops. Maybe others will be different.
Edit: All the way down on the bottom of the documentation page, past the research references, there's a link to HTTP API documentation--literally the last link on the page.
So basically, IBM is charging us to provide it with training data to make Watson useful for practical applications. Makes sense, but I can't help but feel that it would be a smarter move to skip charging entirely for now, or to use drastically reduced pricing tiers that exist only for the purpose of preventing abuse. The idea of releasing a product like this with less than impressive demos is a bit of a risk. It's not going to encourage people to use it if the demos aren't compelling, and the demos won't be compelling until a lot of people are using it. I'd err on the side of optimism here, it'll probably work out for the best, but it will be interesting to see how this goes and provide a good case study.
My other thought is that if IBM can't get sufficient training data on their own, what hope do the rest of us have? Performing classification on arbitrary data is a herculean task. People could throw literally anything at this api and will expect to get common sense results, it's nearly impossible and pushing the boundaries of what even cutting edge software can do. But if a company like IBM spends billions of dollars and their demos still end up generating mostly confusion and complaints... This kind of open ended "AI" might be more difficult than even the most conservative experts thought.
EDIT: As an after thought, the real value here isn't so much software as it is pooled training data. Facebook has been able to identify human faces in photos for years, speech-to-text and concept modelling have all been around for a long time. What's difficult is getting the labelled data necessary to distinguish between "is this a picture of a person or a picture of a cat?". Watson is great and it seems like IBM has made an investment in acquiring and collecting the data necessary to do that. But their big play here might be to build a consumer friendly enough product that their users contribute the rest of that data for them over the next several years, building an aggregate data set that is worth as much or more than the software itself. Again, will be interesting to see how it plays out.
We wanted to get the services into peoples hands early, even though we're still working on them, rather than wait until we had a perfect product. There's a tradeoff here, but we figure that we can improve the services faster and better with public usage and feedback than we could in private isolation.
Since they're free, hopefully people will be able to have some fun playing around with the services, also!
Are they using more than ImageNet? The ImageNet dataset(s) are not hard to get.
>>Speech to Text : This application only works in recent versions of Chrome supporting HTML5 audio capture
http://en.wikipedia.org/wiki/List_of_top_United_States_paten...