The post got pretty popular on HN and I had 5 key questions/predictions, feel free to read all the details (and lots of quotes I transcribed straight from Gruber's video) but now I'll summarize my thoughts quickly below:
1. Very few languages for Siri (where's Spanish?)
2. No API announced for developers to add tasks to Siri
3. It's still named...Siri? What happened to Assistant? I guess it is quicker to say in practice...
4. Siri's in BETA? Is this the first time Apple's released a major iPhone feature with a beta sticker?
5. No payments integration with Siri mentioned. Can it buy stuff for me? as Gruber talked about in 2008?
6. No Facebook partnership for social knowledge on Siri, or even iPad app.
There's also no reason that Safari can't upload photos out of the browser...except the fact that it would allow devs to write web-based apps that use the camera, meaning less mindshare going to iOS dev. Right now (well, as of about a month ago, the last time I looked), that's impossible.
Google has taken a different approach, where your voice sample is uploaded to a Google server, processed, and downloaded back to the device. This takes less CPU power but is also far less accurate, as the voice sample must be very low quality to have a quick response time from Google's server.
Apple/Siri are taking the approach that high quality voice recognition must be done on device in order to provide the level of performance and accuracy that voice recognition requires. I think we will find that Siri actually works and doesn't have as many errors as Google's voicemail transcription.
This is the reason for requiring iPhone 4S.
Voice compresses nicely. Turns out we humans aren't capable of making such varied sound that it can't compress. Our mouth holes are ancient technology.
In practice, I'm in love with Google's voice capabilities. It seems to understand context. Its crazy how accurate it is. I often tease my iphone friends with it. I'm also highly skeptical that an application on a phone can outdo google's massive libraries and server infrastructure. If anything, I'd expect the Apple voice to be worse. Regardless, I can't wait to see this stuff in action. A war for the best voice recognition would be great right now as its been a patent blocked and ignored field for the most part.
The rumor is that Apple is sending the audio to Nuance servers, i.e., they're doing cloud-based speech recognition.
I tried Google Voice Actions on my Nexus One quite a while ago, but it was optimised for the US market. The accuracy for me was so bad that I didn't bother with it.
Then recently, a Google blog on RSS said the latest app had been optimised for my locale. Now, of course, it's spookily accurate.
You can't beat your algorithms in the cloud being bombarded with sample data round the clock.
This takes less CPU power but is also far less accurate,
as the voice sample must be very low quality to have a
quick response time from Google's server.
Has there been a comparison of the accuracy of the two services? This claim seems unsupported.By focusing exclusively on the newest model, they can maximize the power of the feature.
EDIT : I believe > 1 src is needed for fighting noise so you sample the same audio from 2 locations and can distinguish between multiple audio sources in the input signal. This seemed to be one of the selling points of the mic array in the kinect
Particularly as they did just advertise that they added a custom ISP for the camera. I'd be surprised if they wouldn't go to similar lengths for the voice assistant.