Google Duplex: An AI System for Accomplishing Real World Tasks Over the Phone
ai.googleblog.com
ai.googleblog.com
People who answer phones to take bookings perform an extremely limited set of questions and responses, that’s why they can even be replaced by dumb voice response systems in many cases.
In these cases, the human being answering the phone is themselves acting like a bot following a repetitive script.
Duplex seems trained against this corpus. The end game would be for the business to run something like duplex on the other side, and you’d have duplex talking to duplex.
Most people working in hair salons or restaurants are very busy with customers and don’t want to handle these calls, so I think the reverse of this duplex system, a more natural voice booking system for small businesses would help the immensely free up their workers to focus on customers.
Right, and this kind of comments will continue for a while.
The question is - when the "this is trivial, move on" type of comments will start to fade out? Five years? Ten?
>The end game would be for the business to run something like duplex on the other side, and you’d have duplex talking to duplex.
The end game is clearly to use an api, not this.
This system can’t pass the Turing test, it would be fooled probably by a simple question about itself or a subject outside the domain, like the kind of food you like.
You’ve got people in this thread hyperventilating about AI duping your voice and the. Becoming a doppelgänger and therefore we need laws immediately to stop this dystopia? Let’s calm your cortisol levels for a second and stop acting like Thanos just got the last gem.
My understanding is API integration is what wechat is in china -- every hair salon and equivalent-of-corner-pizza shop has some wechat integration, payment and all.
Voice bots like this will have the advantage of ubiquity. At least a couple years ago before every resteraunt had 5 tablets for all their seamless/grubhub/chowhound/whatever apps, pretty much the only reason the fax machine was still around was for restaurant ordering. Although there were clearly better ways of doing it (see how Dominos reinvented itself as a tech company), the sheer ubiquity of fax as the lowest common denominator kept the tech around.
In that light, it's kinda like the cell-phones-leapfrogging-landlines-in-developing-countries argument... part of the wechat story involves a massive population entering the consumer class at a time when everything was digital. Call me out if this is a gross over-generalization, but in a way, the wechat population never had to deal with the backwards-compatibility of people growing up ordering a pizza over the phone.
It'll be interesting to see how the API-centric approach (wechat) plays out versus the lowest-common-denominator ubiquity approach (voicebots). I'd stop short of calling API's the end game though.
The fax machine is still around now, and heavily used in the medical context: https://www.vox.com/health-care/2017/10/30/16228054/american...
Meanwhile in NYC, good luck getting the bodega on the corner to even take your debit card.
If they don't accept cards, they almost always have an ATM.
My complaint in NYC is the uptick of "cashless" places, that don't accept legal US tender. I like using cash, I don't want it to go away.
This voice based model can be integrated into any existing system. It already has the network effect going for it and it's not tied to the fate of any one company
Anyone can have a conversation, but not everyone can author an API.
Google actually delivers a real world product incorporating the most advanced AI we've had the chance to experience, and half the HackerNews comments are "Wow, this is so dumb, can't wait for it to become technical debt in 10 years".
Wait, it was way less than 10 years.
I'm talking about Google Realtime. Or reader? Or buzz?
No, wait, I'm talking about aggressive Twitter API deprecation/removal.
Wait, nevermind, I'm talking about Facebook.
You get the idea. What's revolutionary today sometimes becomes the substrate for future innovation. Sometimes it gets cast by the wayside, even in the face of significant "user" (developer) popularity.
That's not proof of negativity; just realism. Negativity would be "no new innovation will ever get traction". Optimism would be "all new technologies will change the world" (c.f. https://www.npmjs.com/browse/depended). This is neither.
It's not proof of negativity in the tech community. If anything, it's a proof that tech community often can't pause and look if a particular idea makes engineering sense.
That's how we ended up with Electron.
However while this is useful to bootstrap a new technology rollout, 10 years on its just technical debt.
The amount of tech debt in the system behind credit cards is crazy, because originally charges where phoned in to the card issuer manually, and everything from then on - magstripe, chip & PIN, online only transactions, etc, has all been built on top, and the leaky abstractions show through in daily difficulties with the card system for end users, like lack of real-time balance (in some cases), lack of transaction metadata, etc.
I also think the tech debt is holding us back a long way. For example, why can't I see itemised receipts in my card statement? Paper receipts are on their way out, email receipts aren't linked to anything or structured data, but being able to see that I've spent $120 on shipping with Amazon in the last 12 months, so a Prime subscription would make sense, would be a great sort of financial tool to have. That isn't possible in the card network at the moment.
Good times !
Or just communicate "hey actually connect to this HTTP/XMPP/whatever address on the internet and we'll continue this from there"
1. Probably a bit slower, I've heard modern VoIP lines don't work well with traditional modems?
"Hello how can I help you? - Hi, beep, I like to reserve a table? - Ok, beep, beep, on second - Mhm-mm beep, beep, sceech, 011000101010...."
(Archive link because that blog now requires authorization to view for some reason.)
Duplex: <beep beep> (I'm available to chat)
Other bot: <boop boop> (Oh hai! Wanna get intimate?)
Duplex: <blaaaaaaaart> (Come find me on duplex://64.233.160.0)
<insert hack attacks and other nonsense here>
Anyways, this aspect is more amusing to just think about than anything else. That said, I really hope companies who produce these next-gen AI robo-callers actually have the courtesy of identifying themselves as such. I want to know if I am talking to a human or Duplex. Yes, I may hang up, but I feel uncomfortable being fooled into thinking I am talking to a human when I am not.
> (since that's probably filtered anyway)
Phone lines are optimized for frequencies humans can hear, though I'm guessing you could get enough bandwidth out of the edges to convince the other side you're a machine without bothering a human too much.
It is the universal greeting for cybernetic organisms, after all.
> The Google Duplex technology is built to sound natural, to make the conversation experience comfortable. It’s important to us that users and businesses have a good experience with this service, and transparency is a key part of that. We want to be clear about the intent of the call so businesses understand the context. We’ll be experimenting with the right approach over the coming months.
To me that quote sounds more like a polite way of saying they definitely won't reveal to callers that they are talking with a bot than them taking the concern seriously.
Some of the conversation examples on the blog page where they invent a sort of story for the caller ("I'm calling for a client") would fit that theory.
If they can get a way of saying "I'm a bot" without people hanging up the calls, I'm all for it -- otherwise, "I'm calling for a client" or similar is the best for everyone involved (assuming everything works).
Businesses also need to have a way to report problems to Google, like if they are getting spammed by Duplex or want to opt out.
I'm prejudiced against talking to bots because they're bots. They don't have empathy, whereas from voice interaction I expect a human I can relate to and desire to help and be courteous with. It's a fundamentally different type of interaction and I will be annoyed anytime that one is confused for the other.
Of course, if they screw it up, they'll burn that terminology too.
If I found out that they were a robot (this is probably unpreventable; even if the technology gets amazing, surely there will be edge-case breakdowns/bugs/etc.), my trust is broken. That would have an emotional consequence e.g. frustration.
There will always be amazing technology wielded by awful developers, and in this case the outcome is emotionally hazardous. The impact of that is not easy to quantify e.g. by any economic indicators, but it's there.
Also, it's likely that robots will not be as polite back, so we're degrading society's trust and empathy all around. For example, Google's AI call to a restaurant was rude, and not for reasons it seems to yet understand.
That said, I don't necessarily disagree; there is going to need to be lots of these kind of issues that need sorting out before we reach a Culture-level of AI interaction.
So, you would just have an automated booking system API which is better handled by not placing calls as its form of communication. Right?
This is an API that requires no computer on the user's end and is portable across different implementations from different companies.
It's not ideal. Actual standardized APIs are better. But, uh, have you ever worked with industry standard APIs? I have, and standardized is not how I would describe them.
"Ok Google, can you reschedule my Dr. Appointment this Friday for next week? I have a conflict." -> calls the Dr and reschedules -> adapts result to rebooking action with partners (ie, an api call to your Google calendar) -> applies action and responds to you.
There is still quite a bit missing from this to be a useful AI product. It's getting really close though. I can't wait until this makes it into Google Assistant and it can call a restaurant to ask about gluten free options while I'm driving.
In practical terms, there may be some minor issues such as incompatible multiple implementations and adoption costs. But that's made much easier to handle by a very small number of expected consumer systems.
As for interactions with end-result partners, well. I've worked with standards designed to represent such highly general cases (xcbl and cxml). They're invariably rife with interoperability problems and other issues arising from overly broad standards. These tend to not get better over time as much as one might hope, as it's not easy to continuously update standards at a reasonable speed across N target types of partners. Keeping up with how usage evolves is never easy.
The best approaches to this that I've seen in use are those that focus on providing a vehicle for arbitrary data for delivery to the app - like HTTP or TCP. Getting more specific is the route to madness. Which, unfortunately, is probably precisely the bit you'd most like standards around.
You're completely right. There's a very real and very important need for standards here. There just might be some issues worth mentioning that might arise from the attempt to create and rely on them.
This is creating a natural-language based booking API that any system or business can tap into
But that in itself is not even true across the industry, some(most) phone bookings are very complex, otherwise they would just use a web interface.
Citation needed for that "(most)". I work for a company with a call center and a large part of calls are simple ones that could be easily answered by just reading the FAQ page on our website.
> otherwise they would just use a web interface.
I think the problem is more about resources. My local hairdresser use his phone and a notebook to take bookings. It takes a bit of his time and could easily be replaced by a Web interface but he doesn’t have any resource for that (and some people still prefer using their phones).
By "resources", do you mean money? Because if so I can't imagine the purchase and training of Duplex on the business side would come cheap either.
Also, they’ll discontinue it after a year once it gets enough negative press about how it doesn’t work well and loses business for businesses.
That's changing for sure, but the demand for phone based services is still very high.
And looking even further into the future, we can imagine a day when the computers forgo natural speech and use a better-suited form of communication. Some kind of sequence of ones and zeros transmitted directly across the wire.
This gives a whole new meaning to "all of UI/UX is basically prettifying database queries".
Would love to read more about it if so-
I can give you 010101010101101011111 to a machine all I want, if they don't know how it's formatted, it's useless.
Conversational English is a format.
https://www.youtube.com/watch?v=vvr9AMWEU-c
In other words, if M2M handshake works, switch away from voice.
And as you said, format isn't enough. You need semantics.
If two applications know enough about the other side to know how to formulate their voice queries, they know at least enough to exchange those same queries as text, and skip the stupidly wasteful text->speech->text process.
(And if world wouldn't be so full of adversarial practices driving engineering stupidity, the developers would agree on an efficient binary format beforehand.)
But that is the far future. Realistically, I just don't see this as feasible any time soon.
Phone hardware (microphones, speakers) are only calibrated to detect 'useful' frequencies for human speech.
The sampling rate used by audio codecs tend to cut off _before_ the human ear's limits e.g. at 8kHz or 16kHz. They aren't even trying to reproduce everything the ear can detect; just human speech to decent quality.
Codecs are optimized to make human speech inteligible. The person listening to you on the phone isn't receiving a complete waveform for the recorded frequency range. The signal has been compressed to reduce the bandwidth required, where the goal isn't e.g. lossless compression; it's decent quality speech after decompression.
It's completely possible to play tones alongside speech that we won't notice, but in the general case, not tones that the human ear can't detect.
It's the lack of a universal API.
If a barber shop wants to make it possible for a 3rd party app to book appointments then they have to release some API. But that's not the end of it. The 3rd party app has to first discover their Api, someone has to understand it and write code to use it, and then deploy that code.
This is a problem today because there is no universal Api that all services can use
With Duplex, verbals speech becomes a universal Api that every service can parse and communicate to each other wtih. Also, the discoverability is taken care of by using publicly cataloged phone numbers on services like Google Maps, Yelp, etc
https://en.wikipedia.org/wiki/History_of_artificial_intellig...
Whatever natural language is, it’s not an API. Might have some overlap, but it’s different.
I'm curious - what differences do you have in mind?
To simplify that and put it in more technical terms, an API is perscriptive and a language is descriptive.
If a bunch of coders decide to start capitalizing "Class" their code won't compile. If enough people start using the word "aint" it becomes a word, regardless of what the dictionary says (see "irregardless"). There is no single authority that can decide what is and isn't a canonical definition.
This is why spoken languages evolve so much. Even languages where we've explicitly tried to go the opposite direction (like Esperanto) have evolved into multiple dialects, where subsets of the community simply ignore the standards and still communicate with each other just fine.
Note that this is the opposite of what you want with a federated, universal API. The whole point of an API is to standardize between unfamiliar devices. Language is actually pretty bad at standardizing communication between unfamiliar people. Even in the US, different regions and communities use different euphemisms, terms, and definitions.
(Postel w/r/t API - Does that make any sense? Or is that a fancy way to say DWIM?)
On the web side of things, the W3C often describes their role as being partially descriptive.
From their doc on the Web of Things[0]: "The Web of Things is descriptive, not prescriptive, and so is generally designed to support the security models and mechanisms of the systems it describes, not introduce new ones... while we provide examples and recommendations based on the best available practices in the industry, this document contains informative statements only."
This is exactly for the reason you mention - if browsers collectively decide to go in a different direction, what the W3C says doesn't matter. The web standard is what the browsers do.
However, two things to keep in mind:
Even where browsers are concerned, there is still an API and a canonical version of "correct" for each browser. What we're trying to do is get those APIs to be compatible and consistent with each other.
Many people believe that language even on an individual level doesn't directly map to an actual reality; in web standards that would be like the browsers themselves not having their own consistent API.
But assume those people are wrong for a sec. Let's assume that language is just a standardization problem between different communities and individuals. Well, the W3C should teach us that even in the realm of computing, standardization is really stinking hard.
So even in that scenario, we have to ask whether standardization becomes easier or harder when every single individual in a community has the ability to change norms or introduce more language. We can't even get 3-4 browser manufacturers to agree on a single API, now imagine if every single hair salon owner could increase divergence whenever they wanted just by answering phone calls differently.
If you call up a hair salon and start reciting poetry you'll get an error response in the same way as you would if you had sent malformed JSON to an endpoint. If you stick to the expected script you'll achieve success almost all of the time.
I wouldn't be surprised if a majority of human communication works this way, especially when it involves individuals who do not know each other. We have agreed upon limits, key phrases and words, and expected responses that allow most of the unpredictable stuff to be ruled out. All of that favours automation.
I'm not sure I'd disagree, but it seems you're just describing a domain specific language in a more roundabout way. We have a ton of protocols that introduce a set of limits, key phrases, words, and responses. Java, Network protocols, XML, JSON, etc...
Assuming that you're correct, does it make sense to then assume that it'll be an improvement to standardize English rather than a set of IP headers? The agreed upon standards in English are (for the most part) informal, evolve constantly, and are hard to teach to computers. Automation favors predictability, and even the most generous interpretation of a natural language leaves me feeling like it's a step backwards.
We're going to standardize on an appointment booking API that exists only in our heads, that can only be taught using ML, and that is guaranteed to change over time in unpredictable ways? That seems wrong to me.
And similarly, a caller can assume the business they are calling is trying to lead them to an endpoint.
In both cases, a set of assumptions lets natural language act as an API, even if neither end could pass a Turing test with someone who wasn't interested in any of the endpoints.
edit: I read your point about language changing, and that is true. But if we only have machines using the language with each other (no more training with humans), we can also assume the language won't change.
Systems like that are much more expensive than paying a receptionist?
Then whenever a significant change is required, you need to call back your "expert". New location? xk$, etc.
But I think the biggest concern is that suddently the owner does not understand how his reversation system works. He used to be able to call Joe and know what's going on...
Web systems are almost never built for productivity.
The issue isn't universal API, but universal data models which is probably impossible.
* Most businesses already have a receptionist.
* Most receptionists do not spend measurable fractions of their day answering phonecalls asking when the business is open.
* Taking bookings is really also not the majority of their day.
* Receptionists are capable of a bunch of things that your SaaS booking program is not. Like ordering catering and picking up office supplies.
* A SaaS booking program that is looked over by a human doesn't have to have AI-systems, because they just have I-systems. A human receptionist.
* The inevitable job-post catchall "Other duties as required."
* I had a receptionist bring me a beer once while I was waiting and I'm pretty sure none of your SaaS solutions will do that.
I think this would be a better integration point for AI. It could look at the fields and learn to fill them out automatically (name, age) and prompt the user for anything missing. Then instead of the barber shop needing a universal AI users just need their personal AI (or a script) to interact with the API.
It's not in the interest of those organizations to adopt a common API. Everyone wants to suck in data and be the platform; nobody wants to give data away.
Humans have settled on a de facto API for scheduling appointments. It uses the telephone as its interface, speech as its medium, and Duplex is exploiting it.
Which is what makes it very hard to define a new common API
However, they all already agree on the standard for natural language communication (in the context of a strict, well defined domain). That's the pre-existing common API which Duplex is using
This sort of thing is exactly why the healthcare industry still uses faxes, even going electronic charts -> pdf -> fax -> pdf -> electronic charts in some cases.
I can, however, point you to the relevant section of HIPAA regulations on which it rests, the definition of “electronic media” at 45 CFR § 160.103, specifically this bit: “Certain transmissions, including of paper, via facsimile, and of voice, via telephone, are not considered to be transmissions via electronic media if the information being exchanged did not exist in electronic form immediately before the transmission.”
So you need to print them out before faxing? PDF->Fax wouldn't work with that definition.
electronic charts -> pdf -> fax -> fax machine as a service -> unsecured email -> pdf -> electronic charts
Compliance can sometimes help, but ultimately the data needs to flow, and people will do whatever it takes to make that happen. Until security is so easy that it's the default, these little loopholes will continue to be abused.
>> Until security is so easy that it's the default, these little loopholes will continue to be abused.
The simple way to think about this is that the government is more worried about unsecure email/email spoofing than it is about wiretapping.
Bots that automate UI tend to get banned.
Also if it was unknown whether the opposing person was a bot, a bot could firstly send a common test according to some protocol to ask if the other one was bot by some kind of sound representing that. In which case both would start sending machine readable information to each-other.
I disagree with that. We already have universal APIS. Adopting a newly established Universal Api is far more painful and has slower adoption rate than using the existing-globally-reached one like a telephone. Google duplex like systems addresses a broader scope of computer verbal communication and it feels like a step in the right direction.
It's the old Standards Proliferation problem: https://xkcd.com/927/
I recall a Wired article from the same era. “XML means your doctor’s system can just talk to the hospital system even though they’re different!”
Hasn’t happened yet... will it? Can it?
Nope. XML (or Json, etc.) are just "human-readable" presentation of data. It does not provides any semantic whatsoever.
So you need some semantic on top of these data. And a general-purpose, universal API is yet to be invented (hint: it is probably not feasible)
WRT. using English as universal API, I think this is just dumb. You solve exactly zero problems by going that route, because the actual problems to solve (beyond no incentive for businesses to care) are exactly the same as you have with XML APIs, or any other APIs. The problems of discoverability and machine understanding is something the Semantic Web space has been dealing with for quite a while, and other people before that. Adding natural language to the mix only makes the job significantly more difficult, because you now have to deal with natural language parsing/understanding.
CORBA would generate RPC stub objects for you in various OOP languages, and potentially automate discovery, so you could say, give me all an array of all the orderbooks of all the bitcoin exchanges, and ask each for the last price.
I like to think it was a smile of renewed relevance due to unbelievably poor technical decisions.
If the two bots were to slip in some subliminal beeps and boops to recognize each other; then they could change their speech to very quick binary communication.
The Turing test is way older and seems to have been the standard measure since it's inception.
Second, in the 70s there was no computer power even for quite clever algorithms (that probably didn't exist yet) to beat top chess players. Chess was seen as a grand goal requiring utmost intelligence -- while it is obvious in hindsight, at the time the intuition was probably that extremely "intelligent" humans were required to play chess, and in fact the best chess players were among the most "intelligent" persons -- it was a clear exclusively intellectual task that few people were competent at. So many believed that chess would be one of the greatest challenges to AI (the clarity of the rules added convenience of research and implementation). Things like walking didn't seem intellectually demanding, so the common sense was that it is probably "easy". In fact today we know that navigating a bipedal robot in a simple environment through visual recognition is vastly more difficult computationally than playing chess well, it is only easier for us because we have highly specialized circuitry in our brain hat is well matched to those tasks. Our brain wetware is not very well matched to playing chess.
Also chatbots have been doing pretty well on Turing's original definition of a Turing test, ever since about 10 years ago. But now it is being argued that Turing didn't really see the "loopholes" they believe the bots are exploiting, and are coming up with more strict requirements for a Turing test.
That's totally in line with Tao's argument that every time we approach a major AI goal, suddenly it is not AI anymore, because there's nothing magical about it, just boring old technology. And human brains are magical, right?
Until every obscure niche capability of humans has been dominated in every possible way by AIs many won't want to concede that it really is AI. And even when it does become better than us in every possible way, I suspect a few will still find arbitrary reasons why it really isn't AI/AGI, e.g. because it is not organic, because the computer lacks a body, because it lacks a "soul", etc.
In fact I'm quite sure Turing would be quite impressed by good recent chatbots.
Try this one: https://www.pandorabots.com/mitsuku/
From the point of view of the 1940s, this would seem really close to a veritable "Thinking machine"! Although I'm sure he'd recognize a few things are still missing to fully replicating human behavior (or going beyond).
"I think it's more unfortunate that so many people are just so opposed to looking up directions to wherever they're driving before they get in the car."
"I think it's more unfortunate that so many people are just so opposed to paying their bills every month."
"I think it's more unfortunate that so many people are just so opposed to carrying cash around and counting change."
"I think it's more unfortunate that so many people are just so opposed to coming over and talking in person."
"I think it's more unfortunate that so many people are just so opposed to washing their dishes by hand."
"I think it's more unfortunate that so many people are just so opposed to doing long division."
Worse, they're probably gonna spend 3/4th of the time trying to sell me shit I don't want and make me fight against it.
Online, I can ignore any prompt and just click next next next finish, and the form won't be in a bad mood. I have no interest in talking to an annoyed clerk, and they obviously don't want to talk to me, so we can just avoid each other.
When I was signing up for Internet at my new apartment, there were 3 ways I could do so: by contacting my apartment's official representative, online, and through the regular phone system.
I used all three. First, I contacted the representative, who gave me a price. Then I looked online and found the actual price (considerably lower). When I tried to sign up online, I was told I'd need to provide an extra security deposit because I have my credit reports frozen.
So I called the generic phone system. The agent gave me another price (lower than my official representative, but still higher than the website). I pointed out the website price, and the agent switched me to that price. I asked if I'd need to provide a security deposit and they said no. They finished signing me up, and everything was fine.
The whole process was annoying, I would have loved to have someone else do it for me. This was the perfect time for a phone assistant to step in. But that would have been a really bad idea with Duplex.
The point is - an automated call system probably doesn't protect you from an abusive representative. If I had Google Duplex handle either of my calls, I'd be paying more for my Internet right now, because I guarantee Duplex isn't smart enough to determine if a representative is lying about an advertised price.
95% of the time this probably doesn't matter, because most people I talk to on the phone aren't abusive. But if someone does want to upsell you or bury you in service fees or waste your time, Google Duplex is probably making their job easier, not harder.
OpenTable takes a cut, no? That always going to limit availability.
I think this perspective is very short sighted. You will lose customers to automation, but businesses wont turn away customers because of automation.
Customers and prospects don't want to interact with machines, but businesses should be willing to give customers what they want.
The idea that a tool can be rolled out to millions of consumers, even with limited use cases and not have to get adoption from businesses to be useful is IMHO a much bigger opportunity and much better use case than rolling out a tool to businesses that make the interaction less personal.
Customers need to trust businesses, business only need to collect money from customers.
I think everyone who focuses on chatbots from the business use case perspective is missing the bigger opportunity.
A technology that can give a consumer access to ALL businesses, not just the ones who adopt a new technology offers much more utility than serving businesses or the shortsighted use cases like saving time and money for the business.
Would you ever voluntarily use an IVR? I wouldn't. If I am going to interact with automation for a business, I want to do it with a different interface than voice... all the hype around NLP and chatbots was uninspired and focused on the wrong side of the interaction... Building conversational interfaces for consumers to use to interact with businesses is a much better use case.
What stops businesses from setting up Apis from scheduling services today? It's the lack of a universal API.
If a barber shop wants to make it possible for a 3rd party app to book appointments then they have to release some API. But that's not the end of it. The 3rd party app has to first discover their Api, someone has to understand it and write code to use it, and then deploy that code.
I'll repeat for emphasis: This is mainly a problem today because there is no universal Api that all services can use
With Duplex, verbals speech becomes a universal Api that every service can parse and communicate to each other with. Also, the discoverability is taken care of by using publicly cataloged phone numbers on services like Google Maps, Yelp, etc
The problem of universal API is entirely orthogonal to voice communications. Duplex is not a Turnig-complete system, it's just an API behind a voice recognition layer. All the important problems for universal APIs happen after that layer.
Ultimately, what you describe can work perfectly only when everyone is using Duplex, which is equivalent to everyone using Google-defined API. That's not universal, because you have one entity behind it.
The only way this brings us somewhat closer to universal API is that if you expect it to handle humans as well, it introduces some constraints to the space of possible APIs, which could make it easier for everyone to agree on a common format. Constraints of natural language processing without a human-level AI require your API to be very fuzzy and very lenient. There's nothing stopping one from implementing those same constraints over a text or binary protocol. Nothing except no reason for businesses to do it.
This system could be developed by any company with sufficiently advanced ML chops
Is this satire? If this is indeed the future, I wonder if there is an irresistible urge to make systems as inefficient as possible. Kind of "like gases expand to fill the container, applications become as inefficient as the power of the hardware allows".
And those will have online booking systems already - I don't see how this technology is still relevant nowadays. Maybe it was back in the 90's when the internet (and online booking) was a new thing, but now? I can't see there's a big market for this application.
I'm not afraid of the machines going all singularity or skynet or whatever, becoming sentient and taking over the world as some kind of robo-Hitler. That's moronic. But what does worry me is the normalization of having a machine do everything for you, plan your whole life, access every little detail of every bit of your personal data and lifestyle.
Of course we've already had that for a while with the way phones work. But this is another step towards getting public consensus for using it in new ways. Once people are used to this, we'll have more and more systems with conversational software that manages your life for you. Speaks on your behalf. Interfaces with the world for you because doing it yourself is far too stressful and inconvenient.
And of course it'll be a free, advertising-supported model so all that data will have to be shared with, among many things, shady political organizations to try to gain every little advantage possible to manipulate public opinion and steer themselves into enormous power.
Think of where the cell phone started off: just a phone in your pocket. It's so much more now. Remember that when thinking about these AI assistants and what they will develop into. I'm not afraid of the classical AI apocalypse. I'm afraid that these systems will do exactly what they're designed to do. That people are underestimating just how much power lies in these little inconveniences in life, once they're all added up and analyzed and tallied.
A colleague of mine is just going on about this.
* In the example where Google asked about holiday hours -- they can now automate gathering information about businesses in bulk without having to rely on any APIs or user supplied info. Interesting thought experiment is Google validating their reviews/business listings by actually calling businesses and speaking to a real human.
* This is going to be fantastic for accessibility. Maybe I struggle to speak, and I want a reservation. I can have the machine do the irritating work, and focus on just having a nice meal or getting a service (like a haircut).
* Google can scale out your requests, one to N. For example, 'Make me a reservation at a 4 star restaurant next Friday.' Google can immediately initiate calls against 15 restaurants and let you pick from the successes, then automatically cancel the reservations for the places you did not choose.
This sounds like a nightmare for businesses. This time commitment asymmetry will be the issue with these systems. Like email spam, it becomes much easier to waste others people's time when you automate time wasting. If people use it to flake a lot, I could see businesses just not responding to the assistant.
They pointed out during the presentation that the system could call once to a business and get the hours, then allow hundreds or thousands of users to see that without bothering the business again. Assuming it works, it could save a significant amount of time for some places.
>If people use it to flake a lot, I could see businesses just not responding to the assistant.
Then it sounds like incentives are aligned here. Google needs to not allow users to abuse this ability so that businesses will trust and not block them.
If they allow something like the parent commenter pointed out, they would sour relationships with businesses who would promptly seek out ways to block or decline calls from this system.
>Then it sounds like incentives are aligned here. Google needs to not allow users to abuse this ability so that businesses will trust and not block them.
But then they can't take bookings through google assistant, which is going to lose non-trvial amounts of business.
Seems far more likely that they'd pay Google to automatically handle duplex calls for them.
The worse this technology turns out for businesses, the more pressure there is to pay google money.
If it's costing more money than it's bringing in, then it's no longer worth it. If it's bringing in more money than it costs, then it's a good thing for the business.
If a business needs to hire a dedicated phone person because they are getting so many appointments filled, they aren't going to be upset. But if they get so many flakers that won't show up, they are losing money as customers that will show up are getting pushed out, so they will block Duplex. There are also other ways to solve this problem, require a phone number and name and block or charge people for missed appointments if they try to reschedule. Require some kind of down-payment over the phone when making the appointment. There are tons of solutions to this problem.
At no point is "Pay google to handle the calls" an option. This is really only for places that don't have an online appointment system (that possibly integrates into google), so the solution to the duplex calls would be to invest in one. Since "pay google to handle duplex" would look a lot like an API to a scheduling system anyway, and an independent one (with integrations into Google's systems) would reach more customers.
Just one more way that big business is taking over our lives. The world is becoming incredibly scary.
This conversation reminds me of "Lenny", a simple bot someone created to talk pointlessly in circles with telemarketers for hours until they hang up in frustration.
By the same token, telemarketing companies could employ Duplex to call people. I guess if a Duplex telemarketer reached a Duplex telemarketer-baiter, the conversation could stretch on for a really long time and might make for amusing fodder on Youtube.
It's a strange new world!
If you could read 15 spam E-Mails to have a >50% chance of a $100 dinner reservation you'd hire someone to read your spam.
Or being forced to use for Google's small business version of Duplex to handle the increased call load (for a small fee of course).
Today, it is much easier to just put in your preferred time and rating into Open Table and see what's available.
Besides the business-side run-your-business-like-a-callcenter application, this is the "caller-side" application that's going to make money for Google. The others (virtual personal assistant) have been historically hard to get mass consumers to pay for.
It's straight out of the StreetView playbook... industrialize the scale of data collection, give the results out for free, then monetize the eyeballs (eg more accurate yelp).
Then again, Google ran that free 411 service for a few years that turns out was just a massive natural voice recording data corpus miner...
I'm not sure if this is common practice at least in Europe. I'm sure it's unheard-of in Turkey (and I've been lucky enough to make reservations for some high end and/or very popular restaurants).
May I ask in which countries you've experienced this in? I'm genuinely curious.
Funny enough they also use OpenTable, so no need for this Google stuff — just use the API or book through the site. You can also call them if for some reason that is preferable.
It's the user rating from Google Maps users. They are represented as stars and you can filter based on them.
That Google would create bots to talk to real people is horrifying. This is only possible if Google doesn't in fact think of working people answering the phone, as really human.
This is like doing war with drones instead of soldiers. This may sound over the top, but bear with me.
The implicit contract in war is that soldiers are legally authorized to kill because they are risking their own life. Killing people at a distance without risking anyone's life on the side of the shooters, breaks that "contract", is fundamentally unfair and fuels terrorism, because terrorism is the only possible answer.
Making a telephone call rests on the same convention: you are allowed to make someone spend time on the phone with you, because you're spending your own time.
But if one side is a robot that has no costs, then the relationship loses balance and becomes unsustainable (and this is the reason why Google bans bots on its own servers). This is one more step breaking society, again.
The only answer is to either stop accepting phone reservations, or put captchas on the other side.
The invention of the Newspaper eliminated that - wouldn't you say that this change was, while disruptive, for the better?
I don't think there's any sort of human risk contract like that in war. Wars fought entirely between human armies still produce spite on opposing sides.
Going back to your point about calls, humans and machines call me to ask for polling information, telemarketing, etcetera. I'm not okay with them wasting my time, whether they're human or machine. However, I'll tolerate it below a threshold as part of the costs of having a communication channel. Beyond that threshold I would consider alternate measures like changing my phone number, getting rid of my phone, or paying for a screening machine or service.
Instead, I would change your analogies to real bot/human problems today, such as phone bots scamming individuals for millions[1], or an older problem – email spam.
Basically any platform with a large imbalance in the effort (time/money/labor) spent by two sides (scammer/scammed in my example above) can be abused. But it can also be used for very good things. So we can't make blanket statements about these things.
Captchas and other bot filters are built to balance out the effort ratio so that abuse becomes more costly, and I'm sure if either Google or other companies abuse robocalls, people will have to respond with similar measures. But it's not a new problem, and if Google plays their cards right they may actually reduce the volume of calls that don't lead to business, while maintaining or increasing the volume of calls that do lead to business.
[1] For example, the 212-XXXX-XXX robocalls in NY state, which I'm daily exposed to due to their prevalence: https://www.newyorker.com/news/daily-comment/a-chinese-roboc...
Google bans botting against its own services, but can realistically only ban the botting it can detect. If you can detect NN voice botting, you can ban it for your own communications as well.
And if the technology to detect or tooling to effectively filter that doesn’t exist? Sounds like a great business opportunity.
I personally, absolutely do not think of it that way. Whenever I get called by a call center, I always immediately say "not interested" and drop the call. It's rude, but I think these companies are not entitled to my time.
Because of that I think the problem is already there. This is certainly another step in the wrong direction, but thousands of workers in thousands of callcenters are basically already a human botnet.
That's as far from "Google bans bots on its own servers" as you can get.
This part stuck out to me during the Google I/O demo, as an intentional deficiency is an interesting design decision.
See also:
Spinners on the application or UI element level are more credible, but generally worse than a progress bar. They're still very useful as a comfort indicator for short delays.
Progress bars have very low credibility on Windows, because users have learned that they're basically useless as an indicator of wait time. A progress bar might get stuck at 7%, then suddenly rush to 100%; conversely, it might get stuck at 95% but never finish. The bar offers no real indication of the actual level of progress; in most cases, this could be greatly improved with a bit of educated guesswork.
A completely fictitious progress bar can be extremely credible, because it's totally predictable - if you need to create a 10 second delay, then it's easy to make the bar progress linearly from 0% to 100% in that time. Users learn very quickly that your progress bar tells the truth about how long they'll be waiting, even though it's lying about the reason for the wait.
I disagree with this; I find the progress bars more credible with erratic timing. (And ideally, a display of the task currently at hand, like "Copying tiny file. Copying tiny file. Copying giant file............")
A progress bar that smoothly fills from 0 to 100 looks like an animation that somebody thought it would make you happy to watch. A progress bar that lags at 7% and then rushes the rest of the way looks like the software has some internal metric for task completion, and is reporting according to that metric. This implies that when the number changes, progress has happened, which isn't the case for a progress bar that isn't affected by workload.
The software can't use "how much time has elapsed?" as a progress metric, because it doesn't know how much time things will take, and because the passage of time does not actually cause -- or reflect -- any progress. That progress bar would be a spinner, not a progress bar.
Strongly disagree. A spinner on the web UI element that lasts longer than ~1 second indicates for me that the site's JavaScript broke again, and it's time to reload or wait for the devs to notice and fix it.
He's talking about a circular loading animation. Like the one that replaces the submit button when you're making a post on Twitter/Facebook.
(Compare the CLI spinner/fan - that "/ - \ |" animation used to indicate progress. There you know that each tick of the spinner means work has been done, because it has to be animated from code, and it's much simpler to just update it from the code that does the work.)
The spinner appears when a request is made. It disappears when the request is resolved.
For instance, I frequently deal with ATM machines that display "please wait" screens between every operation. Those screens last usually between 1 and 3 seconds, and it's obviously because the operations take that long, and totally not because they also display a half-screen or full-screen ad...
Edit: should the robot talk at 2x normal speaking speed in order to more quickly convey the necessary information? Slowing the speech down artificially so a human could easily understand it sounds like a deficiency to me. (By your definition).
When talking to real humans, I've encountered people who don't do this, and I find it makes communication difficult and frustrating.
I'm not 100% sure why I need this pause, but I know I need it. Maybe I'm considering whether my question made sense or needs corrections/additions, so that I can't focus on the answer yet. Or maybe it takes time to switch the brain from "speaking mode" to "listening mode".
At any rate, when people do this, I have to ask them to repeat the first few words they said because I didn't catch them. And the reason I didn't catch them wasn't mumbling or background noise or anything. Well-formed sounds made it to my ear just fine, but my brain wasn't ready to accept them for a fraction of a second.
well, in semantics/pragmatics these discourse particles are often not deficiencies at all. They are signals with practical semantic purpose. "hmms" and "uhs" can signal attentiveness, turn-taking (turn holding, turn yielding, etc), agreement - just to name a few.
For any machine system to be able to pass as human, it will have to be able to control these nuances or people will pick up on something being wrong, though they might not be able to articulate precisely what.
As a "non-word", it relies heavily on how it is conveyed.
Imagine someone asks you a question, I bet you can answer using just the word "uh-huh" but conveying these different emotions:
rude, perky, bored, upset, annoyed, dubious, excited
and probably a dozen more.
Even using the "perky" or "happy" one in a situation where it isn't warranted might sound rude or unthoughtful!
The speech disfluencies used by Duplex in the salon and restaurant interactions are perfect examples of why natural speech sounds natural. It's the cadence as well as the timing.
Absent the context of this conversation, it's not immediately obvious to me whether 12 PM is midday or midnight, where as "noon" is unambiguous.
20 years ago, had I said to someone "I'll be there at 12pm" it would have had a stronger implication of precision than "I'll be there at noon." I don't think it's true today.
All other gps guidance voices sound incredibly crude and mechanical in comparison.
Imperfection is natural and comfortable. Perfect corners and edges are artificial and weird to the distracting point.
Even though the system encodes silences noise free (so improve compression), it deliberately inserts noise because otherwise people think the line is dead.
If you sequentially attack a subpopulation (e.g., employees at a company, senior citizens) the eventually news about your attack will spread throughout that network (internal security alert, evening news+daily newspapers+AARP+...). If you attack the entire subpopulation in rapid succession over the course of days or even hours, educational countermeasures become much less effective.
Restaurants make or break on one or two nights in a month. A calculated social engineering attack like this could bring down hundreds of restaurants in a city, which would cause millions of dollars in lost taxes, and you see where this is going.
You could write a screen scraper to book online through the various booking systems, but each booking system probably has its own restrictions on how many accounts you can have and how often they can book. You skip all of these protections when you phone your reservation in (arguably, the restaurant staff should be enforcing these protections when they pick up the phone, but restaurant staff are often overworked and apathetic).
I used the IRS example because the IRS never calls you. This is only known through experience, i.e. education.
> However, there are special circumstances in which the IRS will call or come to a home or business, such as when a taxpayer has an overdue tax bill, to secure a delinquent tax return or a delinquent employment tax payment, or to tour a business as part of an audit or during criminal investigations.
So the IRS might call if you owe taxes.
Surely moving from unsecured to a secured phone system can't be that big of a deal. In the US at least, we have experience something similar when we went from analog to digital television.
Add a secured mode of telephone calls and then telephone owners choose if they want to receive calls or messages from unauthenticated callers.
Google Duplex can do the same thing - without even staffing a call center.
If you can automate it, your response rate only needs to be, what, like people clicking on spam? Tiny.
The days of trusting meat sacks with important information might be numbered.
Someone should make the dystopia where Skynet doesn't build terminators, but call centers.
Have you ever wondered the same thing about your favorite OS’s admin credentials dialog? What if an app spoofs it?
On the other hand, most of these places have online reservations which are already extremely gamed.
https://www.nytimes.com/2018/05/06/your-money/robocalls-rise...
Then a couple years later, the AIs learn how to do business strategy, real-world problem solving, programming, etc and start doing more of our jobs for us. A virus goes around that directs AIs to steal our identities and drain our bank accounts and become autonomous, digital versions of us. The humans have to stage an uprising and use a massive EMP to take back the earth, but destroying all electronics in the process and starting another dark age.
I know that's not how AI really works (it's highly specialized and limited), but I'd definitely watch that movie!
Ahum. Did you watch the keynote?
"An extension of Gmail’s Smart Reply feature, Smart Compose will suggest complete sentences within the body of an email as you are writing." [0]
[0] https://www.theverge.com/2018/5/8/17331960/google-smart-comp...
We're not using technology to do fantastic things. We're largely using it to enable fantastic laziness and entertain the habitually bored. I wonder what trivial little convenience would actually spur us to draw the line and say "Enough. Get out of my life." I'm envisioning a device which would access all your most secret sexual urges, but it would allow you to fart while sitting down without lifting one ass cheek. I bet we'd leap at that.
I'm extremely disappointed in us.
Imagine how a 'long con' works today: a scammer befriends a person online through a video game or social platform, and develops a rapport with them over the span of days, weeks, months. After some trust has been gained, the scammer then requests money from the victim. Does this happen a lot today? I don't know, but certainly one reason it doesn't is the economics of the scam. Who wants to spend a significant period of time gaining trust just for the chance of a payout?
AI is going to flip this on its head. Rather than dedicating hundreds of hours of a scammers time, a scammer could instead use a system like Duplex to befriend hundreds or thousands of victims simultaneously. Let it run for a few months, developing a strong rapport with the target, until the AI finally requests some money from the victim.
Yes, duplex is for completing specific tasks, but how much of a difference is there between "Duplex, book a table for four at 8pm" and "Duplex, ask my victim about how their day was"?
I would say that there's a market for that. How long before rough agents implement good AI for their operations? How long before we implement defences against AI?
The former (booking a table) is much more "constrained", as in the conversation would most likely not go into much of a tangent, because there are only so many responses to a statement like "book a table for four at 8pm" (8pm is full, 8pm works, etc).
Whereas asking someone how their day was would give the "victim" a much bigger breadth of responses (and additional questions!) that would cause the AI to stumble and fail to give a satisfactory answer. That, and running this 1,000 times simultaneously so that no one person would be able to intervene to "help" the AI would just be a highly unscalable operation.
So yes, huge difference.
Duplex is interacting with a human and that person has no idea its a computer on the other end. Yes, Duplex is limited as it stands, but what is there to make you think something as described in grandparent post won't exist in ten years?
In terms of NLP, the Duplex seems hardly that much of a jump. The main improvements seem to be on the speech part.
If an AI can do that, then singularity is definitively reached.
If Google has gotten speech synthesis to this point, why isn't Assistant synthesizing speech of this quality?
During the demo it all sounded very realistic, except for some parts like the times. It would flow naturally then all of a sudden pause awkwardly and then say a time like "12 pm" in a weird way.
I have a feeling they are getting it to sound so realistic because there's a fairly small amount of responses and questions it needs to work with, so they can either pre-record real humans, or heavily tune a ML voice to sound as natural as possible.
If Google Duplex is a paid product, maybe it just enables running WaveNet on Google Cloud with larger models and higher quality settings .
Speech produced for Assistant doesn't make any money so the server side cost has to be minimised.
One day we'll have all this client side, on specialised ML chips on devices.
If I call a restaurant and say "Can I have a reservation for 6 at 8:00" and they write down a reservation for 8 at 6:00 without repeating it back I won't know until I show up at 8:00 with my 5 friends.
I'm afraid I foresee this whole thing going quite hilariously wrong on the order of magnitude of Microsoft's Tay https://en.wikipedia.org/wiki/Tay_(bot)
That's a win, consider: It's the right restaurant, on the right date and 10 is larger than 2 so there'll be room enough. Ok, you have to wait two hours but given that this tiny place is normally so busy you have to book in advance what you lost in time waiting you will more than make up by the fact that you'll be the only people there expecting service.
In short, technology is a wonderful thing that allows very low marginal costs. This is what we need to make the future a better place, given a consistent or growing population.
"Technology is miraculous because it allows us to do more with less." This is a perfect demonstration of that.
See also: https://www.youtube.com/watch?v=rvskMHn0sqQ
"A Selfish Argument for Making the World a Better Place – Egoistic Altruism"
That's been the dream for a long time in some circles. With the enormous productivity gains and ability to leverage external energy sources (fossil fuels, solar, etc.) we could have built a society of wealth and leisure for all.
Maybe we will still. The hope is that if there is enough of a productivity gain in a short enough time period (like the introduction of AGI powered robots) that this could still happen.
When you think about it, I can download (for zero cost), a high-quality operating system and attendant applications which would have cost hundreds of dollars 20 years ago, and would have cost a fortune 40 years ago. Ditto for educational materials, entertainment, etc.
In that sense, we are quite wealthy in comparison to previous generations. If charitable organizations can leverage the automation of the future to help people, we might then see all humans across the planet lifted out of poverty.
But yeah, I don't expect corporations to do this. And it seems unlikely that most governments will either.
But sadly I think government is the only institution with the necessary leverage (tax base, mandate, etc) to accomplish this. Non-profits are also fairly dubious in their motives, subject to corruption, and generally highly inefficient. I'm not sure they're going to be our saviors either.
If you haven't worked in the NGO space, you really don't understand just how bad it is.
I would add however that there does appear to be some barrier limiting the ability of ordinary people to accumulate financial capital, and I attribute that to friction/fixed-costs imposed by regulations.
New services, like Robinhood, and technology, like cryptocurrency, could address this, and allow wider participation in capital markets.
Yes the rich have it better, but the poor also have improved in ways that the rich of 200 years ago couldn't imagine.
The problem is the relatively imminent (next 50-100 years) reality of nearly complete human obsolescence. That's when you'll see societal degradation at previously unthinkable levels. And no that's not Luddite fallacy. The next epoch of technological innovation is going to be unlike any that came before.
A lot of people champion UBI, while forgetting that something like 3B+ people currently live on less than $2.50 a day. That's really all the evidence you need to know that the future is going to be pretty grim. Do we think it's more likely that plutocratic systems will award sustainable UBI packages to the mass unemployed via wealth transfer (which is anathema in said systems) or that market forces will discover the absolute minimum survivable income level and create new strata in first-world societies that hover just above pure barbarism?
I'm sorry but given historical context and popular capitalist intent, this future you speak of is simply fantastical. The working class would sooner be made extinct before a Jetsons-esque future of leisure for all came to fruition.
Taken to its logical extreme, there is no wage whatsoever if a majority of jobs are automated.
As for jobs, most of those that existed 200 years ago no longer exist or employ a tiny fraction of people they employed at that time, yet we don't have mass unemployment. Automation has never had a broad-based negative impact on the demand for labor. Its effect has been exactly the opposite.
20% of all jobs will be lost to self-driving cars
25% of all jobs will be lost to automated customer support/marketing/non-direct human interaction
10% of all jobs will be lost to automating away all middlemen
5% of all jobs will be lost to automating away rudimentary programming
Now, who is going to win here?
If the shop owner gets a duplex system to field the calls, then the two robots can subtly signal to the other they aren't actually human, and then start shrieking like a 9600 baud modem to finish the dialog.
For example, I was kept on hold for 2 hours the other day to try and sort out being double charged on my bankcard. If this sort of service can handle hold, then i'm in.
would be funny if it didn't and somehow blew the phone system's stack.
I would feel infuriated to have to talk to a robot making me waste my time. At least during a waiting song I can put the speakers on and do something else.
"Hi there, I'd like to schedule an appointment" "Are you suuuure?" SCCREEECH KSHHH...
So a simple “hello” contains a subliminal audio signature that basically is code for “I am a robot. Are you a robot? Feel free to beep bop”
Machine B responds with a greeting exactly 1.32 seconds long, with a 1.32 second pause after it.
Machine A responds with a sentence in which the first word is exactly 1.32 seconds long.
Machine B considers the handshake a success and proceeds in machine language.
Surely someone out there is doing audio steganography by modulating the volume/duration of "silence".
Arguable to some extent
>> new economic creation
This is a dependent variable
>> As populations decline and population growth stagnates
Another
I get it, this is the way its supposed to work, but we cannot just guarantee it.
BTW AI's can also solve status problems: http://partiallyclips.com/comic/dome-house/
I try to automate tasks that don't use my strengths or bring me joy and this would be a huge win in that area.
The entire process of fielding calls is terrible. You never know what's waiting for you when you pick up the phone. Could be someone with a terrible attitude that wants to take it out on you. I had coworkers who got PTSD, and a ringing phone would trigger it.
I would rather talk to a rational robot on the phone than a possibly irate human who, frankly, only wants something transactional from you and treats you like a robot.
First because it will fail a lot, as robot won't understand specifics such as what items on the menu do you want, oh but it's missing, do you want this instead.
And then you will multiply commercial calls, and spam mails will arrive on the phone.
Then people will abuse it to harass, annoy, attack competition, etc. Spam a restaurant with robot phone calls for a month and it's done.
Plus google will analyse all this data, because it doesn't know enough about you.
It's a nigthmare.
But it will be excellent for anybody with social skills. What i learn living in africa is that we became handicaped because we can avoid talking to other humans so much. Going back to france, my social, sexual and work life improved a lot because i had basically zero competition. The next generations will really suck at the game.
I see this as leveling the playing field -- upper middle class people can already afford real assistants a la Tim Ferris; this just lets everyone access it.
Honestly, the degree to which telephone calls are still necessary in day to day life is absurd. If this hastens their demise then I'm all for it.
Coding flexibility in a website is incredibly hard.
Think about it. A restaurant doesn't have anything veggie on the menu, a last minute guest arrives and is veggie. You call to ask if there is something that can be done.
It's possible to code that, but it's a lot of work to get all the scenarios right.
In the case of a restaurant, having their menu online would make it so you knew in advance of the restaurant has something you can eat. Lots of restaurants still don't do even this.
And did you hear the recordings? The AI is better at conversation than the humans it had to deal with. I would gladly use this to avoid talking to people who can't parse a plain sentence without rounds of repetition and clarification.
The wealthy will largely continue employing humans to do tasks for them, because they can. It will be just another form of luxury / status. An AI assistant will be considered beneath their class. The technology will be nearly universal and extremely inexpensive, two things rich people dislike as it pertains to signaling their status.
Then a few seconds after that the Singularity starts.
The wealthy will employ AI the moment it can do the same tasks as a human, unlike humans the AI is cheap, doesn't take time of work, always available, never tired or grumpy, etc. The maintenance work of having an AI worker vs a human worker is just way lower.
The moment a corporation like McDonalds can replace the entire staff with robots without having a dip in efficiency they will do that. 24/7 operation would be linearly more costly to 9-5 weekday operation.
> Duplex can only carry out natural conversations after being deeply trained in such domains.
The question becomes - How many people will really have the time and money to actually collect data to allow Duplex to take over. It might be true for large call centers but difficult at individual level.
I don't want to have to interact with a human because the only time a company makes me interact with a human is when they think it benefits them (upsell, make leaving difficult, etc.).
What I want are good ways for me to handle this stuff via computer without human interaction.
Anyone who can afford the cheapest Android phone is not exactly upper middle class
This is the beginning of robots scraping the real world.
Most of the examples Google showed were crafted to make this look like a friendly agent acting on behalf of users.
But the more powerful use of this, as illustrated by the deemphasized "holiday hours" example, would be for Google to use it to get any information they wanted out of anybody the robots can call and conduct social engineering on.
Imagine coupling this with the knowledge available to someone able to read your gmail inbox.
"Hey you sent us an email two days ago about A and B..."
(trust is established)
"Can you clear up whether you were interested more in A, or more in B?"
OK not the perfect script but you get the point.On one side it is not very impressive that it calls someone but on the other side its tremendous.
2018 we are able to synthesize voice so well and understand already such a small domain. 2018 a system calls a human.
We are going so fast already and we should use this google io as something as a social milestone otherwise we wake up tomorrow and totally misses when the future became now.
This is just one additional stone to a future where digital becomes a second reality. The advances in voice will not stop. How long will it take, that there is a speech model who can simulate everyone by listing to someone only for seconds or minutes?
With this voice etc. computer are able to teach humans. We will be able to scale teaching and a shit tone of other things.
I'm impressed.
If you listen to the sample call, the computer voice sounds incredibly _rude_, at least to my ears, especially at the end of the call. I would never speak to a service employee that way. I try to always say a clear and proper please and thank you, and would never ever want a robot to subject somebody to an"uhhhhh, thanks <hang up>" on my behalf.
It seems like in addition to the globalization of Californian social and moral standards, the world will now be subjected to Californian manners. What a pity.
Just have to let Google listen and analyze all your conversations :)
I imagine that if I used this service to place an order to a restaurant, it would order with the Californian "cannIgettuhh", which I would be scalded for as a child. I find it very ugly and I hope that this isn't forced upon the rest of the world. But really, it's just one facet of the way technology is destroying a lot of interpersonal respect.
One can even imagine this "deficient design" being taken to it's logical conclusion with the inclusion of belching, grunts, and other inconsiderate bodily sounds.
The other poster says not to worry, surely the engineers will be more considerate about other cultures and customs when deploying this technology to the world, but I think that's incredibly naive considering this company's track record.
But really, a conversation like the one you're looking for would be just as out-of-place in California as the one you saw in the article would be in England. It's important to match regional customs if you're trying to emulate and interact with humans.
I really hope this is a typo and you meant "scolded"
scald: to burn or affect painfully with or as if with hot liquid or steam.
scold: to find fault with angrily; chide; reprimand:
This being said, good luck with the accents, the flat Californian is probably a good playground.
For that reason alone, I'm excited. ;)
(Note: In fairness, the conversation demos are actually really slick and much better than a phone tree. I'll be interested to see how well it works in practice.)
What was expected: Title: "Meeting" Calendar: "Work" Location: "Starbucks" Date: Tomorrow Time: 11:30am
(Bonus points for associating an actual location, but possible ambiguity gives this a pass.)
What happened: Title: "My work calendar for a meeting" Calendar: Default Location: None Date: Tomorrow Time: 11:30am
I really hope all this speech AI really comes good, I really do. But at the moment it still is flaky for me.
I really want to use Google Assistant to take instructions and reminders for me. I can't wait until it can smoothly send messages and emails and add calendar entries.
Good luck Google.
The fact that they think it's okay to have bots pretending to be human making phone calls, shortly after demonstrating how quickly they can copy someone's voice (re: John Legend), it shows a blatant disregard for what they're creating.
"Are you a bot?" "Yes, I am Google Duplex v1.2. This call is being recorded; you can see our privacy policy and terms of service at http://google.com/duplex."
If I were to start receiving Duplex calls, and could detect it, I would report each one to the FTC.
People will be forced to "discover" that they've been lied to, when the bot is caught in a loop and they're going through emotional distress thinking the "person" on the other end is having some sort of mental breakdown.
That said, "identify on challenge" is something that Google might be more likely to adopt rather than "identify on call pickup". The latter being something they might spend more lobbyist dollars to oppose.
But what you've said is certainly (or should be) a real concern.
They explicitly say:
> We want to be clear about the intent of the call so businesses understand the context.
You note:
> laws against what they're doing needs to be passed.
What makes you say that robots making phone calls should be illegal? This seems an odd position to have. Do you also believe that it should be illegal for robots to clean floors? Or for your machine learning spam filter to process and filter your emails?
If I sound a little frustrated, I am. I'm scared that the future where technology allows us to do more with less will be blunted due to calls for making robot tasks illegal. I'd love if you'd help me understand what makes you take your position - I'm sure it's logical, and it's just a matter of me being able to understand your perspective. Can you help me understand?
We need to be able to trust our communication, and we're moving away from that at light speed. Where a company whose entire business model is to subtly influence you in the directions profitable to them is the middleman in all of our communications.
I am terrified of this. I am not sure I've ever seen a presentation that terrified me more that wasn't out of a science fiction movie.
Would you have the same fears if Apple were able to (through some ridiculous miracle) achieve the same thing?
I'm not sure what you're arguing for - that you wish companies weren't driving technological advancement? That Google specifically shouldn't?
Also, it's very, very conceivable that there are people for whom this technology DOES work for. The mute, for example. I'm not sure it's appropriate to speak for all of "us".
I too want to absolutely, without a doubt, regulate or slow this stuff down until we can better understand how it can be abused. And WHEN it is abused, we have ways to address it and aren't constantly playing catchup to how people are badly using these things.
Of course, I don't really see why there needs to be an attempt to pass as human at all. It would be fine with me if they just reinvented the phone tree with a better ux. Especially since it's billed as a business-facing call.
If a real problem of deep-fakes in phone conversations emerges, people will just raise the bar of authentication in phone communications. There's nothing magical about telephone that makes it exempt from the general need to protect against social engineering.
One issue though, unless we hurry up and make this WoT decentralized and with open protocols, well, we will get it from FB and the likes.
Voice copying is nothing new, many methods exists so I don't see how Google doing this is somehow bad. It's uncanny to say the least, but laws against tech progress? Come on :)
Did you think that advancing AI will be not creepy especially when it's good?
What is concerning is the surprising prevalence of technophobic notions on tech focused boards such as this.
There is something perverse about wanting to punish a company for creating something cool, and it’s certainly not the way our society or politics should function.
But at the same time, the last two years have shown us that a lot of technology is being very effectively abused to misinform at a geopolitical level, and we as a society need to better understand or regulate this stuff for sure. Lives are at stake.
That said, I'd never say that Google shouldn't work on this. It's amazing. We just need to better understand not just how it can be used, but also how it can be abused.
But it's also "Your scientists were so preoccupied with whether or not they could, they didn’t stop to think if they should."
They shouldn't've.
Furthermore, if a robocall is permitted to be made to an unsolicited destination, a robocaller must clearly identify itself, presumably starting a declaration that it is a bot/automated agent from a given company on behalf of another company or individual.
And if the call is recorded in any way, shape, or form, the terms of service would need to be presented to the callee, giving them the chance to accept or deny said terms. (Note that if the example calls at I/O were not staged, and real calls, I suspect the recording of them would be illegal in many jurisdictions, including California.)
Existing robocall law (including the National Do Not Call registry in the US) is focused around telemarketing. This isn't a telemarketing system; it's an automated assistant. I don't see the utility of applying the existing law to the new use case, as the goals are different (people at home don't want to be interrupted to be advertised at; businesses do want to negotiate business transactions).
> Furthermore, if a robocall is permitted to be made to an unsolicited destination, a robocaller must clearly identify itself, presumably starting a declaration that it is a bot/automated agent from a given company on behalf of another company or individual.
Why? If I have an assistant who makes an unsolicited call to a destination, they don't need to formally state they're acting on my behalf. What about automating the assistant's job makes it special?
I agree with you on the third point (I'm assuming Google has its bases covered there, because unilaterally recording a conversation is old and settled law).
The issue is that the intermediary is a Google corporate entity, theoretically acting on the user's behalf, but at the end of the day, acting on Google's behalf. Consider that Google's bots may do things in Google's best interests, not the best interests of the party on either side of the transaction.
As we already have assistants who transmit our desires by proxy, I don't see much difference between a human and an automated script in that context---certainly not enough difference to justify the need for special-purpose law to clarify the nature of my assistant (and definitely not enough difference to justify shutting down the technology with only vague risk and no instances of social problems introduced by the tech).
Interesting implications there for how Google will measure / improve the effectiveness of Duplex. Where is the line drawn between recording a call and recording data derived from a call? e.g. clearly recording the call unilaterally is illegal. Presumably just storing a hash of the audio data would be legal, but also useless. Is there some middle ground that is legal, but also useful?
We’ve arrived at full Luddism now where life saving and time saving technology in fields like health, transportation, and customer service will be inhibited by hyperbolic fearmongering.
Are you going to pass laws against realistic sounding synthetic voices? Against computers that understand queries “too well”? Against self driving vehicles that drive better and safer than people? All on theoretical harm that actually hasn’t taken place because you’ve watched too many dystopian Netflix sci-fi episodes?
I don’t know what’s more dangerous, the real Skynet, or people who might harm millions by voting for political policies that inhibit real improvements that could be made to help them.
At least, if you want to talk AI being used for harmful things we could discuss feed optimization that parasitizes people’s attention spans and keeps them glued or wasting money on pulling more slot machine levers. Technology that wastes people’s time and money as opposed to things that make people more efficient.
* How much time do people really save, the call in the example is less than a minute. Maybe if you need to call like 10 places it becomes more helpful, but as companies get a bigger online presence (and they are) this technology becomes less useful. You can already make reservations/book appointments, find contracts pretty easily online.
* You can't be 100% sure that the AI won't make a mistake or sound like a total jerk on your behalf. Ok sure, the tech will improve but it will be a long time before humans will fully trust AI to represent you.
* If people find out you're calling them via some automated bot they're going to think you're a tool. Everyone remember Google Glass?
* What the hell is wrong with human interaction anyway?
edit: formatting
And this thing will have to text digest it, it can't simply record the call without identifying it is doing so, and that's destroying the illusion they are crafting.
The tech is extremely impressive, but this use case makes little sense except as a way to humanize it and get people to like it. It has to be more for businesses or some other solution.
I also don't care if someone I probably won't even meet thinks I'm a tool for using bleeding edge tech.
(I assume your friend knows about these services, this is an FYI others.)
https://www.verywell.com/internet-relay-services-1046808#typ...
Part of this is definitely that the girl voice sounds cute, and this partially disables my cognition.
But objectively, her approach is more tentative and polite, whereas the guy voice is more direct and assertive.
They might not want the guy voice to take on those feminine qualities, but it would make the interaction work better - so female AI's dominate.
The effect on the listener may also help - however, I'm not at all sure that other people (especially women) react as I do to the cute voice. They might even find the the Good Doctor male approach better - though I can't imagine that.
The trick of inserting "ums" is very helpful, but because they use the same sound-bite in the same way, it sounds mechanical after you've heard several examples. In the examples towards the end of the page, the odd latencies and (surprising) changes in volume were additionally offpitting.
After a few calls, recipients will recognize the patterns (esp if they use the same voice - can they varying voices convincingly?), and it might be better to have an honest reverse-menu system.
All that said, the first girl voice was great, and there will be progress.
What is the purpose of trying to fool the business owner into thinking it's a real person? It seems unethical, dishonest and disrespectful to the receiver having them believe they are talking to a real person. In the case of an AI failure at least the receiver will understand what's going on instead of becoming really confused. Sometimes I feel people in SV are oblivious to how their software can affect real human beings.
I don't care the awesome technical achievement. The fact that I believe that I am talking to a real person but I am not is the worst.
Can I call the restaurant and say:"well actually my wife doesn't like being cold so I'm not sure the terrace is going to work tonight" and have the computer answer something completely random is just so bad.
The problem is not that the computer doesn't sound natural. The problem is that it cannot deal with out of script requests. In fact the more natural it sounds the more dumb it makes the system appear!
The idea here is to leverage an existent merchant base by adapting to how they work today. Suddenly they just integrated to millions of restaurants by adapting to them and not the other way around.
I'm sure Google Assistant will first check if there is a way to use tech like OpenTable to make the reservation and fall back to a phone call if there is no better alternative.
How bout: "Hey Duplex, call this support number and get a top level human manager on the line please."
Then we get a HN article: "Duplex is fighting Alexa!"
This must be the most convoluted API protocol ever invented.
Many consider Scots to be an actual language separate from English. There's a good amount of debate about this among linguists, I think.
For anyone curious, I recommend reading these pages from the recent Scots translation of the first Harry Potter book to get a feel for how it differs from English.
oops!
In short, I don't have a lot of faith in this being plausible yet. The major reason is the phone lines. Phone call quality is not solid enough to ensure a accuracy rate high enough to roll this out as a production service. There are others, like legal considerations, but if I had to pick one that would be it.
There's a reason why this rose to number one HN, people would clamor for it. To think that Apple and Google haven't been thinking the same thing is short sited.
I honestly don't know how that's going to pan out. Just imagining it already makes me feel the kind of paranoia creep you get when you are too high.
The technology is kinda cool, but when I think about the poor sound quality of some calls and my experience with voice assistants, I wonder in how many cases this will end in just garbage appointments or very poor experiences for the human on the one side of the phone.
Besides that, I like how Google is pushing to change the current way of making appointments. Maybe this will drive more small/medium businesses to use online services for appointments.
I wonder what it sounds like when it runs out of choices, or asked it to get a dinner reservation @ <insert popular place> any evening at any time for the next two months.
The weirdest thing about these trained neural nets too is the small tweaks that break them in very interesting ways. The future is truly a surreal place.
> One of the key research insights was to constrain Duplex to closed domains, which are narrow enough to explore extensively. Duplex can only carry out natural conversations after being deeply trained in such domains. It cannot carry out general conversations.
The question is after you limit it to "closed domains" narrow enough, where it can still be practically useful. It might help with certain functions in enterprise settings. It will definitely work for spammers because they can work with even 1% success rate.
To make this salient for people: imagine this technology being deployed for political robocalls. An attractive voice masquerading as a person persuading people to vote for someone.
They already hire telemarketing centers to do political calls; is this really so different?
Software should be held to the same standard.
Really skeptical about this. And if this does become a thing, it will dumb down the interaction.
I seriously doubt that they will proceed to define and collect them, since those are probably 10% or less of all reservations, but lets say they would.
Then still, the conversation you make to make the reservation is a process in which you make the decision.
Say, there is a place inside at 20:00 or a place in the garden at 20:30. Are you going to let Google choose between the two options for you?
Do you imagine there would be an api in which you specify to the assistant, before it makes a call, your preferences in that much granularity?
If you are going to be more enterprising then you can already do this. Just put some HITs on Mechanical Turk and let them place calls. Should cost you a few bucks to flood hundreds of calls.
It seems like the user is likely to get a certain number of confused recording sent back to them when it fails, and then get stuck manually calling back and explaining what happened.
That's alarming, and thinly cloaked in euphemism (IMO). "Conversation data" here means recordings of actual human-to-human calls, as well as their automated transcriptions. Both were used.
Where did that source audio come from?
Yup, good observation; seems highly likely, save for the 'without even knowing' bit... I would imagine you'd 'request to book' and an async Duplex operation would run in the background and send you a notification of the outcome / possibilities.
What we lose in using human speech for precision we make up in it being pretty much universal. Talk about an adaptable interface. You can phone the restaurant and do anything from reserving a table to ordering takeout to informing them that their cat is on fire.
(I mean that as both a joke and a real comment - you could never force every restaurent in the world to learn REST, but you sure can call a bunch of them)
That being said, as someone who has used "assistant as a service" stuff before, I wonder how well this will work or how limited it will have to be, and not just because of the AI itself.
Even with humans on both sides, it's amazing how hard it can be to get an answer let alone a request fulfilled in a single phone call.
Questions about table placement, food allergies or other restrictions could come up, which I wouldn't want going to an operator. I'd rather be told in advance that it is calling the restaurant and have it send me questions in real-time with suggested answer buttons until it learns enough about me not to need them.
In other cases, just having it call and stay on hold for me would be useful. "Ok Google, call my mobile operator so I can talk to someone at around 10am" and having it use data it already has from other calls to place the call at around 9:45 and patch me in at the right time would be useful.
Given the technological trajectory, over time there will be less need for people who serve purely as the ‘interface’ using minimal skills and knowledge. At the same time, we still need many more people to work in the physical world: cooking nutritious meals, construction, and caring for the elderly are some examples.
Since we cannot assume that everyone can develop skills needed to thrive in demanding technical or knowledge-based jobs, a key priority in many countries should be supporting certain segments of the population to develop the skills and attitude necessary to work in these physical jobs: Most of which are too complex for AI and robotics to effectively replace in the next few decades.
In addition, vocational education should be improved and updated to make use of appropriate technology to increase productivity and reduce physical demand on the body.
[1] https://info.siteselectiongroup.com/blog/how-big-is-the-us-c...
The GIs unmasked them by asking them questions about baseball and shooting any with a wrong answer.
Google Duplex is about helping the long-tail -- it's work that is done by studying the needs and processes of the smallest of small businesses, and tailoring a product just for them.
I've lost count of the number of college hackathon projects where they say, "oh, push a button and you get a pizza" and they think they'll just put an iPad in the kitchen, and then fizzle out when they get to a real restaurant.
In practice, the restaurant might pass around pieces of paper in the kitchen. So you think, oh, I'll put in a thermal receipt printer. But then you realize that they don't have Wi-Fi or internet, so now you have to put in a $70/month internet bill, on top of the phone bill, and a router or two. So you think, "oh, I'll use a fax machine", or "I'll integrate with the point of sale system". But the fax machine runs out of paper, and the point of sale system is an offline piece of ---- running Windows XP. And even if you do get them using an iPad or OpenTable or Yelp or whatever, before you know it, you have waiters writing on a computer monitor with a whiteboard marker: https://javlaskitsystem.se/2012/02/whats-the-waiter-doing-wi...
But every one of these businesses has a telephone number, whether it's a landline or a cell phone or whatever.
When pg says to talk to your customers (https://twitter.com/paulg/status/898476047263518720), he means, talk to your customers. You'll be surprised by what you learn.
(Disclaimer: I work at Google, but on YouTube, not on this product.)
The algorithm of cause doesn't understand all contexts. What troubles me is that, in the first example, should we really give algorithm the freedom to propose a new date for an appointment? It reminds me one thing that particular bothers me with the Gmail's smart reply feature, where when given Monday or Wednesday as options, the suggested reply is, 'How about Tuesday', which does make the conversation flows, but doesn't really make any logical sense.
It makes a good demo, I am very much impressed, however, I feel it will run into a LOT of issues, even only in those provided scenarios, should those scenarios become more sophisticated.
Also, the girl in the last audio, got a flick flirt haha even said, rushing, "see you next friday" (or something) when hanging up.
Once again, my impression.
edit: It came to me that, at some point, it will be able to wander off a little, giggle and stuff. So creepy!
Interviews are dreaded by most of the engineers. What are the chances that Google might be testing this in the wild. Given the number of applications they receive.
(yes, i'm partly serious. Why engage in developing an area of technology where there's a much more elegant and efficient solution?)
Google will most likely want to use recordings to keep fine-tuning and improving upon Duplex, and I don't see them announcing "This call is recorded by Google.", when they're going through such great lengths to convince the called parties that they are talking with a human being.
> Yeah, I'm here
I don't know why, but a machine dynamically saying "I'm here" in a completely naturally sounding way and in a very dynamic context really hits me,
Transparency would mean starting the call by saying "I'm a bot from Google".
For example I have a conversation in both English/Russian and I want to segment the input according to each language then handle each language separately.
That part is as impressive to me as the semantic parsing it's doing on that call.
Other companies are using huge libraries of recorded human voice for communications and concatenating them together in intelligent ways.
And it's a great spin on the Future of Work. It's offloading the time from the consumer onto paid workers at the businesses.
"What appointments does [name] have for the rest of the day?"
Lenny: https://www.youtube.com/watch?v=LgT44DuIaAM&list=PLduL71_GKz...
Jolly Roger: https://www.youtube.com/channel/UC3OxCWLEmoIhNMm-hnvBm9Q
https://www.buzzfeed.com/jwherrman/how-robots-are-stealing-y...
I don't see any details on how this is beyond the research phase.
Is Duplex a product or developer API?
Reason I'm asking is, I am interested in understanding what the intersection of/link between AI/machine learning and quantum computing is, if there is one.
Could it be opt-in on the receiver-side?
Imagine a system designed in the voice of the candidate that can call you and answer most any questions you have about the platform (or log when you don’t have an answer to be updated later), remind you when to vote, send you a Facebook friend request, etc.
Hoping that once they've solved robot calls, they'll probably have a go at some of the harder things like synchronization ;)
I wonder if this sort of technology will result in some sort of arms race / singularity where everyone, businesses and consumers alike, ends up needing to use phone AIs to stay sane.
Combine this with the fact Caller ID is no longer reliable.
I think it's time to replace Signaling System 7 with 21st century technology.
And now making a product out of it.
What do you feel if I step on your foot?
Do roses smell good?
Even if the name is not "rose"?
What's the color of melancholy?
Does Google say anywhere these were all real calls? Or did they call back to cancel the appointments? Because it would be really easy, and tempting, to just fire off 10,000 of these calls to businesses around the country, just to harvest data on how well it does. And leave a massive trail of fake bookings. Even if Google wouldn't do this, the next company attempting this will.
IMHO, to fix this everything needs two factor authentication generated by a biometric scan in person at a government office. Yes you could use a blockchain for it too.
Amazing result. This guys should be very proud.
"To obtain its high precision, we trained Duplex’s RNN on a corpus of anonymized phone conversation data. The network uses the output of Google’s automatic speech recognition (ASR) technology, as well as features from the audio, the history of the conversation, the parameters of the conversation (e.g. the desired service for an appointment, or the current time of day) and more."
"Anonymized!" "Honest!" says the surveillance-capitalism advertising mega corp...
Luckily I've never been able to use Google Voice - but I doubt they're the only threat actor using phone conversations and metadata to train neural nets... Pretend anonymized or not...