This week
while I was having a conversation of Amazon Echo over beer, someone mentioned
the reluctance in adopting the technology at this time. "I want to have a conversation with
Alexa. Maybe, even give her a name and voice that I can fall in love with. Just
like I love my vintage car". This was not the Big Bang Theory kind of fall
in love with Siri, it was admiration and passion for the product that worked
with an intelligent human being with a voice.
The
profundity of this statement pushed the boundary of the voice opportunity, much further than my
earlier realization that, people want technology to work for them. Not in a
" I know it all" kind of snobbish way, but genuinely helpful
technology. It is not the efficiency of
how the task is achieved, but the manner in which machines put in the effort to
gracefully and patiently interact with you to get you the information,
accomplish the tasks or just plainly chat with you. It is not intelligence that
people look for, they look for normal.
Artificial
Intelligence
The concept
where the machines will be smart enough to work with us using speech has
several dimensions. In my research on voice, several books point to one simple
fact. We do not completely and accurately know how we-humans acquire
conversation skills. We start the process as babies and by the time we can
experiment non-intrusively, it is too late. Just the thought of any scientist
experimenting with babies would be repulsive to anyone. Therefore we cannot
model the machines based on us, we don’t know how we do it.
Designing for Intelligence
There is a
very large effort behind developing intelligence, building models, trying and
testing them out to convert data into intelligence. We are going for a world
where machines can have intelligent conversations with us.
I believe
that the technology today exists that can be leveraged to design for
intelligence within the limits of current capabilities. At the very least, we
can get a better understanding into how we converse, under various real
circumstances. We can understand what is "Normal" better.
Infant Voice
When we
talk to infants, we use simple bite sized chunks of words, phrases even to
interact. We do so in thousands of
scenarios within the first year of the baby's life. Each of this gives the baby
the skills, tools and data to start with words and graduate to phrases-
mimicking adults in a matter of months.
I believe we are in this stage of the development of the voice
technology.
We would
work with voice better if the voice was that of an infant. One that is eager to
do things for you and please you - just for the opportunity to interact with
you. An infant, that would make mistakes and understands little but whatever
he/she understands, is able to act on with gusto. Who would not think this is
cute? This would perfectly align with what current skills for voice systems.
Not
everyone would like a child's voice but maybe- just maybe we can set the right
expectations with the users. Clearly, this is not available at this time.
Child's Voice
As we learn
more, develop more capabilities, we can frame simple sentences. Crisp, clear,
concise and sometimes insensitive - this is what we are designing current voice
interactions today.
"Alexa,
bring me my paper."
"Hmm,
I 'm not sure what you meant by that question?"
Responding
to an ask beyond ability should be an awkward situation and not something
graceful. We do it all the time. To me this response from Alexa is not a
technology limitation, it is a VUI gap.
A child could respond to this in several ways like
" I
cant do that" or "That sounds hard" or a simple chuckle with a
"Nooooo". Something that will elicit
"Awww,
that’s cute" or "she can't do this" to maybe even "I should
do something to help her do this" the situation most systems gravitate to.
Now, that
to me is being on the way to developing a relationship with technology. Its not
mere upsell and cross sell, it is building a relationship. This might seem like
a sound argument but this option is not available to us.
Adult Voice
Alexa is an
adult, so is Siri, Cortana and Google Now/Assistant. We started with the robotic press # to
continue and n-level deep menu, which put every single human being on the
defensive with voice technology, to Siri a fully grown adult, which would
constantly misunderstand utterings- leading to frustrations. I struggled to get Siri to call my wife and
went out of the way to add a digit to all my contacts whose names started with
a V or a W. and there was one contact in the address book for
"Wife".
As the
voice of an adult, our expectations of technology is very different from our
actual experience. Even though the technology delivers a marvelous advance in
science, we are not impressed- mismatched expectations.
The
Time is NOW.
As the
conversation over beer continued, we discussed how Samsung's latest 2016
generation TV got the TV interface
right. It took us many generations,
co-evolution, iterations and incremental clunky advances to get the user
experience right. The same will be true for Voice. We have to start now, learn
how conversations work and then evolve, adapt, improve and perfect it.
Join
me on this journey of taking technology to everyone- the natural way.
Feedback
and your experience with VUI is really appreciated.
No comments:
Post a Comment