Method for multi-channel audio synchronization for task automation
Abstract
A method for coordinating actions between an audio channel and a synchronized non-audio channel includes receiving an indication of a start of a session associated with a user and having an audio channel that is synchronized with a non-audio channel. Thereafter, repeated determinations are made as to whether a prompt on the non-audio channel has been received from the user. In response to each determination that the prompt on the non-audio channel has not been received from the user, a signal is sent to cause an inaudible output on the audio channel to the user. In response to a determination that the prompt on the non-audio channel has been received from the user, an audible output is selected based on an activity by the user on the non-audio channel, and a signal is sent to cause the audible output to be output on the audio channel.
Claims
exact text as granted — not AI-modified1 . An apparatus, comprising:
a processor; and a memory operably coupled to the processor, the memory storing instructions to cause the processor to:
receive an indication of a start of a session associated with a user and having an audio channel that is synchronized with a non-audio channel;
repeatedly determine, after the receiving, whether a prompt on the non-audio channel has been received from the user; and
send a signal to cause the audio channel to remain synchronized with the non-audio channel in response to each determination that the prompt on the non-audio channel has not been received from the user.
2 . The apparatus of claim 1 , wherein the signal to cause the audio channel to remain synchronized with the non-audio channel further causes an inaudible output on the audio channel to the user in response to each determination that the prompt on the non-audio channel has not been received from the user.
3 . The apparatus of claim 1 , wherein the memory further stores instructions to cause the processor to:
in response to a determination that the prompt on the non-audio channel has been received from the user, select an audible output based on an activity by the user on the non-audio channel; select, at a first time, a first language from a plurality of languages; and select, at a second time after the first time, a second language from the plurality of languages, the selecting the audible output being based on the second language.
4 . The apparatus of claim 1 , wherein:
the audio channel is associated with a first device type from a plurality of device types, the non-audio channel is associated with a second device type from the plurality of device types, and the first device type includes a phone, a smart speaker, an earphone or an Internet of Things (IoT) device, the second device type includes a phone, a smart speaker, an earphone or an IoT device, the first device type being different from the second device type.
5 . The apparatus of claim 1 , wherein the memory further stores instructions to cause the processor to:
in response to a determination that the prompt on the non-audio channel has been received from the user, select an audible output based on an activity by the user on the non-audio channel, and receive, via an application programming interface (API), a signal from a device for the non-audio channel, the selecting the audible output being based on the signal from the device for the non-audio channel.
6 . The apparatus of claim 4 , wherein:
during a first time period, the non-audio channel is associated with and the selecting is performed with respect to a first digital non-audio channel, and during a second time period after the first time period, the non-audio channel is associated with and the selecting is performed with respect to a second digital non-audio channel different from the first digital non-audio channel.
7 . The apparatus of claim 1 , wherein the repeatedly determining and the sending the signal is repeated until an end of the session, the method further comprising:
after the start of the session and before the end of the session, performing at least one of:
determine that a prompt on the audio channel received from the user includes an indication that the user would like to discontinue the non-audio channel, or
determine that the prompt on the non-audio channel includes an indication that the user would like to discontinue the non-audio channel;
terminate the non-audio channel of the session, in response to the indication that the user would like to discontinue the non-audio channel; and send, after the terminating, a signal to connect a communication device of the user with a communication device of a live agent.
8 . A non-transitory, processor-readable medium storing instructions to cause a processor to:
initiate a request for a session associated with a user to cause an audio channel associated with the session to synchronize with a non-audio channel associated with the session; repeatedly determine whether a prompt on the non-audio channel has been received from the user; and cause the audio channel to remain active during the session in response to each determination that the prompt on the non-audio channel has not been received from the user.
9 . The non-transitory, processor-readable medium of claim 8 , wherein the instructions to cause the processor to cause the audio channel to remain active during the session further includes instructions to cause the processor to cause an inaudible output on the audio channel to the user in response to each determination that the prompt on the non-audio channel has not been received from the user.
10 . The non-transitory, processor-readable medium of claim 8 , wherein the instructions includes instructions to cause the processor to:
cause an audible output to be output on the audio channel in response to a determination that the prompt on the non-audio channel has been received from the user, the audible output includes a first portion associated with a first voice and a second portion associated with a second voice different than the first voice.
11 . The non-transitory, processor-readable medium of claim 8 , wherein the instructions includes instructions to cause the processor to:
receive an indication from the user to end the session; and cause a connection to a compute device associated with at least one of a live chat or a live agent.
12 . The non-transitory, processor-readable medium of claim 8 , wherein the audio channel is associated with a first compute device, and the non-audio channel is associated with a second compute device different than the first compute device.
13 . The non-transitory, processor-readable medium of claim 8 , wherein:
the initiating of the request, the repeatedly determining, and the causing is performed by a first compute device, and the initiating of the request includes calling, via the first compute device, a phone number associated with a second compute device to cause the second compute device to generate the session.
14 . The non-transitory, processor-readable medium of claim 8 , wherein the initiating of the request, the repeatedly determining, and the causing are performed by a voice assistant device, the instructions includes instructions to cause the processor to:
receive, by the voice assistant device, a voice command from the user that includes an indication of the request, the initiating of the request performed automatically in response to the receiving of the voice command.
15 . A method, comprising:
receiving a representation of a request from a compute device associated with a user to complete a task; causing an audio channel associated with the user to synchronize with at least one non-audio channel associated with the user; sending a first signal to cause a first audible output to be output by the audio channel; repeatedly determining whether a prompt on the at least one non-audio channel has been received from the user; sending a second signal to cause the audio channel to remain open in response to each determination that the prompt on the at least one non-audio channel has not been received from the user; and in response to a determination that the prompt on the at least one non-audio channel has been received from the user:
selecting a second audible output based on a determination that the prompt is in accordance with the task,
selecting a third audible output based on a determination that the prompt is not in accordance with the task, and
sending a third signal to cause one of the second audible output or the third audible output to be output on the audio channel.
16 . The method of claim 15 , wherein the second signal further causes an inaudible output to be output on the audio channel to the user in response to each determination that the prompt on the non-audio channel has not been received from the user.
17 . The method of claim 15 , wherein the prompt is a first prompt, the method further comprising:
repeatedly determining whether a second prompt on the at least one non-audio channel has been received from the user; sending a fourth signal to cause the inaudible output on the audio channel to the user in response to each determination that the second prompt on the at least one non-audio channel has not been received from the user; and in response to the determination that the second prompt on the at least one non-audio channel has been received from the user:
selecting a fourth audible output based on an activity by the user on the at least one non-audio channel, and
sending a fourth signal to cause the fourth audible output to be output on the audio channel.
18 . The method of claim 15 , wherein the compute device is a mobile device, the method further comprising:
transmitting a hyperlink to the mobile device via at least one of a text message or an email, the causing of the audio channel associated with the user to synchronize with the at least one non-audio channel associated with the user performed automatically in response to the user selecting the hyperlink.
19 . The method of claim 15 , wherein:
the audio channel is associated with a first device type from a plurality of device types, the at least one non-audio channel is associated with a second device type from the plurality of device types, the first device type includes a phone, a smart speaker, a speaker, an earphone or an Internet of Things (IoT) device, the second device type includes a phone, a smart speaker, a speaker, an earphone or an IoT device, the first device type is different from the second device type.
20 . The method of claim 15 , wherein the compute device is a first compute device, the method further comprising:
causing a connection to a second compute device associated with at least one of a live chat or a live agent in response to an indication from the user to connect with at least one of the live chat or the live agent.Join the waitlist — get patent alerts
Track US2023274103A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.