Combining device or assistant-specific hotwords in a single utterance
Abstract
A method for combining hotwords in a single utterance receives, at a first assistant-enabled device (AED), audio data corresponding to an utterance directed toward the first AED and a second AED among two or more AEDs where the audio data includes a query specifying an operation to perform. The method also detects, using a hotword detector, a first hotword assigned to the first AED that is different than a second hotword assigned to the second AED In response to detecting the first hotword, the method initiates processing on the audio data to determine that the audio data includes a term preceding the query that at least partially matches the second hotword assigned. Based on the at least partial match, the method executes a collaboration routine to cause the first AED and the second AED to collaborate with one another to fulfill the query.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A method comprising:
receiving, at data processing hardware of a first assistant-enabled device (AED), audio data corresponding to an utterance spoken by the user and directed toward the first AED and a second AED among two or more AEDs associated with the user, the audio data comprising a single query directed to both the first AED and the second AED that specifies an operation for both the first AED and the second AED to perform at the same time;
detecting, by the data processing hardware, using a hotword detection model running on the first AED, a first hotword in the audio data that proceeds the single query without hotword detection model detecting a second hotword in the audio data, the first hotword assigned to the first AED and different than the second hotword assigned to the second AED;
in response to detecting the first hotword assigned to the first AED in the audio data, initiating, by the data processing hardware, processing on the audio data to determine that the audio data comprises one or more terms preceding the single query that at least partially match the second hotword assigned to the second AED; and
based on detecting the first hotword in the audio data that precedes the single query and the determination that the audio data comprises the one or more terms preceding the single query that at least partially match the second hotword, executing, by the data processing hardware, a collaboration routine to cause the first AED and the second AED to collaborate with one another to fulfill performance of the operation specified by the single query directed to both the first AED and the second AED.
2. The method of claim 1 , wherein initiating processing on the audio data in response to determining that the audio data includes the first hotword comprises:
instructing a speech recognizer to perform speech recognition on the audio data to generate a speech recognition result for the audio data; and
determining, using the speech recognition result for the audio data, the one or more terms that at least partially match the second hotword are recognized in the audio data.
3. The method of claim 2 , wherein instructing the speech recognizer to perform speech recognition on the audio data comprises one of:
instructing a server-side speech recognizer to perform speech recognition on the audio data; or
instructing the speech recognizer to execute on the data processing hardware of the first AED to perform speech recognition on the audio data.
4. The method of claim 1 , wherein determining that the audio data comprises the one or more terms preceding the single query that at least partially match the second hotword assigned to the second AED comprises:
accessing a hotword registry containing a respective list of one or more hotwords assigned to each of the two or more AEDs associated with the user; and
recognizing the one or more terms in the audio data that match or partially match the second hotword in the respective list of one or more hotwords assigned to the second AED.
5. The method of claim 4 , wherein:
the respective list of one or more hotwords assigned to each of the two or more AEDs in the hotword registry further comprises one or more variants associated with each hotword; and
determining that the audio data comprises the one or more terms preceding the single query that at least partially match the second hotword comprises determining that the one or more terms recognized in the audio data match one of the one or more variants associated with the second hotword.
6. The method of claim 4 , wherein the hotword registry is stored on at least one of:
the first AED;
the second AED;
a third AED among the two or more AEDs associated with the user; or
a server in communication with the two or more AEDs associated with the user.
7. The method of claim 1 , wherein determining that the audio data comprises the one or more terms preceding the single query that at least partially match the second hotword comprises providing the audio data as input to a machine learning model trained to determine a likelihood of whether a user intended to speak the second hotword assigned to the user device.
8. The method of claim 1 , wherein, when the one or more terms in the audio data preceding the single query only partially match the second hotword, executing the collaboration routine causes the first AED to invoke the second AED to wake-up and collaborate with the first AED to fulfill performance of the operation specified by the single query.
9. The method of claim 1 , wherein, during execution of the collaboration routine, the first AED and the second AED collaborate with one another by designating one of the first AED or the second AED to:
generate a speech recognition result for the audio data;
perform query interpretation on the speech recognition result to determine that the speech recognition result identifies the single query specifying the operation to perform; and
share the query interpretation performed on the speech recognition result with the other one of the first AED or the second AED.
10. The method of claim 1 , wherein, during execution of the collaboration routine, the first AED and the second AED collaborate with one another by each independently:
generating a speech recognition result for the audio data; and
performing query interpretation on the speech recognition result to determine that the speech recognition result identifies the single query specifying the operation to perform.
11. The method of claim 1 , wherein:
the operation specified by the single query comprises a device-level operation to perform on each of the first AED and the second AED; and
during execution of the collaboration routine, the first AED and the second AED collaborate with one another by fulfilling performance of the device-level operation independently.
12. The method of claim 1 , wherein:
the single query specifying the operation to perform comprises a single query for the first AED and the second AED to both perform a long-standing operation at the same time; and
during executing of the collaboration routine, the first AED and the second AED collaborate with one another by:
pairing with one another for a duration of the long-standing operation; and
coordinating performance of sub-actions related to the long-standing operation between first AED and the second AED to perform.
13. A first assistant-enabled (AED) device comprising:
data processing hardware; and
memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
receiving audio data corresponding to an utterance spoken by the user and directed toward the first AED and a second AED among two or more AEDs associated with the user, the audio data comprising a single query directed toward both the first AED and the second AED that specifies an operation for both the first AED and the second AED to perform at the same time;
detecting, using a hotword detection model running on the first AED, a first hotword in the audio data that proceeds the single query without hotword detection model detecting a second hotword in the audio data, the first hotword assigned to the first AED and different than the second hotword assigned to the second AED;
in response to detecting the first hotword assigned to the first AED in the audio data, initiating processing on the audio data to determine that the audio data comprises one or more terms preceding the single query that at least partially match the second hotword assigned to the second AED; and
based on detecting the first hotword in the audio data that precedes the single query and the determination that the audio data comprises the one or more terms preceding the single query that at least partially match the second hotword, executing a collaboration routine to cause the first AED and the second AED to collaborate with one another to fulfill performance of the operation specified by the single query directed to both the first AED and the second AED.
14. The device of claim 13 , wherein initiating processing on the audio data in response to determining that the audio data includes the first hotword comprises:
instructing a speech recognizer to perform speech recognition on the audio data to generate a speech recognition result for the audio data; and
determining, using the speech recognition result for the audio data, the one or more terms that at least partially match the second hotword are recognized in the audio data.
15. The device of claim 14 , wherein instructing the speech recognizer to perform speech recognition on the audio data comprises one of:
instructing a server-side speech recognizer to perform speech recognition on the audio data; or
instructing the speech recognizer to execute on the data processing hardware of the first AED to perform speech recognition on the audio data.
16. The device of claim 13 , wherein determining that the audio data comprises the one or more terms preceding the single query that at least partially match the second hotword assigned to the second AED comprises:
accessing a hotword registry containing a respective list of one or more hotwords assigned to each of the two or more AEDs associated with the user; and
recognizing the one or more terms in the audio data that match or partially match the second hotword in the respective list of one or more hotwords assigned to the second AED.
17. The device of claim 16 , wherein:
the respective list of one or more hotwords assigned to each of the two or more AEDs in the hotword registry further comprises one or more variants associated with each hotword; and
determining that the audio data comprises the one or more terms preceding the single query that at least partially match the second hotword comprises determining that the one or more terms recognized in the audio data match one of the one or more variants associated with the second hotword.
18. The device of claim 16 , wherein the hotword registry is stored on at least one of:
the first AED;
the second AED;
a third AED among the two or more AEDs associated with the user; or
a server in communication with the two or more AEDs associated with the user.
19. The device of claim 13 , wherein determining that the audio data comprises the one or more terms preceding the single query that at least partially match the second hotword comprises providing the audio data as input to a machine learning model trained to determine a likelihood of whether a user intended to speak the second hotword assigned to the user device.
20. The device of claim 13 , wherein, when the one or more terms in the audio data preceding the single query only partially match the second hotword, executing the collaboration routine causes the first AED to invoke the second AED to wake-up and collaborate with the first AED to fulfill performance of the operation specified by the single query.
21. The device of claim 13 , wherein, during execution of the collaboration routine, the first AED and the second AED collaborate with one another by designating one of the first AED or the second AED to:
generate a speech recognition result for the audio data;
perform query interpretation on the speech recognition result to determine that the speech recognition result identifies the single query specifying the operation to perform; and
share the query interpretation performed on the speech recognition result with the other one of the first AED or the second AED.
22. The device of claim 13 , wherein, during execution of the collaboration routine, the first AED and the second AED collaborate with one another by each independently:
generating a speech recognition result for the audio data; and
performing query interpretation on the speech recognition result to determine that the speech recognition result identifies the single query specifying the operation to perform.
23. The device of claim 13 , wherein:
the operation specified by the single query comprises a device-level operation to perform on each of the first AED and the second AED; and
during execution of the collaboration routine, the first AED and the second AED collaborate with one another by fulfilling performance of the device-level operation independently.
24. The device of claim 13 , wherein:
the single query specifying the operation to perform comprises a single query for the first AED and the second AED to both perform a long-standing operation at the same time; and
during executing of the collaboration routine, the first AED and the second AED collaborate with one another by:
pairing with one another for a duration of the long-standing operation; and
coordinating performance of sub-actions related to the long-standing operation between first AED and the second AED to perform.Join the waitlist — get patent alerts
Track US11948565B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.