Machine-trained network detecting context-sensitive wake expressions for a digital assistant
Abstract
Some embodiments provide a method of training a MT network to detect a wake expression that directs a digital assistant to perform an operation based on a request that follows the expression. The MT network includes processing nodes with configurable parameters. The method iteratively selects different sets of input values with known sets of output values. Each of a first group of input value sets includes a vocative use of the expression. Each of a second group of input value sets includes a non-vocative use of the expression. For each set of input values, the method uses the MT network to process the input set to produce an output value set and computes an error value that expresses an error between the produced output value set and the known output value set. Based on the error values, the method adjusts configurable parameters of the processing nodes of the MT network.
Claims
exact text as granted — not AI-modified1 - 19 . (canceled)
20 . A non-transitory machine readable medium storing a program for execution by at least one hardware processing unit of a digital assistant of a device to direct the digital assistant of the device to perform an operation based on a vocative use of a wake expression, the program comprising sets of instructions for:
capturing an audio input comprising (i) a vocative use of a wake expression associated with the digital assistant and (ii) a set of one or more instructions for the device; processing the audio input with a machine-trained network comprising a plurality of layers of processing nodes that are trained through machine learning to differentiate vocative uses of the wake expression from non-vocative uses of the wake expression, wherein vocative uses of the wake expression include a name associated with the digital assistant with a tonal inflection of the name that is associated with an invocation of the digital assistant; based on an output of the machine-trained network specifying that the audio input includes the vocative use of the wake expression, directing the digital assistant of the device to perform an operation associated with the set of one or more instructions in the audio input.
21 . The non-transitory machine readable medium of claim 20 , wherein the non-vocative uses of the wake expression comprise at least one of an ablative use of the wake expression, a dative use of the wake expression, and a genitive use of the wake expression.
22 . The non-transitory machine readable medium of claim 20 , wherein input value sets used to train the machine-trained network include first and second groups that comprise different prosodic utterances of the wake expression, with the input value sets of the first group including prosodic utterances of the wake expression associated with the vocative use of the wake expression while the input value sets of the second group include prosodic utterances of the wake expression associated with the non-vocative use of the wake expression.
23 . The non-transitory machine readable medium of claim 22 , wherein the different prosodic utterances of the first and second groups differentiate prosodic utterances associated with the vocative use of the expression from prosodic utterances associated with at least one of an ablative use of the wake expression, a dative use of the wake expression, and a genitive use of the wake expression.
24 . The non-transitory machine readable medium of claim 20 , wherein the vocative use further includes the use of the wake expression in a particular sentence structure.
25 . The non-transitory machine readable medium of claim 20 , wherein the vocative uses further include a plurality of different uses of the wake expression in a plurality of different sentence structures.
26 . The non-transitory machine readable medium of claim 20 , wherein the machine learning trains the processing nodes by using a plurality of input/output pairs of a plurality of training sets that are selected to differentiate a plurality of different vocative uses of the wake expression from a plurality of different non-vocative uses of the wake expression.
27 . The non-transitory machine readable medium of claim 26 , wherein the input/output pairs are selected to differentiate syntactical components and tonal components in the plurality of different vocative uses of the wake expression from the syntactical and tonal components of the plurality of different non-vocative uses of the wake expression.
28 . The non-transitory machine readable medium of claim 26 , wherein the machine learning trains the machine-trained network to detect vocative uses of the wake expressions while ignoring the non-vocative uses of the wake expressions.
29 . The non-transitory machine readable medium of claim 20 , wherein the machine learning trains the processing nodes by using grammar and prosody of uses of wake expressions to differentiate vocative uses of the wake expression from the non-vocative uses of the wake expression.
30 . The non-transitory machine readable medium of claim 20 , wherein the audio input is a first audio input, the method further comprising:
capturing a second audio input that includes a non-vocative use of the wake expression; processing the second audio input with the machine-trained network to determine that the second audio input does not include the vocative use of the wake expression; and discarding the second audio input without directing the digital assistant to perform an operation based on any input that follows the second audio input.
31 . The non-transitory machine readable medium of claim 20 , wherein the machine trained network comprises a recurrent neural network.
32 . The non-transitory machine readable medium of claim 20 , wherein the machine trained network comprises an LSTM (long short term memory) network.
33 . The non-transitory machine readable medium of claim 31 , wherein the input value sets in the first and second groups use the wake expression differently in different syntactical sentence structures associated with the vocative uses of the wake expression and the non-vocative uses of the wake expression.
34 . The non-transitory machine readable medium of claim 33 , wherein the different syntactical sentence structures of the first and second groups differentiate syntactical sentence structures associated with the vocative uses of the wake expression from syntactical sentence structures associated with at least one of an ablative use of the wake expression, a dative use of the wake expression, and a genitive use of the wake expression.
35 . A method of operating a digital assistant of a device, the method comprising:
capturing an audio input comprising (i) a vocative use of a wake expression associated with the digital assistant and (ii) a set of one or more instructions for the device; processing the audio input with a machine-trained network comprising a plurality of layers of processing nodes that are trained through machine learning to differentiate vocative uses of the wake expression from non-vocative uses of the wake expression, wherein vocative uses of the wake expression include a name associated with the digital assistant with a tonal inflection of the name that is associated with an invocation of the digital assistant; based on an output of the machine-trained network specifying that the audio input includes the vocative use of the wake expression, determining an operation for the device to perform based on the set of one or more instructions in the audio input; and after determining the operation, directing the digital assistant of the device to perform the operation.
36 . The method of claim 35 , wherein the vocative use further includes the use of the wake expression in a particular sentence structure.
37 . The method of claim 35 , wherein the non-vocative uses of the wake expression comprise at least one of an ablative use of the wake expression, a dative use of the wake expression, and a genitive use of the wake expression.
38 . The method of claim 35 , wherein the input/output pairs of the plurality of training sets include first and second groups that comprise different prosodic utterances of the wake expression, with the input/output pairs of the first group including prosodic utterances of the wake expression associated with the vocative use of the wake expression while the input value sets of the second group include prosodic utterances of the wake expression associated with the non-vocative use of the wake expression.
39 . The method of claim 38 , wherein the different prosodic utterances of the first and second groups differentiate prosodic utterances associated with the vocative uses of the wake expression from prosodic utterances associated with at least one of an ablative use of the wake expression, a dative use of the wake expression, and a genitive use of the wake expression.Join the waitlist — get patent alerts
Track US2023419955A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.