Voice profile updating
Abstract
Techniques for updating voice profiles used to perform user recognition are described. A system may use clustering techniques to update voice profiles. When the system receives audio data representing a spoken user input, the system may store the audio data. Periodically, the system may recall, from storage, audio data (representing previous user inputs). The system may identify clusters of the audio data, with each cluster including similar or identical speech characteristics. The system may determine a cluster is substantially similar to an existing voice profile. If this occurs, the system may create an updated voice profile using the original voice profile and the cluster of audio data.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . A computer-implemented method, comprising:
receiving first audio data representing a voice of a first user; processing the first audio data to determine a voice profile corresponding to the first user; receiving second audio data; processing the second audio data to determine the second audio data represents the voice of the first user; and in response to the second audio data representing the voice of the first user, processing the second audio data and the voice profile to determine an updated voice profile corresponding to the first user.
22 . The computer-implemented method of claim 21 , wherein processing the second audio data to determine the second audio data represents the voice of the first user comprises processing the second audio data with respect to the voice profile.
23 . The computer-implemented method of claim 21 , further comprising:
determining the second audio data corresponds to a first speech characteristic; and determining the voice profile does not represent the first speech characteristic.
24 . The computer-implemented method of claim 21 , further comprising:
performing speech processing using the second audio data to determine output data responsive to a command represented in the second audio data; and sending the output data to a device associated with the second audio data.
25 . The computer-implemented method of claim 21 , further comprising:
determining the second audio data represents a wakeword; and in response to the second audio data representing a wakeword, using the second audio data to determine the updated voice profile.
26 . The computer-implemented method of claim 21 , further comprising:
determining a cluster of data corresponding to voice inputs of the first user; and determining an updated cluster of data corresponding to voice inputs of the first user, wherein the updated cluster of data includes a representation of the second audio data.
27 . The computer-implemented method of claim 21 , further comprising:
determining user verification information corresponding to the second audio data; and determining the user verification information corresponds to the first user, wherein determining the second audio data represents the voice of the first user is based at least in part on the user verification information.
28 . The computer-implemented method of claim 21 , wherein the second audio data is received from a first device and the method further comprises:
determining a plurality of voice profiles corresponding to the first device, the plurality of voice profiles including the voice profile corresponding to the first user and a second voice profile corresponding to a second user; and determining the second audio data more closely corresponds to the voice profile than the second voice profile.
29 . The computer-implemented method of claim 21 , wherein processing the second audio data to determine the second audio data represents the voice of the first user comprises:
determining first audio characteristic data corresponding to the second audio data; determining second audio characteristic data corresponding to the voice profile; and determining the first audio characteristic data is similar to the second audio characteristic data.
30 . The computer-implemented method of claim 21 , wherein processing the second audio data to determine the second audio data represents the voice of the first user comprises:
generating a first feature vector representing the second audio data; processing the first feature vector with respect to a second feature vector representing the voice profile; and determining the first feature vector is similar to the second feature vector.
31 . A system comprising:
at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
receive first audio data representing a voice of a first user;
process the first audio data to determine a voice profile corresponding to the first user;
receive second audio data;
process the second audio data to determine the second audio data represents the voice of the first user; and
in response to the second audio data representing the voice of the first user, process the second audio data and the voice profile to determine an updated voice profile corresponding to the first user.
32 . The system of claim 31 , wherein the instructions that cause the system to process the second audio data to determine the second audio data represents the voice of the first user comprise instructions that, when executed by the at least one processor, further cause the system to process the second audio data with respect to the voice profile.
33 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine the second audio data corresponds to a first speech characteristic; and determine the voice profile does not represent the first speech characteristic.
34 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
performing speech processing using the second audio data to determine output data responsive to a command represented in the second audio data; and send the output data to a device associated with the second audio data.
35 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine the second audio data represents a wakeword; and in response to the second audio data representing a wakeword, use the second audio data to determine the updated voice profile.
36 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine a cluster of data corresponding to voice inputs of the first user; and determine an updated cluster of data corresponding to voice inputs of the first user, wherein the updated cluster of data includes a representation of the second audio data.
37 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine user verification information corresponding to the second audio data; and determine the user verification information corresponds to the first user, wherein determination that the second audio data represents the voice of the first user is based at least in part on the user verification information.
38 . The system of claim 31 , wherein the second audio data is received from a first device and wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine a plurality of voice profiles corresponding to the first device, the plurality of voice profiles including the voice profile corresponding to the first user and a second voice profile corresponding to a second user; and determine the second audio data more closely corresponds to the voice profile than the second voice profile.
39 . The system of claim 31 , wherein the instructions that process the second audio data to determine the second audio data represents the voice of the first user comprise instructions that, when executed by the at least one processor, further cause the system to:
determine first audio characteristic data corresponding to the second audio data; determine second audio characteristic data corresponding to the voice profile; and determine the first audio characteristic data is similar to the second audio characteristic data.
40 . The system of claim 31 , wherein the instructions that process the second audio data to determine the second audio data represents the voice of the first user comprise instructions that, when executed by the at least one processor, further cause the system to:
generate a first feature vector representing the second audio data; process the first feature vector with respect to a second feature vector representing the voice profile; and determine the first feature vector is similar to the second feature vector.Join the waitlist — get patent alerts
Track US2021304774A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.