US2023352031A1PendingUtilityA1
Sample-efficient representation learning for real-time latent speaker state characterisation
Est. expiryAug 4, 2040(~14 yrs left)· nominal 20-yr term from priority
G06N 3/0464G06N 3/09G10L 17/18G10L 17/02G06N 3/049G06N 3/08G06N 3/045G06N 3/048G10L 17/08
61
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Systems, methods, and non-transitory computer-readable media can provide audio waveform data that corresponds to a voice sample to a temporal convolutional network for evaluation. The temporal convolutional network can pre-process the audio waveform data and can output an identity embedding associated with the audio waveform data. The identity embedding associated with the voice sample can be obtained from the temporal convolutional network. Information describing a speaker associated with the voice sample can be determined based at least in part on the identity embedding.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A system comprising:
one or more computer processors; one or more computer memories; a set of instructions stored into the one or more computer memories, the set of instructions configuring the one or more processors to perform operations for implementing a neural network architecture for predicting markers for speaker properties or speaker states, the operations comprising: combining a plurality of data sets from a plurality of sources; creating a plurality of trained identity embeddings from the combined plurality of data sets during a first stage of the neural network architecture, the first stage being a preprocessing stage; deriving the markers in a second stage of the neural network architecture, the second stage being a pre-trained marker classification stage applied on top of the preprocessing stage, the deriving of the markers including training an identity model based on the trained identify embeddings; and performing the predicting of the markers using the trained identity model.
3 . The system of claim 2 , wherein the first stage includes training the model using a loss function wherein a baseline input is compared to a positive input and a negative input.
4 . The system of claim 3 , wherein triplets of samples are chosen using a semi-hard triplet mining process.
5 . The system of claim 2 , wherein the second stage includes converting time slices of the plurality of trained identify embeddings of variable lengths into the markers.
6 . The system of claim 2 , wherein the markers are human-readable markers.
7 . The system of claim 2 , wherein the markers pertain to one or more of emotions, arousal level, gender, age, spoken language, native language, or accent.
8 . The system of claim 2 , wherein the speaker states include one or more of laughter, crying, sighing, or coughing.
9 . A method comprising:
combining a plurality of data sets from a plurality of sources; creating a plurality of trained identity embeddings from the combined plurality of data sets during a first stage of the neural network architecture, the first stage being a preprocessing stage; deriving the markers in a second stage of the neural network architecture, the second stage being a pre-trained marker classification stage applied on top of the preprocessing stage, the deriving of the markers including training an identity model based on the trained identify embeddings; and performing the predicting of the markers using the trained identity model.
10 . The method of claim 9 , wherein the first stage includes training the model using a loss function wherein a baseline input is compared to a positive input and a negative input.
11 . The method of claim 10 , wherein triplets of samples are chosen using a semi-hard triplet mining process.
12 . The method of claim 9 , wherein the second stage includes converting time slices of the plurality of trained identify embeddings of variable lengths into the markers.
13 . The method of claim 9 , wherein the markers are human-readable markers.
14 . The method of claim 9 , wherein the markers pertain to one or more of emotions, arousal level, gender, age, spoken language, native language, or accent.
15 . The method of claim 9 , wherein the speaker states include one or more of laughter, crying, sighing, or coughing.
16 . A non-transitory computer-readable storage medium storing a set of instructions that, when executed by one or more computer processors, causes the one or more computer processors to perform operations, the operations comprising:
combining a plurality of data sets from a plurality of sources; creating a plurality of trained identity embeddings from the combined plurality of data sets during a first stage of the neural network architecture, the first stage being a preprocessing stage; deriving the markers in a second stage of the neural network architecture, the second stage being a pre-trained marker classification stage applied on top of the preprocessing stage, the deriving of the markers including training an identity model based on the trained identify embeddings; and performing the predicting of the markers using the trained identity model.
17 . The non-transitory computer-readable storage medium of claim 16 , wherein the first stage includes training the model using a loss function wherein a baseline input is compared to a positive input and a negative input.
18 . The non-transitory computer-readable storage medium of claim 17 , wherein triplets of samples are chosen using a semi-hard triplet mining process.
19 . The non-transitory computer-readable storage medium of claim 16 , wherein the second stage includes converting time slices of the plurality of trained identify embeddings of variable lengths into the markers.
20 . The non-transitory computer-readable storage medium of claim 16 , wherein the markers are human-readable markers.
21 . The non-transitory computer-readable storage medium of claim 16 , wherein the markers pertain to one or more of emotions, arousal level, gender, age, spoken language, native language, or accent.Join the waitlist — get patent alerts
Track US2023352031A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.