Determining speaker changes in audio input
Abstract
Intelligent assistant systems, methods and computing devices are disclosed for identifying a speaker change. A method comprises receiving audio input comprising a speech fragment. A first voice model is trained with a first sub-fragment from the speech fragment. A second voice model is trained with a second sub-fragment from the speech fragment. The first sub-fragment is analyzed with the second voice model to yield a first confidence value. The second sub-fragment is analyzed with the first voice model to yield a second confidence value. Based at least on the first and second confidence values, the method determines if a speaker of the first sub-fragment is the speaker of the second sub-fragment.
Claims
exact text as granted — not AI-modified1 . An intelligent digital assistant system, comprising:
a logic processor; and a storage device holding instructions executable by the logic processor to:
receive audio input comprising a speech fragment;
train a first voice model with a first sub-fragment from the speech fragment;
train a second voice model with a second sub-fragment from the speech fragment;
analyze the first sub-fragment with the second voice model to yield a first confidence value;
analyze the second sub-fragment with the first voice model to yield a second confidence value; and
based at least on the first confidence value and the second confidence value, determine if a speaker of the first sub-fragment is the speaker of the second sub-fragment.
2 . The intelligent digital assistant system of claim of 1 , wherein the instructions are executable to, based at least on determining that the speaker of the first sub-fragment is the speaker of the second sub-fragment, utilize at least the first sub-fragment and the second sub-fragment to determine a user intent of the speaker.
3 . The intelligent digital assistant system of claim of 1 , wherein the instructions are executable to, based at least on determining that the speaker of the first sub-fragment is not the speaker of the second sub-fragment, utilize at least the first sub-fragment and forego utilizing the second sub-fragment to determine a user intent of the speaker of the first sub-fragment.
4 . The intelligent digital assistant system of claim 1 , wherein the instructions are executable to:
generate the first voice model and second voice model from a universal background model; and based at least on determining that the speaker of the first sub-fragment is the speaker of the second sub-fragment, update the universal background model to an updated universal background model using the first sub-fragment and the second sub-fragment.
5 . The intelligent digital assistant system of claim of 4 , wherein the instructions are executable to generate a third voice model from the updated universal background model by training the updated universal background model with another sub-fragment of speech.
6 . The intelligent digital assistant system of claim 1 , wherein the first sub-fragment and the second sub-fragment have unequal temporal lengths.
7 . The intelligent digital assistant system of claim 1 , wherein the instructions are executable to, based at least on the first confidence value and the second confidence value exceeding a predetermined threshold, determine that the speaker of the first sub-fragment is the speaker of the second sub-fragment.
8 . The intelligent digital assistant system of claim 1 , wherein the instructions are executable to, based at least on the first confidence value and the second confidence value being less than or equal to a predetermined threshold, determine that the speaker of the first sub-fragment is not the speaker of the second sub-fragment.
9 . The intelligent digital assistant system of claim 1 , wherein determining if the speaker of the first sub-fragment is the speaker of the second sub-fragment comprises:
computing an average of the first confidence value and the second confidence value; and if the average exceeds a predetermined threshold, then determining that the speaker of the first sub-fragment is the speaker of the second sub-fragment.
10 . At a computing device, a method for identifying a speaker change, the method comprising:
receiving audio input comprising a speech fragment; training a first voice model with a first sub-fragment from the speech fragment; training a second voice model with a second sub-fragment from the speech fragment; analyzing the first sub-fragment with the second voice model to yield a first confidence value; analyzing the second sub-fragment with the first voice model to yield a second confidence value; and based at least on the first confidence value and the second confidence value, determining if a speaker of the first sub-fragment is the speaker of the second sub-fragment.
11 . The method of claim 10 , further comprising, based at least on determining that the speaker of the first sub-fragment is the speaker of the second sub-fragment, utilizing at least the first sub-fragment and the second sub-fragment to determine a user intent of the speaker.
12 . The method of claim 10 , further comprising, based at least on determining that the speaker of the first sub-fragment is not the speaker of the second sub-fragment, utilizing at least the first sub-fragment and foregoing utilizing the second sub-fragment to determine a user intent of the speaker of the first sub-fragment.
13 . The method of claim 10 , further comprising:
generating the first voice model and second voice model from a universal background model; and based at least on determining that the speaker of the first sub-fragment is the speaker of the second sub-fragment, updating the universal background model to an updated universal background model using the first sub-fragment and the second sub-fragment.
14 . The method of claim 13 , further comprising generating a third voice model from the updated universal background model by training the updated universal background model with another sub-fragment of speech.
15 . The method of claim 10 , wherein the first sub-fragment and the second sub-fragment have unequal temporal lengths.
16 . The method of claim 10 , further comprising, based at least on the first confidence value and the second confidence value exceeding a predetermined threshold, determining that the speaker of the first sub-fragment is the speaker of the second sub-fragment.
17 . The method of claim 10 , further comprising, based at least on the first confidence value and the second confidence value being less than or equal to a predetermined threshold, determining that the speaker of the first sub-fragment is not the speaker of the second sub-fragment.
18 . The method of claim 10 , wherein determining if the speaker of the first sub-fragment is the speaker of the second sub-fragment comprises:
computing an average of the first confidence value and the second confidence value; and if the average exceeds a predetermined threshold, then determining that the speaker of the first sub-fragment is the speaker of the second sub-fragment.
19 . A computing device, comprising:
at least one microphone; a logic processor; and a storage device holding instructions executable by the logic processor to:
via the at least one microphone, receive audio input comprising a speech fragment;
generate a first sub-fragment and a second sub-fragment from the speech fragment;
train a first voice model with the first sub-fragment;
train a second voice model with the second sub-fragment;
analyze the first sub-fragment with the second voice model to yield a first confidence value;
analyze the second sub-fragment with the first voice model to yield a second confidence value; and
based at least on the first confidence value and the second confidence value, determine if a speaker of the first sub-fragment is the speaker of the second sub-fragment.
20 . The computing device of claim 19 , wherein the instructions are executable to, based at least on determining that the speaker of the first sub-fragment is the speaker of the second sub-fragment, utilize at least the first sub-fragment and the second sub-fragment to determine a user intent of the speaker.Join the waitlist — get patent alerts
Track US2018233140A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.