Speech processing techniques
Abstract
Techniques for an interactive turn-based reading experience are described. A system may take turns reading content, such as a book, with a user. The system may process audio data representing a user reading a portion of the content, determine reading evaluation data, and determine how to proceed for the next turn based on the reading evaluation data. For example, based on the reading evaluation data, the system may read a portion of the content by outputting synthesized speech representing the content, may ask the user re-read a portion of the content, or may ask the user to read a different, smaller portion of the content.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 - 20 . (canceled)
21 . A computer-implemented method, comprising:
receiving first input audio data corresponding to speech representing a first portion of content; determining that the first input audio data corresponds to an entirety of the first portion of the content; determining, using a first trained machine learning (ML) model, first reading evaluation data based on the first input audio data and the first portion of the content; based at least in part on the first reading evaluation data, determining output data; and causing a device to present the output data.
22 . The computer-implemented method of claim 21 , further comprising:
performing automatic speech processing (ASR) using the first input audio data to determine ASR results data; and receiving first data representing the first portion of content, wherein determining the first reading evaluation data comprises processing the ASR results data and the first data using the first trained ML model.
23 . The computer-implemented method of claim 22 , wherein the ASR results data includes ASR confidence data.
24 . The computer-implemented method of claim 21 , wherein determining the first reading evaluation data comprises:
processing a representation of the first input audio data and a representation of the first portion of content using the first trained ML model to determine the first reading evaluation data, wherein the first reading evaluation data represents a pronunciation accuracy.
25 . The computer-implemented method of claim 21 , wherein:
determining the output data comprises determining text data representing the first reading evaluation data; and causing the device to present the output data comprises causing the device to present, on a display, a representation of the text data.
26 . The computer-implemented method of claim 21 , wherein:
determining the output data comprises, based at least in part on the first reading evaluation data, performing text-to-speech (TTS) processing to generate first output audio data including synthesized speech corresponding to the first portion of the content; and causing the device to present the output data comprises causing the device to output audio corresponding to the first output audio data.
27 . The computer-implemented method of claim 21 , wherein:
determining the output data comprises, based at least in part on the first reading evaluation data, determining a request to reread at least a subset of the first portion of content; and causing the device to present the output data comprises causing the device to present the request.
28 . The computer-implemented method of claim 21 , further comprising:
determining, based at least in part on the first reading evaluation data, a request for the user to read a second portion of content different from the first portion of content, wherein the first portion of content comprises a first number of words and the second portion of content comprises a second number of words, wherein the second number of words is less than the first number of words; and including the request in the output data.
29 . The computer-implemented method of claim 21 , further comprising:
prior to receiving the first input audio data, receiving input data representing a request to read content; receiving data representing the content; enabling a listening mode to capture speech; and disabling the listening mode after determining that the first input audio data corresponds to the entirety of the first portion of the content.
30 . The computer-implemented method of claim 21 , further comprising:
receiving first data representing text of the first portion of the content; and causing the device to present the text on a display.
31 . A system comprising:
at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
receive first input audio data corresponding to speech representing a first portion of content;
determine that the first input audio data corresponds to an entirety of the first portion of the content;
determine, using a first trained machine learning (ML) model, first reading evaluation data based on the first input audio data and the first portion of the content;
based at least in part on the first reading evaluation data, determine output data; and
cause a device to present the output data.
32 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
perform automatic speech processing (ASR) using the first input audio data to determine ASR results data; and receive first data representing the first portion of content, wherein the instructions that cause the system to determine the first reading evaluation data comprise instructions that, when executed by the at least one processor, cause the system to process the ASR results data and the first data using the first trained ML model.
33 . The system of claim 32 , wherein the ASR results data includes ASR confidence data.
34 . The system of claim 31 , wherein the instructions that cause the system to determine the first reading evaluation data comprise instructions that, when executed by the at least one processor, cause the system to:
process a representation of the first input audio data and a representation of the first portion of content using the first trained ML model to determine the first reading evaluation data, wherein the first reading evaluation data represents a pronunciation accuracy.
35 . The system of claim 31 , wherein:
the instructions that cause the system to determine the output data comprise instructions that, when executed by the at least one processor, cause the system to determine text data representing the first reading evaluation data; and the instructions that cause the system to cause the device to present the output data comprise instructions that, when executed by the at least one processor, cause the system to cause the device to present, on a display, a representation of the text data.
36 . The system of claim 31 , wherein:
the instructions that cause the system to determine the output data comprise instructions that, when executed by the at least one processor, cause the system to, based at least in part on the first reading evaluation data, perform text-to-speech (TTS) processing to generate first output audio data including synthesized speech corresponding to the first portion of the content; and the instructions that cause the system to cause the device to present the output data comprise instructions that, when executed by the at least one processor, cause the system to cause the device to output audio corresponding to the first output audio data.
37 . The system of claim 31 , wherein:
the instructions that cause the system to determine the output data comprise instructions that, when executed by the at least one processor, cause the system to, based at least in part on the first reading evaluation data, determine a request to reread at least a subset of the first portion of content; and the instructions that cause the system to cause the device to present the output data comprise instructions that, when executed by the at least one processor, cause the system to cause the device to present the request.
38 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine, based at least in part on the first reading evaluation data, a request for the user to read a second portion of content different from the first portion of content, wherein the first portion of content comprises a first number of words and the second portion of content comprises a second number of words, wherein the second number of words is less than the first number of words; and include the request in the output data.
39 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
prior to receipt of the first input audio data, receive input data representing a request to read content; receive data representing the content; enable a listening mode to capture speech; and disable the listening mode after determining that the first input audio data corresponds to the entirety of the first portion of the content.
40 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
receive first data representing text of the first portion of the content; and cause the device to present the text on a display.Join the waitlist — get patent alerts
Track US2023360633A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.