US2016183867A1PendingUtilityA1
Method and system for online and remote speech disorders therapy
Est. expiryDec 31, 2034(~8.4 yrs left)· nominal 20-yr term from priority
H04L 65/1069A61B 5/486A61B 5/7465G10L 25/66G16H 20/40H04L 67/10G16H 40/63G09B 5/02G09B 19/04A61B 5/0022G09B 7/00A61B 5/4803A61B 5/742A61B 5/7282A61B 5/0077A61B 7/003A61B 2560/0475A61B 2560/0223G16Z 99/00G16H 40/67G16H 20/30
38
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A method and device for enabling remote speech disorder therapy are presented. The method includes setting a first device with at least one exercise to be performed during a current therapy session, wherein each exercise includes at least a difficulty parameter; receiving a voice production of a user of the first device; processing the received voice production to evaluate a correct execution of the voice production respective of the at least one difficulty parameter; generating a feedback based on the analysis; and outputting the generated feedback to the first device.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for enabling remote speech disorder therapy, comprising:
setting a first device with at least one exercise to be performed during a current therapy session, wherein each exercise includes at least a difficulty parameter; receiving a voice production of a user of the first device; processing the received voice production to evaluate a correct execution of the voice production respective of the at least one difficulty parameter; generating a feedback based on the analysis; and outputting the generated feedback to the first device.
2 . The method of claim 1 , further comprising:
establishing a network communication channel between the first device and a second device; and outputting the generated feedback to the second device, thereby enabling a user of the second device to remotely monitor the execution of the at least one exercise.
3 . The method of claim 2 , further comprising:
receiving instructions from the second device, wherein the instructions include at least one of: a video stream, a video clip, a text file, an image, and an audio clip.
4 . The method of claim 2 , wherein the user of the first device is a patient and the user of the second device is a therapist.
5 . The method of claim 1 , wherein the generated feedback is at least a visual feedback.
6 . The method of claim 5 , further comprising:
rendering a target template respective of the at least one exercise and the received voice production; and displaying the target template at least on the first device corresponding to the received voice production.
7 . The method of claim 6 , wherein the displayed target template includes at least one of: a start boundary, a finish boundary, and a top boundary.
8 . The method of claim 7 , wherein the target template and at least the start boundary are displayed as the voice production is received.
9 . The method of claim 5 , wherein generating feedback based on the analysis further comprises:
coloring the voice production using at least a first color and a second color, wherein the first color represents a loud sound produced by the user and the second color represents a soft sound produced by the user; and displaying at least one of: a positive indication upon performing a correct execution, and an instructive indication upon performing an incorrect execution.
10 . The method of claim 5 , wherein the at least one exercise includes a sequence having a plurality of target templates that require the user to produce a sequence of voice productions.
11 . The method of claim 10 , further comprising:
providing a breathing indicator, wherein the breathing indicator represents a duration of time that the user needs to breathe before trying a subsequent target template, wherein the duration of time is determined based on the at least one difficulty parameter.
12 . The method of claim 1 , further comprising:
measuring a speech rate respective of the analysis; and displaying a speech-rate meter respective of the measured speech rate.
13 . The method of claim 1 , further comprising:
performing an audio calibration process for the first user device, wherein the audio calibration process provides at least a normal speech energy level, a silence energy level, and a calibration energy level.
14 . The method of claim 13 , wherein processing the received voice production further comprises:
sampling the received voice production to create voice samples; buffering the voice samples to create voice chunks; converting the voice chunks from a time domain to a frequency domain; extracting spectrum features from each of the frequency domain voice chunks, wherein the spectrum features include at least dominant frequencies, wherein each dominant frequency corresponds to a voice chunk; computing, for each voice chunk, the energy level of the corresponding dominant frequency; and determining, for each voice chunk, an energy level of the voice chunk based on the energy level of the corresponding dominant frequency.
15 . The method of claim 14 , further comprising:
determining a correctness of the execution of the voice production based on the energy levels of the voice chunks and at least one of: the normal speech energy level, the silence energy level, and the calibration energy level.
16 . The method of claim 15 , wherein the correctness determination results in at least one error related to an incorrect execution of the voice production, wherein each error is any of: a gentle onset, a soft peak, a gentle offset, a volume control, a pattern usage, a miss of a subsequent voice production, an asymmetry of the voice production, a short inhale, a slow voice production, a fast voice production, a short voice production, a long voice production, and an intense peak voice production.
17 . The method of claim 1 , wherein the at least one exercise is related to fluency shaping.
18 . The method of claim 17 , wherein the at least one exercise is related to customized content.
19 . The method of claim 1 , further comprising:
generating a reporting summarizing the execution of the voice production throughout the current therapy session; and saving the report.
20 . The method of claim 1 , wherein the speech disorder therapy is for at least one of:
stuttering, cluttering, and diction.
21 . A non-transitory computer readable medium having stored thereon instructions for causing one or more processing units to execute the method according to claim 1 .
22 . A device for enabling remote speech disorder therapy, comprising:
an interface for receiving a voice production of a user of a first user device; a processing unit; and a memory coupled to the processing unit, the memory containing instructions that, when executed by the processing unit, configure the device to: set the first user device with at least one exercise to be performed during a current therapy session, wherein each exercise includes at least a difficulty parameter; receive the voice production of the user of the first device; analyze the received voice production to evaluate a correct execution of the voice production respective of the at least one difficulty parameter; generate a feedback respective of the analysis; and output the generated feedback to the first device.
23 . The device of claim 22 , wherein the device is further configured to:
establish a network communication channel between the first device and a second device; and output the generated feedback to the second device, thereby enabling a user of the second device to remotely monitor the execution of the at least one exercise.
24 . The device of claim 23 , wherein the device is further configured to:
receive instructions from the second device, wherein the instructions include at least one of: a video stream, a video clip, a text file, an image, and an audio clip.
25 . The device of claim 23 , wherein the user of the first device is a patient and the user of the second device is a therapist.
26 . The device of claim 22 , wherein the generated feedback is at least a visual feedback.
27 . The device of claim 26 , wherein the device is further configured to:
render a target template respective of the at least one exercise and the received voice production; and display the target template at least on the first device corresponding to the received voice production.
28 . The device of claim 27 , wherein the displayed target template includes at least one of: a start boundary, a finish boundary, and a top boundary.
29 . The device of claim 28 , wherein the target template and at least the start boundary are displayed as the voice production is received.
30 . The device of claim 26 , wherein the device is further configured to:
color the voice production using at least a first color and a second color, wherein the first color represents a loud sound produced by the user and the second color represents a soft sound produced by the user; and display at least one of: a positive indication upon performing a correct execution, and an instructive indication upon performing an incorrect execution.
31 . The device of claim 26 , wherein the at least one exercise includes a sequence having a plurality of target templates that require the user to produce a sequence of voice productions.
32 . The device of claim 31 , wherein the device is further configured to:
provide a breathing indicator, wherein the breathing indicator represents a duration of time that the user needs to breathe before trying a subsequent target template, wherein the duration of time is determined based on the at least one difficulty parameter.
33 . The device of claim 22 , wherein the device is further configured to:
measure a speech rate respective of the analysis; and display a speech-rate meter respective of the measured speech rate.
34 . The device of claim 22 , wherein the device is further configured to:
perform an audio calibration process for the first user device, wherein the audio calibration process provides at least a normal speech energy level, a silence energy level, and a calibration energy level.
35 . The device of claim 34 , wherein the device is further configured to:
sample the received voice production to create voice samples; buffer the voice samples to create voice chunks; convert the voice chunks from a time domain to a frequency domain; extract spectrum features from each of the frequency domain voice chunks, wherein the spectrum features include at least dominant frequencies, wherein each dominant frequency corresponds to a voice chunk; compute, for each voice chunk, the energy level of the corresponding dominant frequency; and determine, for each voice chunk, an energy level of the voice chunk based on the energy level of the corresponding dominant frequency.
36 . The device of claim 35 , wherein the device is further configured to:
determine a correctness of the execution of the voice production based on the energy levels of the voice chunks and at least one of: the normal speech energy level, the silence energy level, and the calibration energy level.
37 . The device of claim 36 , wherein the correctness determination results in at least one error related to an incorrect execution of the voice production, wherein each error is any of: a gentle onset, a soft peak, a gentle offset, a volume control, a pattern usage, a miss of a subsequent voice production, an asymmetry of the voice production, a short inhale, a slow voice production, a fast voice production, a short voice production, a long voice production, and an intense peak voice production.
38 . The device of claim 22 , wherein the at least one exercise is related to fluency shaping.
39 . The device of claim 38 , wherein the at least one exercise is related to customized content.
40 . The device of claim 22 , wherein the device is further configured to:
generate a reporting summarizing the execution of the voice production throughout the current therapy session; and save the report.
41 . The device of claim 22 , wherein the speech disorder therapy is for at least one of: stuttering, cluttering, and diction
42 . A method for monitoring a speech of a user, comprising:
capturing, by a user device, a voice production during a conversation of the user; analyzing the voice production to detect at least a fluency shaping error; and upon detecting the fluency shaping error, generating an instructive notification for improving the speech of the user during the conversation.
43 . The method of claim 42 , wherein the fluency shaping error is an abnormal speech rate.
44 . The method of claim 43 , further comprising:
analyzing the voice production to measure a speech rate of the user; comparing the measured speech rate to a threshold indicating a normal speech rate to determine whether the measured speech rate meets the threshold; and upon determining that the measured speech rate does not meet the threshold, generating the instructive notification to indicate the measured speech rate respective of the threshold.
45 . The method of claim 42 , further comprising:
triggering the user to practice a fluency shaping exercise respective of the detected error.
46 . The method of claim 42 , wherein the fluency shaping error is any of: a gentle onset, a soft peak, a gentle offset, a volume control, a pattern usage, a miss of a subsequent voice production, an asymmetry of the voice production, a short inhale, a slow voice production, a fast voice production, a short voice production, a long voice production, and an intense peak voice production.
47 . A non-transitory computer readable medium having stored thereon instructions for causing one or more processing units to execute the method according to claim 42 .
48 . A device for monitoring a speech of a user, comprising:
an interface for receiving a voice production of a user of a first user device; a processing unit; and a memory coupled to the processing unit, the memory containing instructions that, when executed by the processing unit, configure the device to: set the first user device with at least one exercise to be performed during a current therapy session, wherein each exercise includes at least a difficulty parameter; capture, by a user device, the voice production during a conversation of the user; analyze the voice production to detect at least a fluency shaping error; upon detecting the fluency shaping error, generate an instructive notification for improving the speech of the user during the conversation.
49 . The device of claim 48 , wherein the fluency shaping error is an abnormal speech rate.
50 . The device of claim 48 , wherein the device is further configured to:
analyze the voice production to measure a speech rate of the user; compare the measured speech rate to a threshold indicating a normal speech rate to determine whether the measured speech rate meets the threshold; and upon determining that the measured speech rate does not meet the threshold, generate the instructive notification to indicate the measured speech rate respective of the threshold.
51 . The device of claim 48 , wherein the device is further configured to:
trigger the user to practice a fluency shaping exercise respective of the detected error.
52 . The device of claim 18 , wherein the fluency shaping error is any of: a gentle onset, a soft peak, a gentle offset, a volume control, a pattern usage, a miss of a subsequent voice production, an asymmetry of the voice production, a short inhale, a slow voice production, a fast voice production, a short voice production, a long voice production, and an intense peak voice production.Join the waitlist — get patent alerts
Track US2016183867A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.