US2016183867A1PendingUtilityA1

Method and system for online and remote speech disorders therapy

Assignee: NOVOTALK LTDPriority: Dec 31, 2014Filed: Dec 22, 2015Published: Jun 30, 2016
Est. expiryDec 31, 2034(~8.4 yrs left)· nominal 20-yr term from priority
H04L 65/1069A61B 5/486A61B 5/7465G10L 25/66G16H 20/40H04L 67/10G16H 40/63G09B 5/02G09B 19/04A61B 5/0022G09B 7/00A61B 5/4803A61B 5/742A61B 5/7282A61B 5/0077A61B 7/003A61B 2560/0475A61B 2560/0223G16Z 99/00G16H 40/67G16H 20/30
38
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method and device for enabling remote speech disorder therapy are presented. The method includes setting a first device with at least one exercise to be performed during a current therapy session, wherein each exercise includes at least a difficulty parameter; receiving a voice production of a user of the first device; processing the received voice production to evaluate a correct execution of the voice production respective of the at least one difficulty parameter; generating a feedback based on the analysis; and outputting the generated feedback to the first device.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for enabling remote speech disorder therapy, comprising:
 setting a first device with at least one exercise to be performed during a current therapy session, wherein each exercise includes at least a difficulty parameter;   receiving a voice production of a user of the first device;   processing the received voice production to evaluate a correct execution of the voice production respective of the at least one difficulty parameter;   generating a feedback based on the analysis; and   outputting the generated feedback to the first device.   
     
     
         2 . The method of  claim 1 , further comprising:
 establishing a network communication channel between the first device and a second device; and   outputting the generated feedback to the second device, thereby enabling a user of the second device to remotely monitor the execution of the at least one exercise.   
     
     
         3 . The method of  claim 2 , further comprising:
 receiving instructions from the second device, wherein the instructions include at least one of: a video stream, a video clip, a text file, an image, and an audio clip.   
     
     
         4 . The method of  claim 2 , wherein the user of the first device is a patient and the user of the second device is a therapist. 
     
     
         5 . The method of  claim 1 , wherein the generated feedback is at least a visual feedback. 
     
     
         6 . The method of  claim 5 , further comprising:
 rendering a target template respective of the at least one exercise and the received voice production; and   displaying the target template at least on the first device corresponding to the received voice production.   
     
     
         7 . The method of  claim 6 , wherein the displayed target template includes at least one of: a start boundary, a finish boundary, and a top boundary. 
     
     
         8 . The method of  claim 7 , wherein the target template and at least the start boundary are displayed as the voice production is received. 
     
     
         9 . The method of  claim 5 , wherein generating feedback based on the analysis further comprises:
 coloring the voice production using at least a first color and a second color, wherein the first color represents a loud sound produced by the user and the second color represents a soft sound produced by the user; and   displaying at least one of: a positive indication upon performing a correct execution, and an instructive indication upon performing an incorrect execution.   
     
     
         10 . The method of  claim 5 , wherein the at least one exercise includes a sequence having a plurality of target templates that require the user to produce a sequence of voice productions. 
     
     
         11 . The method of  claim 10 , further comprising:
 providing a breathing indicator, wherein the breathing indicator represents a duration of time that the user needs to breathe before trying a subsequent target template, wherein the duration of time is determined based on the at least one difficulty parameter.   
     
     
         12 . The method of  claim 1 , further comprising:
 measuring a speech rate respective of the analysis; and   displaying a speech-rate meter respective of the measured speech rate.   
     
     
         13 . The method of  claim 1 , further comprising:
 performing an audio calibration process for the first user device, wherein the audio calibration process provides at least a normal speech energy level, a silence energy level, and a calibration energy level.   
     
     
         14 . The method of  claim 13 , wherein processing the received voice production further comprises:
 sampling the received voice production to create voice samples;   buffering the voice samples to create voice chunks;   converting the voice chunks from a time domain to a frequency domain;   extracting spectrum features from each of the frequency domain voice chunks, wherein the spectrum features include at least dominant frequencies, wherein each dominant frequency corresponds to a voice chunk;   computing, for each voice chunk, the energy level of the corresponding dominant frequency; and   determining, for each voice chunk, an energy level of the voice chunk based on the energy level of the corresponding dominant frequency.   
     
     
         15 . The method of  claim 14 , further comprising:
 determining a correctness of the execution of the voice production based on the energy levels of the voice chunks and at least one of: the normal speech energy level, the silence energy level, and the calibration energy level.   
     
     
         16 . The method of  claim 15 , wherein the correctness determination results in at least one error related to an incorrect execution of the voice production, wherein each error is any of: a gentle onset, a soft peak, a gentle offset, a volume control, a pattern usage, a miss of a subsequent voice production, an asymmetry of the voice production, a short inhale, a slow voice production, a fast voice production, a short voice production, a long voice production, and an intense peak voice production. 
     
     
         17 . The method of  claim 1 , wherein the at least one exercise is related to fluency shaping. 
     
     
         18 . The method of  claim 17 , wherein the at least one exercise is related to customized content. 
     
     
         19 . The method of  claim 1 , further comprising:
 generating a reporting summarizing the execution of the voice production throughout the current therapy session; and   saving the report.   
     
     
         20 . The method of  claim 1 , wherein the speech disorder therapy is for at least one of:
 stuttering, cluttering, and diction.   
     
     
         21 . A non-transitory computer readable medium having stored thereon instructions for causing one or more processing units to execute the method according to  claim 1 . 
     
     
         22 . A device for enabling remote speech disorder therapy, comprising:
 an interface for receiving a voice production of a user of a first user device;   a processing unit; and   a memory coupled to the processing unit, the memory containing instructions that, when executed by the processing unit, configure the device to:   set the first user device with at least one exercise to be performed during a current therapy session, wherein each exercise includes at least a difficulty parameter;   receive the voice production of the user of the first device;   analyze the received voice production to evaluate a correct execution of the voice production respective of the at least one difficulty parameter;   generate a feedback respective of the analysis; and   output the generated feedback to the first device.   
     
     
         23 . The device of  claim 22 , wherein the device is further configured to:
 establish a network communication channel between the first device and a second device; and   output the generated feedback to the second device, thereby enabling a user of the second device to remotely monitor the execution of the at least one exercise.   
     
     
         24 . The device of  claim 23 , wherein the device is further configured to:
 receive instructions from the second device, wherein the instructions include at least one of: a video stream, a video clip, a text file, an image, and an audio clip.   
     
     
         25 . The device of  claim 23 , wherein the user of the first device is a patient and the user of the second device is a therapist. 
     
     
         26 . The device of  claim 22 , wherein the generated feedback is at least a visual feedback. 
     
     
         27 . The device of  claim 26 , wherein the device is further configured to:
 render a target template respective of the at least one exercise and the received voice production; and   display the target template at least on the first device corresponding to the received voice production.   
     
     
         28 . The device of  claim 27 , wherein the displayed target template includes at least one of: a start boundary, a finish boundary, and a top boundary. 
     
     
         29 . The device of  claim 28 , wherein the target template and at least the start boundary are displayed as the voice production is received. 
     
     
         30 . The device of  claim 26 , wherein the device is further configured to:
 color the voice production using at least a first color and a second color, wherein the first color represents a loud sound produced by the user and the second color represents a soft sound produced by the user; and   display at least one of: a positive indication upon performing a correct execution, and an instructive indication upon performing an incorrect execution.   
     
     
         31 . The device of  claim 26 , wherein the at least one exercise includes a sequence having a plurality of target templates that require the user to produce a sequence of voice productions. 
     
     
         32 . The device of  claim 31 , wherein the device is further configured to:
 provide a breathing indicator, wherein the breathing indicator represents a duration of time that the user needs to breathe before trying a subsequent target template, wherein the duration of time is determined based on the at least one difficulty parameter.   
     
     
         33 . The device of  claim 22 , wherein the device is further configured to:
 measure a speech rate respective of the analysis; and   display a speech-rate meter respective of the measured speech rate.   
     
     
         34 . The device of  claim 22 , wherein the device is further configured to:
 perform an audio calibration process for the first user device, wherein the audio calibration process provides at least a normal speech energy level, a silence energy level, and a calibration energy level.   
     
     
         35 . The device of  claim 34 , wherein the device is further configured to:
 sample the received voice production to create voice samples;   buffer the voice samples to create voice chunks;   convert the voice chunks from a time domain to a frequency domain;   extract spectrum features from each of the frequency domain voice chunks, wherein the spectrum features include at least dominant frequencies, wherein each dominant frequency corresponds to a voice chunk;   compute, for each voice chunk, the energy level of the corresponding dominant frequency; and   determine, for each voice chunk, an energy level of the voice chunk based on the energy level of the corresponding dominant frequency.   
     
     
         36 . The device of  claim 35 , wherein the device is further configured to:
 determine a correctness of the execution of the voice production based on the energy levels of the voice chunks and at least one of: the normal speech energy level, the silence energy level, and the calibration energy level.   
     
     
         37 . The device of  claim 36 , wherein the correctness determination results in at least one error related to an incorrect execution of the voice production, wherein each error is any of: a gentle onset, a soft peak, a gentle offset, a volume control, a pattern usage, a miss of a subsequent voice production, an asymmetry of the voice production, a short inhale, a slow voice production, a fast voice production, a short voice production, a long voice production, and an intense peak voice production. 
     
     
         38 . The device of  claim 22 , wherein the at least one exercise is related to fluency shaping. 
     
     
         39 . The device of  claim 38 , wherein the at least one exercise is related to customized content. 
     
     
         40 . The device of  claim 22 , wherein the device is further configured to:
 generate a reporting summarizing the execution of the voice production throughout the current therapy session; and   save the report.   
     
     
         41 . The device of  claim 22 , wherein the speech disorder therapy is for at least one of: stuttering, cluttering, and diction 
     
     
         42 . A method for monitoring a speech of a user, comprising:
 capturing, by a user device, a voice production during a conversation of the user;   analyzing the voice production to detect at least a fluency shaping error; and   upon detecting the fluency shaping error, generating an instructive notification for improving the speech of the user during the conversation.   
     
     
         43 . The method of  claim 42 , wherein the fluency shaping error is an abnormal speech rate. 
     
     
         44 . The method of  claim 43 , further comprising:
 analyzing the voice production to measure a speech rate of the user;   comparing the measured speech rate to a threshold indicating a normal speech rate to determine whether the measured speech rate meets the threshold; and   upon determining that the measured speech rate does not meet the threshold, generating the instructive notification to indicate the measured speech rate respective of the threshold.   
     
     
         45 . The method of  claim 42 , further comprising:
 triggering the user to practice a fluency shaping exercise respective of the detected error.   
     
     
         46 . The method of  claim 42 , wherein the fluency shaping error is any of: a gentle onset, a soft peak, a gentle offset, a volume control, a pattern usage, a miss of a subsequent voice production, an asymmetry of the voice production, a short inhale, a slow voice production, a fast voice production, a short voice production, a long voice production, and an intense peak voice production. 
     
     
         47 . A non-transitory computer readable medium having stored thereon instructions for causing one or more processing units to execute the method according to  claim 42 . 
     
     
         48 . A device for monitoring a speech of a user, comprising:
 an interface for receiving a voice production of a user of a first user device;   a processing unit; and   a memory coupled to the processing unit, the memory containing instructions that, when executed by the processing unit, configure the device to:   set the first user device with at least one exercise to be performed during a current therapy session, wherein each exercise includes at least a difficulty parameter;   capture, by a user device, the voice production during a conversation of the user;   analyze the voice production to detect at least a fluency shaping error;   upon detecting the fluency shaping error, generate an instructive notification for improving the speech of the user during the conversation.   
     
     
         49 . The device of  claim 48 , wherein the fluency shaping error is an abnormal speech rate. 
     
     
         50 . The device of  claim 48 , wherein the device is further configured to:
 analyze the voice production to measure a speech rate of the user;   compare the measured speech rate to a threshold indicating a normal speech rate to determine whether the measured speech rate meets the threshold; and   upon determining that the measured speech rate does not meet the threshold, generate the instructive notification to indicate the measured speech rate respective of the threshold.   
     
     
         51 . The device of  claim 48 , wherein the device is further configured to:
 trigger the user to practice a fluency shaping exercise respective of the detected error.   
     
     
         52 . The device of  claim 18 , wherein the fluency shaping error is any of: a gentle onset, a soft peak, a gentle offset, a volume control, a pattern usage, a miss of a subsequent voice production, an asymmetry of the voice production, a short inhale, a slow voice production, a fast voice production, a short voice production, a long voice production, and an intense peak voice production.

Join the waitlist — get patent alerts

Track US2016183867A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.