US2023359812A1PendingUtilityA1

Digitally aware neural dictation interface

Assignee: WELLS FARGO BANK NAPriority: Oct 11, 2019Filed: Jul 18, 2023Published: Nov 9, 2023
Est. expiryOct 11, 2039(~13.2 yrs left)· nominal 20-yr term from priority
G06F 40/174G06F 3/167G10L 17/02G10L 17/00G10L 17/06G06F 40/40G10L 15/26G10L 2015/027
73
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for populating the elements of content are disclosed. One method includes determining a plurality of elements of a document and receiving a first speech input from a user to enable a mode of operation. The method further includes authenticating the user by comparing the first speech input with at least one voice sample of the user and enabling the mode of operation. The method further includes receiving, in the mode of operation, a second speech input for filling out a first element of the document and determining an irregularity or distortion in the second speech input based on the first element and identifying a missing syllable or a distorted syllable. The method further includes refining the second speech input into at least one matching syllable, converting the refined second speech input, and providing the text to populate the first element with the text.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method, comprising:
 determining, by one or more processing circuits, a plurality of elements of a document;   receiving, by the one or more processing circuits, a first speech input from a user to enable a mode of operation;   authenticating, by the one or more processing circuits, the user by comparing the first speech input from the user with at least one voice sample of the user;   in response to authenticating the first speech input, enabling, by the one or more processing circuits, the mode of operation;   receiving, by the one or more processing circuits in the mode of operation, a second speech input for filling out a first element of the document, wherein the first element is selected based on a priority order;   determining, by the one or more processing circuits, an irregularity or distortion in the second speech input based on the first element and identifying a missing syllable or a distorted syllable in the second speech input, wherein identifying the missing syllable or the distorted syllable comprises executing an analysis of the second speech input, wherein either (1) the missing syllable is determined based on other syllables identified in the analysis, or (2) the distorted syllable is determined based on failing to recognize a syllable in the analysis;   refining, by the one or more processing circuits, the second speech input into at least one matching syllable by extrapolating the missing syllable or the distorted syllable based on stored syllables of a plurality of speech inputs, wherein the at least one matching syllable is determined at least in part on an expected element value associated with the first element;   converting, by the one or more processing circuits, the refined second speech input comprising the at least one matching syllable into text; and   providing, by the one or more processing circuits, the text to a user device to populate the first element with the text.   
     
     
         2 . The method of  claim 1 , wherein the irregularity is a first irregularity and the method further comprises:
 determining, by the one or more processing circuits, a second irregularity that is a non-English language speech input;   identifying, by the one or more processing circuits, a language of the non-English language speech input of the second irregularity; and   translating, by the one or more processing circuits, the non-English language speech input into English language.   
     
     
         3 . The method of  claim 1 , wherein determining the irregularity in the second speech input is based on identifying the distorted syllable in the second speech input, and the method further comprises:
 determining, by the one or more processing circuits, that the distorted syllable is due to at least one of an attenuation of the second speech input, a presence of background noise, or an accent in the second speech input.   
     
     
         4 . The method of  claim 3 , further comprising:
 transmitting, by the one or more processing circuits, the second speech input to a speech enhancement circuit to at least partially mitigate the irregularity in the second speech input.   
     
     
         5 . The method of  claim 4 , further comprising:
 receiving, by the one or more processing circuits, a mitigated speech output from the speech enhancement circuit as a refinement to at least partially mitigate the irregularity in the second speech input.   
     
     
         6 . The method of  claim 1 , further comprising:
 highlighting, by the one or more processing circuits on a display screen of the user device, a second element of the plurality of elements in response to determining that the first element is populated with the text.   
     
     
         7 . The method of  claim 1 , further comprising:
 determining, by the one or more processing circuits, the expected element value for the first element based on metadata;   comparing, by the one or more processing circuits, the second speech input to the expected element value; and   determining, by the one or more processing circuits, that the second speech input does not match the expected element value of the first element.   
     
     
         8 . The method of  claim 1 , further comprising:
 correcting, by the one or more processing circuits, an error in the first element by disregarding the received second speech input for a second value of the first element in favor of information that matches the expected element value of the first element.   
     
     
         9 . The method of  claim 1 , further comprising:
 filtering, by the one or more processing circuits through at least one digital processing technique, the second speech input to remove at least a portion of the irregularity.   
     
     
         10 . The method of  claim 1 , wherein the refinement of the second speech input comprises:
 executing, by the one or more processing circuits, at least one artificial intelligence algorithm to compare each syllable in the second speech input to the stored syllables in a database to find a closest match for the missing syllable or the distorted syllable in the second speech input; and   providing, by the one or more processing circuits, at least one user specific auto-complete suggestion based on information stored in the database associated with the user, wherein the information represents stored values corresponding to multiple elements of documents previously filled by the user.   
     
     
         11 . A system, comprising:
 one or more processing circuits configured to:
 determine a plurality of elements of a document; 
 receive a first speech input from a user to enable a mode of operation; 
 authenticate the user by comparing the first speech input from the user with at least one voice sample of the user; 
 in response to authenticating the first speech input, enable the mode of operation; 
 receive, in the mode of operation, a second speech input for filling out a first element of the document, wherein the first element is selected based on a priority order; 
 determine an irregularity or distortion in the second speech input based on the first element and identifying a missing syllable or a distorted syllable in the second speech input, wherein identifying the missing syllable or the distorted syllable comprises executing an analysis of the second speech input, wherein either (1) the missing syllable is determined based on other syllables identified in the analysis, or (2) the distorted syllable is determined based on failing to recognize a syllable in the analysis; 
 refine the second speech input into at least one matching syllable by extrapolating the missing syllable or the distorted syllable based on stored syllables of a plurality of speech inputs, wherein the at least one matching syllable is determined at least in part on an expected element value associated with the first element; 
 convert the refined second speech input comprising the at least one matching syllable into text; and 
 provide the text to a user device to populate the first element with the text. 
   
     
     
         12 . The system of  claim 11 , wherein the irregularity is a first irregularity and the one or more processing circuits are further configured to:
 determine a second irregularity that is a non-English language speech input;   identify a language of the non-English language speech input; and   translate the non-English language speech input into English language.   
     
     
         13 . The system of  claim 11 , wherein determining the irregularity in the second speech input is based on identifying the distorted syllable in the second speech input, and the one or more processing circuits are further configured to:
 determine that the distorted syllable is due to at least one of an attenuation of the second speech input, a presence of background noise, or an accent in the second speech input.   
     
     
         14 . The system of  claim 13 , wherein the one or more processing circuits are further configured to:
 transmit the second speech input to a speech enhancement circuit to at least partially mitigate the irregularity in the second speech input.   
     
     
         15 . The system of  claim 14 , wherein the one or more processing circuits are further configured to:
 receive a mitigated speech output from the speech enhancement circuit as a refinement to at least partially mitigate the irregularity in the second speech input.   
     
     
         16 . The system of  claim 11 , wherein the one or more processing circuits are further configured to:
 highlight a second element of the plurality of elements in response to determining that the first element is populated with the text.   
     
     
         17 . One or more non-transitory computer-readable storage media having instructions stored thereon that, when executed by one or more processing circuits, cause the one or more processing circuits to perform operations comprising:
 determining a plurality of elements of a document;   receiving a first speech input from a user to enable a mode of operation;   authenticating the user by comparing the first speech input from the user with at least one voice sample of the user;   in response to authenticating the first speech input, enabling the mode of operation;   receiving, in the mode of operation, a second speech input for filling out a first element of the document, wherein the first element is selected based on a priority order;   determining an irregularity or distortion in the second speech input based on the first element and identifying a missing syllable or a distorted syllable in the second speech input, wherein identifying the missing syllable or the distorted syllable comprises executing an analysis of the second speech input, wherein either (1) the missing syllable is determined based on other syllables identified in the analysis, or (2) the distorted syllable is determined based on failing to recognize a syllable in the analysis;   refining the second speech input into at least one matching syllable by extrapolating the missing syllable or the distorted syllable based on stored syllables of a plurality of speech inputs, wherein the at least one matching syllable is determined at least in part on an expected element value associated with the first element;   converting the refined second speech input comprising the at least one matching syllable into text; and   providing the text to a user device to populate the first element with the text.   
     
     
         18 . The one or more non-transitory computer-readable storage media of  claim 17 , wherein the instructions, when executed by the one or more processing circuits, further cause the one or more processing circuits to perform operations comprising:
 highlighting, on a display screen of the user device, a second element of the plurality of elements in response to determining that the first element is populated with the text.   
     
     
         19 . The one or more non-transitory computer-readable storage media of  claim 17 , wherein the instructions, when executed by the one or more processing circuits, further cause the one or more processing circuits to perform operations comprising:
 determining the expected element value for the first element based on metadata;   comparing the second speech input to the expected element value; and   determining that the second speech input does not match the expected element value of the first element.   
     
     
         20 . The one or more non-transitory computer-readable storage media of  claim 17 , wherein the instructions, when executed by the one or more processing circuits, further cause the one or more processing circuits to perform operations comprising:
 correcting an error in the first element by disregarding the received second speech input for a second value of the first element in favor of information that matches the expected element value of the first element.

Join the waitlist — get patent alerts

Track US2023359812A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.