US2022059077A1PendingUtilityA1

Training speech recognition systems using word sequences

Assignee: SORENSON IP HOLDINGS LLCPriority: Aug 19, 2020Filed: Aug 19, 2020Published: Feb 24, 2022
Est. expiryAug 19, 2040(~14.1 yrs left)· nominal 20-yr term from priority
Inventors:David Thomson
H04L 9/0869G10L 15/065G06F 21/602G10L 15/19G10L 15/197G10L 25/51G10L 15/063
56
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method may include obtaining a text string that is a transcription of audio data and selecting a sequence of words from the text string as a first word sequence. The method may further include encrypting the first word sequence and comparing the encrypted first word sequence to multiple encrypted word sequences. Each of the multiple encrypted word sequences may be associated with a corresponding one of multiple counters. The method may also include in response to the encrypted first word sequence corresponding to one of the multiple encrypted word sequences based on the comparison, incrementing a counter of the multiple counters associated with the one of the multiple encrypted word sequences and adapting a language model of an automatic transcription system using the multiple encrypted word sequences and the multiple counters.

Claims

exact text as granted — not AI-modified
1 . A method comprising:
 obtaining a text string that is a transcription of audio data;   selecting a sequence of words from the text string as a first word sequence;   encrypting the first word sequence;   comparing the encrypted first word sequence to a plurality of encrypted word sequences, each of the plurality of encrypted word sequences associated with a corresponding one of a plurality of counters;   in response to the encrypted first word sequence corresponding to one of the plurality of encrypted word sequences based on the comparison, incrementing a counter of the plurality of counters associated with the one of the plurality of encrypted word sequences; and   adapting a language model of an automatic transcription system using the plurality of encrypted word sequences and the plurality of counters.   
     
     
         2 . The method of  claim 1 , wherein the plurality of counters are encrypted and the counter associated with the one of the plurality of encrypted word sequences is incremented while being encrypted. 
     
     
         3 . The method of  claim 2 , wherein a first encryption key for the plurality of encrypted word sequences is different from a second encryption key for the plurality of encrypted counters. 
     
     
         4 . The method of  claim 1 , wherein the plurality of counters are initialized with random numbers. 
     
     
         5 . The method of  claim 1 , wherein before obtaining the text string, the plurality of encrypted word sequences are generated from random text strings generated from another plurality of word sequences or a second language model. 
     
     
         6 . The method of  claim 1 , further comprising:
 obtaining second audio data originating at a plurality of first devices;   obtaining a plurality of second text strings that are transcriptions of the second audio data; and   before obtaining the text string, generating the plurality of encrypted word sequences from the plurality of second text strings,   wherein the audio data originates at a plurality of second devices and the plurality of second devices do not include the plurality of first devices.   
     
     
         7 . The method of  claim 1 , further comprising after incrementing the counter of the plurality of counters, removing a second word sequence of the plurality of encrypted word sequences from the plurality of encrypted word sequences based on a second counter of the plurality of counters associated with the second word sequence satisfying a threshold. 
     
     
         8 . The method of  claim 7 , wherein before obtaining the text string, the first word sequence is generated from random text strings generated from another plurality of word sequences or a second language model. 
     
     
         9 . The method of  claim 8 , further comprising:
 after removing the first word sequence, generating a second word sequence to include in the plurality of encrypted word sequences using the plurality of encrypted word sequences.   
     
     
         10 . The method of  claim 1 , further comprising decrypting the plurality of encrypted word sequences, wherein the language model is adapted using the decrypted plurality of word sequence and the plurality of counters. 
     
     
         11 . The method of  claim 1 , wherein each one of the plurality of counters indicates a number of occurrences that a corresponding one of the plurality of encrypted words sequences is included in a plurality of transcriptions of a plurality of communication sessions that occur between a plurality of devices. 
     
     
         12 . A non-transitory computer-readable medium configured to store instructions that when executed by a computer system perform the method of  claim 1 . 
     
     
         13 . A method comprising:
 generating a plurality of word sequences from random text strings generated from another plurality of word sequences or language model;   obtaining a text string that is a transcription of audio data;   selecting a sequence of words from the text string as a first word sequence;   comparing the first word sequence to the plurality of word sequences, each of the plurality of word sequences associated with a corresponding one of a plurality of counters;   in response to the first word sequence corresponding to one of the plurality of word sequences based on the comparison, incrementing a counter of the plurality of counters associated with the one of the plurality of word sequences;   removing a second word sequence of the plurality of word sequences from the plurality of word sequences based on a second counter of the plurality of counters associated with the second word sequence satisfying a threshold; and   after removing the second word sequence, adapting a language model of an automatic transcription system using the plurality of word sequences and the plurality of counters.   
     
     
         14 . The method of  claim 13 , further comprising:
 encrypting the first word sequence; and   encrypting the plurality of word sequences, wherein the first word sequence and the plurality of word sequences are both encrypted when compared.   
     
     
         15 . The method of  claim 13 , wherein the plurality of counters are encrypted and the counter associated with the one of the plurality of encrypted word sequences is incremented while being encrypted. 
     
     
         16 . The method of  claim 13 , wherein the plurality of counters are initialized with random numbers. 
     
     
         17 . The method of  claim 13 , further comprising after removing the second word sequence, generating a third word sequence to include in the plurality of word sequences using the plurality of word sequences. 
     
     
         18 . The method of  claim 13 , further comprising:
 encrypting the first word sequence using a first encryption key;   encrypting the plurality of word sequences using the first encryption key, wherein the first word sequence and the plurality of word sequences are both encrypted when compared; and   encrypting the plurality of counters using a second encryption key that is different from the first encryption key, wherein the counter is incremented while being encrypted.   
     
     
         19 . A non-transitory computer-readable medium configured to store instructions that when executed by a computer system perform the method of  claim 13 . 
     
     
         20 . A system comprising:
 at least one computer-readable media configured to store instructions; and
 at least one processor coupled to the one computer-readable media, the processor configured to execute the instructions to cause the system to perform operations, the operations comprising: 
 obtain a text string that is a transcription of audio data; 
 select a sequence of words from the text string as a first word sequence; 
 encrypt the first word sequence; 
 compare the encrypted first word sequence to a plurality of encrypted word sequences, each of the plurality of encrypted word sequences associated with a corresponding one of a plurality of counters; 
 in response to the encrypted first word sequence corresponding to one of the plurality of encrypted word sequences based on the comparison, increment a counter of the plurality of counters associated with the one of the plurality of encrypted word sequences; and 
 adapt a language model of an automatic transcription system using the plurality of encrypted word sequences and the plurality of counters.

Join the waitlist — get patent alerts

Track US2022059077A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.