US2019244606A1PendingUtilityA1

Closed captioning through language detection

Assignee: IBMPriority: May 26, 2017Filed: Apr 18, 2019Published: Aug 8, 2019
Est. expiryMay 26, 2037(~10.8 yrs left)· nominal 20-yr term from priority
G10L 15/063G10L 15/183G10L 15/005
53
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In an approach for acoustic modeling with a language model, a computer isolates an audio stream. The computer identifies one or more language models based at least in part on the isolated audio stream. The computer selects a language model from the identified one or more language models. The computer creates a text based on the selected language model and the isolated audio stream. The computer creates an acoustic model based on the created text. The computer generates a confidence level associated with the created acoustic model. The computer selects a highest ranked language model based at least in part on the generated confidence level.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for acoustic modeling with a language model, the method comprising:
 isolating, by one or more computer processors, an audio stream;   identifying, by one or more computer processors, one or more language models based at least in part on the isolated audio stream;   selecting, by one or more computer processors, a language model from the identified one or more language models based on the isolated audio stream, wherein selecting, by one or more computer processors, the language model from the identified one or more language models based on the isolated audio stream further comprises;
 receiving, by one or more computer processors, a latency delay; 
 analyzing, by one or more computer processors, a sample of the isolated audio stream based on the received latency delay; 
 identifying, by one or more computer processors, one or more words within the isolated audio stream; 
 identifying, by one or more computer processors, a number of overall words for the identified one or more words that are included within each instance of the identified one or more language models; and 
 selecting, by one or more computer processors, the language model from the identified one or more language models based on the identified number of overall words, wherein the selected language model is based on a highest number of included words; 
   creating, by one or more computer processor, a text based on the selected language model and the isolated audio stream;   creating, by one or more computer processors, an acoustic model based on the created text;   generating, by one or more computer processors, a confidence level associated with the created acoustic model, wherein generating a confidence level associated with the created acoustic model further comprises:
 comparing, by one or more computer processors, the created acoustic model and the isolated audio stream, based on a probability of the created acoustic model to the isolated audio stream, a word error rate of the created acoustic model to the isolated audio stream, and a semantic error rate of the created acoustic model to the isolated audio stream; 
 selecting, by one or more computer processors, a highest ranked language model based at least in part on the generated confidence level, wherein selecting the highest ranked language model based at least in part on the generated confidence level further comprises: 
 comparing, by one or more computer process, the generated confidence level to an another generated confidence level; 
 identifying, by one or more computer processors, a higher confidence level based on the comparison of the generated confidence level to the another generated confidence level; and 
 selecting, by one or more computer processors, the highest ranked language model based on the identified higher confidence level; 
   performing, by one or more computer processors, machine learning for the selected language model based on the isolated audio stream;   updating, by one or more computer processors, the selected language model with the performed machine learning, wherein the performed machine learning comprises pattern recognition, data mining and a neural network;   determining, by one or more computer processors, whether an another language model exists within the identified one or more language models, wherein the another language model is a different language model than the selected language model;   responsive to determining the another language model exists, selecting, by one or more computer processors, the another language model from the identified one or more language models;   creating, by one or more computer processor, an another text based on the selected another language model and the isolated audio stream;   creating, by one or more computer processors, an another acoustic model based on the created another text;   generating, by one or more computer processors, an another confidence level associated with the created another acoustic model;   generating, by one or more computer processors, a curve based on a comparison of the created acoustic model and the isolated audio stream;   determining, by one or more computer processors, the confidence level based on interval estimates of the generated curve;   determining, by one or more computer processors, a best language model based at least in part upon the determined confidence level, wherein the best language model exceeds a minimum rating;   providing, by one or more computer processors, closed captioning for the isolated audio stream based on the determined best language model; and   displaying, by one or more computer processors, the provided closed captioning with a streaming video.

Join the waitlist — get patent alerts

Track US2019244606A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.