Integrated language model, related systems and methods
Abstract
An integrated language model includes an upper-level language model component and a lower-level language model component, with the upper-level language model component including a non-terminal and the lower-level language model component being applied to the non-terminal. The upper-level and lower-level language model components can be of the same or different language model formats, including finite state grammar (FSG) and statistical language model (SLM) formats. Systems and methods for making integrated language models allow designation of language model formats for the upper-level and lower-level components and identification of non-terminals. Automatic non-terminal replacement and retention criteria can be used to facilitate the generation of one or both language model components, which can include the modification of existing language models.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of making an integrated language model for a speech recognition engine, the method comprising:
identifying a first language model format for an upper-level language model component; identifying a plurality of text elements to be represented by a non-terminal in the upper-level language model component; generating the upper-level language model component including the non-terminal; identifying a second language model format for a lower-level language model component to be applied to the non-terminal of the upper-level language model component; and generating the lower-level language model component.
2 . The method of claim 1 , wherein the first language model format is a finite state grammar format.
3 . The method of claim 2 , wherein the second language model format is a statistical language model format.
4 . The method of claim 1 , wherein a plurality of non-terminals are included in the upper-level language model.
5 . The method of claim 4 , wherein a plurality of lower-level language model components are applied to the plurality of non-terminals.
6 . The method of claim 5 , wherein the plurality of lower-level language model components include at least two language model components having different language model formats.
7 . The method of claim 5 , further comprising identifying to which of the plurality of non-terminals each of the plurality of lower-level language models is to be applied during operation of the speech recognition engine.
8 . The method of claim 1 , wherein the lower-level language model component also includes a non-terminal, and the method further comprises identifying a third language model format for an additional lower-level language model component to be applied to the non-terminal of the lower-level language model component, and generating the additional language model component.
9 . The method of claim 1 , wherein generating the upper-level language model component includes modifying an existing language model.
10 . The method of claim 9 , wherein the existing language model is a finite state grammar format language model.
11 . The method of claim 9 , wherein the existing language model is a statistical language model format language model.
12 . The method of claim 9 , wherein identifying the plurality of text elements to be represented by the non-terminal in the upper-level language model component includes applying an automatic text element replacement criterion to the existing language model.
13 . The method of claim 12 , wherein the text element replacement criterion is to replace text element sequences having a definable length.
14 . The method of claim 13 , wherein the definable length is a number of digits.
15 . The method of claim 12 , wherein the text element replacement criterion is to replace text element sequences having a definable value range.
16 . The method of claim 12 , wherein a plurality of automatic text element replacement criteria are applied.
17 . The method of claim 12 , wherein identifying the plurality of text elements to be represented by the non-terminal in the upper-level language model component further includes applying an automatic text element retention criterion to the existing language model.
18 . The method of claim 1 , wherein identifying the plurality of text elements to be represented by the non-terminal in the upper-level language model component includes receiving a user-supplied list of text elements.
19 . The method of claim 1 , wherein generating the lower-level language model component includes modifying an existing language model.
20 . The method of claim 19 , wherein modifying the existing language model includes automatically eliminating text elements that are determined not to be relevant to the non-terminal, the determination being based on a non-terminal replacement criterion applied to identify the plurality of text elements to be represented by the non-terminal.
21 . The method of claim 1 , further comprising generating instructions for a speech recognition engine decoder to apply the upper-level and lower-level language model components during operation.
22 . The method of claim 1 , wherein the text elements are words.
23 . A method for identifying text elements to be represented by non-terminals in an integrated language model for a speech recognition engine, the method comprising:
determining a text element replacement criterion allowing automatic identification of the text elements to be represented by the non-terminals within an existing language model or textual corpus; and applying the text element replacement criterion to the existing language model or textual corpus.
24 . The method of claim 23 , further comprising:
receiving a user-supplied list of text elements to be represented by the non-terminals to the existing language model or textual corpus; and applying the user-supplied list to the existing language model or textual corpus.
25 . The method of claim 23 , further comprising:
determining a text element retention criterion allowing automatic identification of text elements falling within the text element replacement criterion that should be retained within the existing language model or textual corpus; and applying the text element retention criterion to the existing language model or textual corpus.
26 . The method of claim 23 , wherein a plurality of text element replacement criteria are identified and applied.
27 . The method of claim 23 , wherein the text element replacement criterion is to replace text element sequences having a definable length.
28 . The method of claim 27 , wherein the definable length is a number of digits.
29 . The method of claim 23 , wherein the text element replacement criterion is to replace text element sequences having a definable value range.
30 . The method of claim 23 , wherein the text element replacement criterion is applied to a finite state grammar format language model.
31 . A system for making an integrated language model, the system comprising at least one processor and machine-readable memory configured to execute:
a language model integration control module adapted to receive user inputs regarding language model integration options and to generate language model modification rules and application rules based thereon; a language model generation module adapted to modify existing language models based upon the language model generation rules to generate upper-level and lower-level language model components for the integrated language model.
32 . The system of claim 31 , wherein the language model integration options include at least one of:
identification of an upper-level language model format; identification of a lower-level language model format; and identification of text element replacement criteria.Join the waitlist — get patent alerts
Track US2014163989A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.