US2009216739A1PendingUtilityA1
Boosting extraction accuracy by handling training data bias
Est. expiryFeb 22, 2028(~1.6 yrs left)· nominal 20-yr term from priority
G06F 16/313
31
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Methods and apparatus are described for use with information extraction techniques based on sequential models. Additional statistics are maintained during inference and employed to boost the accuracy of the extraction algorithm and mitigate the effects of training bias.
Claims
exact text as granted — not AI-modified1 . A computer-implemented method for extracting information from sequential data, the sequential data comprising a plurality of sequentially arranged tokens, the method comprising:
generating a plurality of label sequences with reference to the sequential data and a sequential model, each label sequence comprising a plurality of attribute labels, at least some of the attribute labels corresponding to attributes of interest, the attribute labels in each label sequence being sequentially arranged and corresponding to the tokens of the sequential data; generating an output sequence using selected ones of the attribute labels from different ones of the label sequences, each of the selected attribute labels corresponding to one of the attributes of interest, each selected attribute label occupying a same position in the output sequence as in a corresponding one of the label sequences from which the selected attribute label originated; and generating a representation of selected ones of the tokens corresponding to the selected attribute labels.
2 . The method of claim 1 wherein each label sequence has a confidence level associated therewith, and wherein the plurality of label sequences correspond to the k highest confidence levels, where k is a natural number which is fewer than a total number of sequences generated for the sequential data.
3 . The method of claim 1 wherein the plurality of label sequences includes a highest confidence sequence for each of the attributes of interest.
4 . The method of claim 1 wherein the sequential model comprises one of a Conditional Random Field model, a Hidden Markov model, or a Maximum Entropy Markov model.
5 . The method of claim 1 wherein the sequential data represents one of a web page, a portion of a genome, recorded speech, or text.
6 . The method of claim 1 further comprising transmitting the representation of the selected tokens in response to a search query relating to at least one of the attributes of interest.
7 . The method of claim 1 wherein each label sequence has a confidence level associated therewith, and wherein the confidence level associated with the label sequence from which each of the selected attribute labels is selected for inclusion in the output sequence is highest among all sequences including the corresponding selected attribute label.
8 . A computer program product for extracting information from sequential data, the sequential data comprising a plurality of sequentially arranged tokens, the computer program product comprising at least one computer-readable medium having computer program instructions stored therein configured to cause at least one computing device to:
generate a plurality of label sequences with reference to the sequential data and a sequential model, each label sequence comprising a plurality of attribute labels, at least some of the attribute labels corresponding to attributes of interest, the attribute labels in each label sequence being sequentially arranged and corresponding to the tokens of the sequential data; generate an output sequence using selected ones of the attribute labels from different ones of the label sequences, each of the selected attribute labels corresponding to one of the attributes of interest, each selected attribute label occupying a same position in the output sequence as in a corresponding one of the label sequences from which the selected attribute label originated; and generate a representation of selected ones of the tokens corresponding to the selected attribute labels.
9 . The computer program product of claim 8 wherein each label sequence has a confidence level associated therewith, and wherein the plurality of label sequences correspond to the k highest confidence levels, where k is a natural number which is fewer than a total number of sequences generated for the sequential data.
10 . The computer program product of claim 8 wherein the plurality of label sequences includes a highest confidence level sequence for each of the attributes of interest.
11 . The computer program product of claim 8 wherein the sequential model comprises one of a Conditional Random Field model, a Hidden Markov model, or a Maximum Entropy Markov model.
12 . A computer-implemented method for presenting information extracted from sequential data, the sequential data comprising a plurality of sequentially arranged tokens, the method comprising facilitating presentation of a representation of selected ones of the tokens in a user interface, the selected tokens corresponding to selected ones of a plurality of attribute labels, each selected attribute label corresponding to one of a plurality of attributes of interest and having been selected for inclusion in an output sequence from a corresponding one of a plurality of label sequences, each selected attribute label having occupied a same position in the output sequence as in the corresponding label sequence from which the selected attribute label originated, the label sequences having been generated with reference to the sequential data and a sequential model, each label sequence having included at least some of the plurality of attribute labels, the attribute labels in each label sequence having been sequentially arranged and having corresponded to the tokens of the sequential data.
13 . The method of claim 12 wherein the sequential data represented one of a web page, a portion of a genome, recorded speech, or text.
14 . The method of claim 12 wherein presentation of the representation of the selected tokens is facilitated in response to a search query relating to at least one of the attributes of interest.
15 . At least one computer-readable medium having a data structure stored therein representing information extracted from sequential data, the sequential data comprising a plurality of sequentially arranged tokens, the data structure comprising an output sequence comprising selected ones of a plurality of attribute labels, each selected attribute label corresponding to one of a plurality of attributes of interest and having been selected for inclusion in the output sequence from a corresponding one of a plurality of label sequences, each selected attribute label occupying a same position in the output sequence as in the corresponding label sequence from which the selected attribute label originated, the label sequences having been generated with reference to the sequential data and a sequential model, each label sequence having included at least some of the plurality of attribute labels, the attribute labels in each label sequence having been sequentially arranged and having corresponded to the tokens of the sequential data, wherein the output sequence is configured to facilitate presentation of a representation of selected ones of the tokens in a user interface, the selected tokens corresponding to the selected attribute labels.
16 . The at least one computer-readable medium of claim 15 wherein each label sequence had a confidence level associated therewith, and wherein the plurality of label sequences corresponded to the k highest confidence levels, where k is a natural number which is fewer than a total number of sequences generated for the sequential data.
17 . The at least one computer-readable medium of claim 15 wherein the plurality of label sequences included a highest confidence level sequence for each of the attributes of interest.
18 . The at least one computer-readable medium of claim 15 wherein the sequential model comprises one of a Conditional Random Field model, a Hidden Markov model, or a Maximum Entropy Markov model.Join the waitlist — get patent alerts
Track US2009216739A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.