US2009216739A1PendingUtilityA1

Boosting extraction accuracy by handling training data bias

Assignee: YAHOO INCPriority: Feb 22, 2008Filed: Feb 22, 2008Published: Aug 27, 2009
Est. expiryFeb 22, 2028(~1.6 yrs left)· nominal 20-yr term from priority
G06F 16/313
31
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods and apparatus are described for use with information extraction techniques based on sequential models. Additional statistics are maintained during inference and employed to boost the accuracy of the extraction algorithm and mitigate the effects of training bias.

Claims

exact text as granted — not AI-modified
1 . A computer-implemented method for extracting information from sequential data, the sequential data comprising a plurality of sequentially arranged tokens, the method comprising:
 generating a plurality of label sequences with reference to the sequential data and a sequential model, each label sequence comprising a plurality of attribute labels, at least some of the attribute labels corresponding to attributes of interest, the attribute labels in each label sequence being sequentially arranged and corresponding to the tokens of the sequential data;   generating an output sequence using selected ones of the attribute labels from different ones of the label sequences, each of the selected attribute labels corresponding to one of the attributes of interest, each selected attribute label occupying a same position in the output sequence as in a corresponding one of the label sequences from which the selected attribute label originated; and   generating a representation of selected ones of the tokens corresponding to the selected attribute labels.   
   
   
       2 . The method of  claim 1  wherein each label sequence has a confidence level associated therewith, and wherein the plurality of label sequences correspond to the k highest confidence levels, where k is a natural number which is fewer than a total number of sequences generated for the sequential data. 
   
   
       3 . The method of  claim 1  wherein the plurality of label sequences includes a highest confidence sequence for each of the attributes of interest. 
   
   
       4 . The method of  claim 1  wherein the sequential model comprises one of a Conditional Random Field model, a Hidden Markov model, or a Maximum Entropy Markov model. 
   
   
       5 . The method of  claim 1  wherein the sequential data represents one of a web page, a portion of a genome, recorded speech, or text. 
   
   
       6 . The method of  claim 1  further comprising transmitting the representation of the selected tokens in response to a search query relating to at least one of the attributes of interest. 
   
   
       7 . The method of  claim 1  wherein each label sequence has a confidence level associated therewith, and wherein the confidence level associated with the label sequence from which each of the selected attribute labels is selected for inclusion in the output sequence is highest among all sequences including the corresponding selected attribute label. 
   
   
       8 . A computer program product for extracting information from sequential data, the sequential data comprising a plurality of sequentially arranged tokens, the computer program product comprising at least one computer-readable medium having computer program instructions stored therein configured to cause at least one computing device to:
 generate a plurality of label sequences with reference to the sequential data and a sequential model, each label sequence comprising a plurality of attribute labels, at least some of the attribute labels corresponding to attributes of interest, the attribute labels in each label sequence being sequentially arranged and corresponding to the tokens of the sequential data;   generate an output sequence using selected ones of the attribute labels from different ones of the label sequences, each of the selected attribute labels corresponding to one of the attributes of interest, each selected attribute label occupying a same position in the output sequence as in a corresponding one of the label sequences from which the selected attribute label originated; and   generate a representation of selected ones of the tokens corresponding to the selected attribute labels.   
   
   
       9 . The computer program product of  claim 8  wherein each label sequence has a confidence level associated therewith, and wherein the plurality of label sequences correspond to the k highest confidence levels, where k is a natural number which is fewer than a total number of sequences generated for the sequential data. 
   
   
       10 . The computer program product of  claim 8  wherein the plurality of label sequences includes a highest confidence level sequence for each of the attributes of interest. 
   
   
       11 . The computer program product of  claim 8  wherein the sequential model comprises one of a Conditional Random Field model, a Hidden Markov model, or a Maximum Entropy Markov model. 
   
   
       12 . A computer-implemented method for presenting information extracted from sequential data, the sequential data comprising a plurality of sequentially arranged tokens, the method comprising facilitating presentation of a representation of selected ones of the tokens in a user interface, the selected tokens corresponding to selected ones of a plurality of attribute labels, each selected attribute label corresponding to one of a plurality of attributes of interest and having been selected for inclusion in an output sequence from a corresponding one of a plurality of label sequences, each selected attribute label having occupied a same position in the output sequence as in the corresponding label sequence from which the selected attribute label originated, the label sequences having been generated with reference to the sequential data and a sequential model, each label sequence having included at least some of the plurality of attribute labels, the attribute labels in each label sequence having been sequentially arranged and having corresponded to the tokens of the sequential data. 
   
   
       13 . The method of  claim 12  wherein the sequential data represented one of a web page, a portion of a genome, recorded speech, or text. 
   
   
       14 . The method of  claim 12  wherein presentation of the representation of the selected tokens is facilitated in response to a search query relating to at least one of the attributes of interest. 
   
   
       15 . At least one computer-readable medium having a data structure stored therein representing information extracted from sequential data, the sequential data comprising a plurality of sequentially arranged tokens, the data structure comprising an output sequence comprising selected ones of a plurality of attribute labels, each selected attribute label corresponding to one of a plurality of attributes of interest and having been selected for inclusion in the output sequence from a corresponding one of a plurality of label sequences, each selected attribute label occupying a same position in the output sequence as in the corresponding label sequence from which the selected attribute label originated, the label sequences having been generated with reference to the sequential data and a sequential model, each label sequence having included at least some of the plurality of attribute labels, the attribute labels in each label sequence having been sequentially arranged and having corresponded to the tokens of the sequential data, wherein the output sequence is configured to facilitate presentation of a representation of selected ones of the tokens in a user interface, the selected tokens corresponding to the selected attribute labels. 
   
   
       16 . The at least one computer-readable medium of  claim 15  wherein each label sequence had a confidence level associated therewith, and wherein the plurality of label sequences corresponded to the k highest confidence levels, where k is a natural number which is fewer than a total number of sequences generated for the sequential data. 
   
   
       17 . The at least one computer-readable medium of  claim 15  wherein the plurality of label sequences included a highest confidence level sequence for each of the attributes of interest. 
   
   
       18 . The at least one computer-readable medium of  claim 15  wherein the sequential model comprises one of a Conditional Random Field model, a Hidden Markov model, or a Maximum Entropy Markov model.

Join the waitlist — get patent alerts

Track US2009216739A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.