US2010049772A1PendingUtilityA1

Extraction of anchor explanatory text by mining repeated patterns

Assignee: MICROSOFT CORPPriority: Mar 31, 2006Filed: Oct 30, 2009Published: Feb 25, 2010
Est. expiryMar 31, 2026(expired)· nominal 20-yr term from priority
G06F 16/951Y10S707/99936
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method and system for identifying explanatory text for a referenced web page based on a reference to the referenced web page contained in a repeated pattern of a referencing web page is provided. An anchor explanatory text (“AET”) system uses the hierarchical organization of the web page to identify a repeated pattern of hierarchical elements that contain references to other display pages. After the AET system identifies a repeated pattern, it identifies the dominant reference or anchor within each occurrence of the pattern. The AET system uses the explanatory text surrounding a dominant anchor as a description of the referenced web page.

Claims

exact text as granted — not AI-modified
1 - 11 . (canceled) 
   
   
       12 . A computer-readable storage medium containing instructions for controlling a computer system to identify explanatory text for a referenced web page from a referencing web page, by a method comprising:
 identifying repeated patterns of elements within the referencing web page, an element of a repeated pattern having a reference to a web page along with text surrounding the reference; and   for each occurrence of a repeated pattern, identifying a dominant reference to a web page; and
 extracting the text surrounding the dominant reference as explanatory text for the referenced web page. 
   
   
   
       13 . The computer-readable storage medium of  claim 12  wherein the elements of the referencing web page are hierarchically organized as nodes and wherein the identifying of repeated patterns identifies a reference explanatory text node as a collection of adjacent, sibling nodes with a subtree of one node containing a reference node with surrounding text. 
   
   
       14 . The computer-readable storage medium of  claim 13  wherein the identifying of repeated patterns identifies a reference explanatory text region as a collection of adjacent, sibling reference explanatory text nodes that have the same length and that are within a threshold edit distance. 
   
   
       15 . The computer-readable storage medium of  claim 14  wherein the threshold edit distance varies based on number of block nodes within the reference explanatory text nodes. 
   
   
       16 . The computer-readable storage medium of  claim 12  wherein a reference explanatory text node has a dominant reference node when it has only one reference node that is a block node with a unique subtree structure. 
   
   
       17 . The computer-readable storage medium of  claim 12  including generating a summary of the referenced web page from the extracted text that surrounds references to the referenced web page. 
   
   
       18 . A computer system for identifying explanatory text for a referenced web page from a referencing web page, comprising:
 a memory storing computer-executable instructions of:
 a component that identifies repeated patterns of elements within the referencing web page, an occurrence of a repeated pattern having a reference to a web page along with text surrounding the reference; 
 a component that identifies a dominant reference for each repeated pattern; and 
 a component that extracts text surrounding the dominant reference as explanatory text for the referenced web page, and 
   a processor executing the computer-executable instructions stored in the memory.   
   
   
       19 . The computer system of  claim 18  wherein the occurrences of a repeated pattern have a similarity that is within a similarity threshold that varies based on whether an occurrence contains a block element. 
   
   
       20 . The computer system of  claim 18  wherein elements of the referencing web page are hierarchically organized as nodes and wherein the identifying of repeated patterns identifies a reference explanatory text node as a collection of adjacent, sibling nodes with a subtree of one node containing a reference node with surrounding text and identifies a reference explanatory text region as a collection of adjacent, sibling reference explanatory text nodes that have the same length and are similar. 
   
   
       21 . A method in a computing device for identifying explanatory text for a referenced display page from a referencing display page, comprising:
 identifying by the computing device repeated patterns of elements within a display page by comparing elements of the display page to other elements of the display page, a repeated pattern having a reference to a referenced display page along with text associated with the reference;   for each identified repeated pattern a dominant anchor, identifying a dominant anchor of the identified repeated pattern that is a reference to a referenced display page along with text associated with the reference; and   for each identified dominant anchor, extracting the text associated with the identified dominant anchor, wherein the extracted text represents explanatory text for the display page referenced by the identified dominant anchor.   
   
   
       22 . The method of  claim 21  wherein patterns of elements are considered to be repeated when the patterns have the same number of elements and the patterns have an edit distance that is within a threshold. 
   
   
       23 . The method of  claim 21  wherein a display page is represented as a tag tree with nodes representing elements and the identifying of repeated patterns identifies a reference explanatory text node as a collection of adjacent, sibling nodes with a subtree of one node containing a reference node with associated text. 
   
   
       24 . The method of  claim 23  wherein the identifying of repeated patterns includes identifying a reference explanatory text region as a collection of adjacent, sibling reference explanatory text nodes that have the same length and are similar. 
   
   
       25 . The method of  claim 21  including ranking the display page based on the identified explanatory text. 
   
   
       26 . The method of  claim 21  wherein when an occurrence of a repeated pattern includes multiple references with associated text, designating that the occurrence does not have a dominant anchor. 
   
   
       27 . The method of  claim 21  wherein when an occurrence of a repeated pattern includes only one reference with associated text, designating that the occurrence has a dominant anchor. 
   
   
       28 . The method of  claim 21  including using the identified explanatory text when crawling the display page. 
   
   
       29 . The method of  claim 21  including using the identified explanatory text for query refinement. 
   
   
       30 . The method of  claim 21  wherein the display page is a web page.

Join the waitlist — get patent alerts

Track US2010049772A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.