US2006242266A1PendingUtilityA1

Rules-based extraction of data from web pages

Assignee: KEEZER PAULAPriority: Feb 27, 2001Filed: Apr 24, 2006Published: Oct 26, 2006
Est. expiryFeb 27, 2021(expired)· nominal 20-yr term from priority
G06Q 30/0641G06F 16/958
46
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A rule creation application uses a reference web page, and user input regarding information displayed thereon, to generate a rule for extracting such information from the web page. The rule uses a structured graph representation of the web page, such as the page's Document Object Model (DOM), to extract the information. In addition to being applicable to the reference web page, the rule may be used to extract information from other web pages that have a similar structure.

Claims

exact text as granted — not AI-modified
1 - 20 . (canceled)  
   
   
       21 . A computer program comprising executable instructions represented in computer storage, said computer program adapted to be executed on a computer, and being capable of causing the computer to: 
 receive user input specifying a selected data element of a web page displayed on the computer;    identify, in a structural graph representation of said web page, a node that corresponds to the selected data element;    identify a path to said node, said path specifying how to traverse the structural graph representation of the web page to reach the node; and    generate a rule that is adapted to be applied to the web page, and to other web pages of similar structure, to extract web page data corresponding to the selected data element, the rule specifying said path.    
   
   
       22 . The computer program of  claim 21 , wherein the structural graph representation of the web page is a Document Object Model (DOM) representation of the web page.  
   
   
       23 . The computer program of  claim 21 , wherein the computer program is additionally capable of causing the computer to identify, among a set of attributes of said node, an attribute that contains the data element, and to include an identifier of said attribute in the rule.  
   
   
       24 . The computer program of  claim 21 , wherein the computer program is additionally capable of causing the network-connected computer to receive user input of, and to include in said rule, an expected value associated with the selected data element, said expected value being a value that must be present for the rule to extract associated data from a web page.  
   
   
       25 . The computer program of  claim 21 , wherein the computer program is additionally capable of causing the computer to automatically transmit the rule over a network to a server.  
   
   
       26 . The computer program of  claim 25 , in combination with the server, wherein the server is configured to store the rule in association with an address of an associated web site, and to serve the rule to user computers to enable the user computers to extract data from web pages of the web site.  
   
   
       27 . The computer program of  claim 26 , wherein the server is additionally configured to use the data extracted by a user computer via said rule to identify an item represented on a web page, and to transmit supplemental information about the item to the user computer for display to a user.  
   
   
       28 . The computer program of  claim 21 , in combination with a client program that is adapted to run in conjunction with a web browser, said client program adapted to apply the rule to a web page to extract item-identifying data therefrom, and to use the extracted item-identifying information to retrieve and display supplemental information about an identified item.  
   
   
       29 . The computer program of  claim 21 , wherein the computer program comprises JavaScript code.  
   
   
       30 . The computer program of  claim 21 , wherein the computer program is a browser plug-in.  
   
   
       31 . The computer program of  claim 21 , wherein the computer program is configured to display a rule-creation user interface that provides functionality for creating rules based on web pages loaded in a browser window.  
   
   
       32 . A computer-implemented method, comprising: 
 receiving user input specifying selected content of a web page displayed on a computer;    identifying, in a structural graph representation of said web page, a node that corresponds to the selected content;    identifying a path to said node, said path specifying how the structural graph representation of the web page is traversable to reach the node; and    generating a rule that is adapted to be applied to the web page, and to other web pages of similar structure, to extract data corresponding to said selected content, said rule specifying said path.    
   
   
       33 . The method of  claim 32 , wherein the structural graph representation of the web page is a Document Object Model (DOM) representation of the web page.  
   
   
       34 . The method of  claim 32 , further comprising applying the rule to at least one additional web page to extract data therefrom.  
   
   
       35 . The method of  claim 32 , further comprising identifying, among a set of attributes of said node, an attribute that contains the selected content, and including an identifier of said attribute in the rule.  
   
   
       36 . The method of  claim 32 , further comprising receiving user input of, and including in said rule, an expected value associated with the selected content.  
   
   
       37 . The method of  claim 32 , further comprising transmitting the rule over a network to a server, and storing the rule on the server in association with an address of a corresponding web site  
   
   
       38 . The method of  claim 37 , further comprising transmitting the rule from said server to a user computer, and on said user computer, applying the rule to a second web page to extract item-identifying data therefrom.  
   
   
       39 . The method of  claim 38 , further comprising using the item-identifying data to retrieve, and display on the user computer, supplemental information about an item represented on the second web page.  
   
   
       40 . The method of  claim 32 , wherein the method is implemented, at least in part, via execution of JavaScript code executed by a web browser.  
   
   
       41 . The method  claim 32 , wherein the method is implemented, at least in part, via execution of a browser plug-in.  
   
   
       42 . A rule generated according to the method of  claim 32  represented in computer storage.  
   
   
       43 . A computer system, comprising: 
 a rule generation component that runs on a user computer in conjunction with a browser program, the rule generation component configured to receive user input specifying a data element of a web page displayed on the user computer by the browser program, and to use a Document Object Model (DOM) representation of the web page to generate a rule for extracting the data element from the web page; and    a server system that communicates with the rule generation component over a network, and stores rules generated by the rule generation component in association with web site addresses to which such rules correspond.    
   
   
       44 . The computer system of  claim 43 , wherein the rule generation component is capable of generating a rule that is adapted to be applied to a plurality of different web pages of similar DOM structure.  
   
   
       45 . The computer system of  claim 43 , further comprising a data extraction component that retrieves rules from the server system, and applies the rules to web pages to extract data therefrom.

Join the waitlist — get patent alerts

Track US2006242266A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.