US2005273706A1PendingUtilityA1

Systems and methods for identifying and extracting data from HTML pages

Assignee: YAHOO INCPriority: Aug 24, 2000Filed: May 4, 2005Published: Dec 8, 2005
Est. expiryAug 24, 2020(expired)· nominal 20-yr term from priority
Inventors:Udi ManberQi Lu
G06F 16/951Y10S707/99931Y10S707/99936
48
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for analyzing HTML formatted web pages to automatically identify and extract desired information. A computer algorithm identifies and extracts different pieces of information from different web pages automatically after minimal manual setup. The algorithm automatically analyzes pages with different content if they have the same, or similar, formats. The algorithm is fast and efficient and performs the extraction process quickly in real-time. The systems and methods are useful to build databases from unstructured web information. The algorithm can be used as an agent that captures information about products, and compares prices or other characteristics. It can also be used to populate structured databases that, given the different pieces of information, can analyze products and their characteristics. And it can also be used for data mining applications looking for patterns useful for marketing analyses, or other uses.

Claims

exact text as granted — not AI-modified
1 . A computer implemented method of identifying desired content in HTML formatted web pages, comprising the steps of: 
 selecting a model page, wherein the model page includes content data and a plurality of HTML tags for formatting the content data;    identifying a first area of interest in the model page;    parsing the model page to generate a first string of symbols for the plurality of HTML tags, the generated symbols in the first string representing only HTML tags, wherein the first area of interest is identified by a first portion of the first string of symbols;    retrieving a second web page associated with a different URL than the model page;    parsing the second web page to generate a second string of symbols for a plurality of HTML tags of the second web page, the generated symbols in the second string representing only HTML tags; and    comparing the first and second symbol strings to determine whether the second string includes a second portion similar to the first portion of the first string, wherein the second portion corresponds to a second area of interest in the second page.    
   
   
       2 - 28 . (canceled)

Join the waitlist — get patent alerts

Track US2005273706A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.