Information retrieval system and information retrieving method therefor
Abstract
To provide an information retrieval system capable of easily finding a site similar to a users favorite site without any difference in retrieval result obtained for each user and in steps of obtaining information. HTML file obtaining means obtains an HTML file from a Web site in an Internet. Retrieval key extraction means analyzes contents of the HTML file indicated by a URL specified by the user, and extracts a keyword as a retrieval key. Retrieval result storage means retrieves an index table based on the extracted retrieval key, and stores the retrieval result. Retrieval result display means reforms the retrieval result for visibility for the user and outputs the result. Score computation means computes the scores of the HTML tag and the keyword. Index table storage means stores an extracted index.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An information retrieval system which retrieves a record site of contents represented by a hypertext file, comprising:
extraction means for extracting keywords from an externally specified hypertext file; and retrieval means for retrieving a record site of the contents using said keywords extracted by said extraction means.
2 . The information retrieval system according to claim 1 , wherein said extraction means extracts said keywords from character strings specified by predetermined control information contained in said externally specified hypertext file.
3 . The information retrieval system according to claim 1 , further comprising computation means for computing scores indicating priorities for said keywords extracted by said extraction means.
4 . The information retrieval system according to claim 3 , wherein said computation means selects the keywords to be used as a retrieval key from said extracted keywords by assigning said scores by assigning predetermined weights to predetermined control information and said keywords extracted from character strings specified by the control information.
5 . The information retrieval system according to claim 4 , further comprising storage means for storing the control information and said keywords for which said scores are computed by said computation means after associating said keywords with the hypertext file from which said keywords are extracted,
wherein said retrieval means retrieves a record site of the contents by searching said storage means.
6 . The information retrieval system according to claim 2 , wherein said extraction means extracts tag information contained in said hypertext file as said control information, and extracts said keywords from the character strings specified by the tag information.
7 . An information retrieving method which retrieves a record site of contents represented by a hypertext file, comprising the steps of:
extracting keywords from an externally specified hypertext file; and retrieving a record site of the contents using said extracted keywords.
8 . The information retrieving method according to claim 7 , further comprising a computation step of computing scores indicating priorities for said extracted keywords and tag information contained in said externally specified hypertext file.
9 . The information retrieving method according to claim 8 , wherein said computation step assigns higher scores to more important HTML (hypertext markup language) tags and keywords, and lower scores to less important HTML tags and keywords so that a retrieval key can be selected as a significant index.
10 . The information retrieving method according to claim 9 , wherein storage means storing said HTML tags and said keywords assigned said scores after associating said keywords with the HTML file from which said keywords are extracted is searched so that a record site of the contents can be retrieved.Join the waitlist — get patent alerts
Track US2003088559A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.