System and Methodology for Real-time Content Aggregation and Syndication
Abstract
A system and methodology for real-time content aggregation and syndication is described. In one embodiment, for example, a method is described for assisting a user with extracting items relevant to search queries from documents including items of various types, the method comprises steps of: receiving a search query specifying a search phrase and a particular item type; identifying documents matching the search phrase; for each matching document, determining whether the document includes an item having the particular item type; and extracting items having the particular item type from the matching documents for display to the user. The solution enables a user to aggregate and syndicate content without a professional content manager or complicated content management software tools.
Claims
exact text as granted — not AI-modified1 . A method for assisting a user with extracting items relevant to search queries from documents including items of various types, the method comprising:
receiving a search query specifying a search phrase and a particular item type; identifying documents matching said search phrase; for each matching document, determining whether the document includes an item having said particular item type; and extracting items having said particular item type from the matching documents for display to the user.
2 . The method of claim 1 , wherein said documents comprise Web pages having searchable text.
3 . The method of claim 2 , wherein said Web pages include items of various types which may or may not have searchable text.
4 . The method of claim 1 , wherein said particular item type comprises a selected one of a headline, text, an article, a graphic object, an image, a byline, and a button.
5 . The method of claim 1 , wherein said identifying step includes generating a list of URLs identifying documents available on the Internet using one of an Internet search engine and a Web directory.
6 . The method of claim 1 , wherein said receiving step includes receiving a search phrase including one or more keywords.
7 . The method of claim 1 , wherein said determining step includes parsing a plurality of matching documents using a plurality of threads, so as to speed return of search results.
8 . The method of claim 1 , wherein a matching document comprises a Web page and said determining step includes parsing container objects of the Web page to determine attributes of each item included in the Web page.
9 . The method of claim 8 , wherein said determining step includes calculating a score based on attributes of each item for determining whether the item has said particular item type.
10 . The method of claim 1 , wherein said identifying step includes identifying documents matching said search phrase, without regard to whether those documents themselves comprise the particular item type.
11 . The method of claim 1 , wherein said extracting step includes aggregating a plurality of items extracted from the matching documents in a single document for display.
12 . The method of claim 11 , further comprising:
inserting additional items of content into the single document, the additional items of content selected based on the search query.
13 . The method of claim 12 , wherein said step of inserting additional items of content includes inserting advertising into the single document between items extracted from the matching documents.
14 . The method of claim 11 , wherein said single document is displayed to the user in a Web browser.
15 . A computer-readable medium having processor-executable instructions for performing the method of claim 1 .
16 . A method for generating a single document displaying items of content retrieved from one or more Web pages, the method comprising:
receiving a request for items of content, the request including keywords and extended attributes of items to be obtained; retrieving one or more Web pages based on the keywords; parsing each of the one or more Web pages into its component objects, each object representing an item of content from the given Web page; selecting particular objects matching the extended attributes of the request; and aggregating items of content corresponding to said particular objects into a single document for display.
17 . The method of claim 16 , wherein said method is performed at a client device.
18 . The method of claim 16 , wherein said method is performed by a Web browser application.
19 . The method of claim 16 , wherein said retrieving step includes retrieving Web pages using one of an Internet search engine and a Web directory to identify Web pages which may include requested items of content.
20 . The method of claim 16 , wherein said extended attributes include type of item that is requested.
21 . The method of claim 20 , wherein said type of item comprises a selected one of a headline, text, an article, a graphic object, an image, a byline, and a button.
22 . The method of claim 16 , wherein said extended attributes include item size.
23 . The method of claim 16 , wherein said parsing step includes parsing container objects of the Web page.
24 . The method of claim 23 , wherein said step of parsing container objects includes creating feature extraction objects for elements of the container objects based on attributes of said elements.
25 . The method of claim 24 , wherein said selecting step includes calculating a score for an item of content based on matching attributes of said feature extraction objects and extended attributes of the request.
26 . The method of claim 16 , wherein said single document is displayed to a user in a Web browser application.
27 . A computer-readable medium having processor-executable instructions for performing the method of claim 16 .
28 . A Web browser system for dynamically generating a page displaying items of content extracted from sources of content available on a network, the system comprising:
a user interface module for a user to navigate to sources of content available on the network, select particular items of content, and build a page composed of the particular items; a feature extraction module for automatically creating objects representing the particular items of content on the page built by the user; and a content collection module for dynamically generating the page by extracting the particular items of content from the sources of content via the network using the objects and aggregating the particular items for display on the page.
29 . The system of claim 28 , wherein the network comprises the Internet and the sources of content comprise Web pages available on the Internet.
30 . The system of claim 28 , further comprising:
a syndication module for sending the page built by the user to a given device, so as to enable the page to be dynamically generated on the given device.
31 . The system of claim 28 , wherein said feature extraction module generates an object based on attributes of a particular item of content.
32 . The system of claim 31 , wherein said feature extraction module parses container objects of a Web page to determine attributes of the particular item of content.
33 . The system of claim 32 , wherein said feature extraction module creates an object based on attributes of the particular item, the object facilitating dynamic access to the particular item via the network.
34 . The system of claim 28 , wherein the particular items comprise selected ones of headlines, text, articles, graphic objects, images, bylines, and buttons.
35 . The system of claim 28 , further comprising:
a search module for obtaining particular items of content available via the network in response to a search query and displaying said items in the user interface.
36 . The system of claim 35 , wherein said search query includes a search phrase and extended attributes and said search module locates a source of content based on said search phrase and obtains particular items of content based on said extended attributes.
37 . The system of claim 28 , wherein said Web browser system is stored on a computer-readable medium.
38 . A system for extracting items of content from documents available on the Internet in response to a search query, the system comprising:
means for receiving a search query comprising a search phrase and specified attributes of items of to be obtained; means for obtaining a list of relevant documents in response to the search query based on matching terms of the search phrase to terms contained in the documents; means for retrieving a relevant document on the list and parsing it into a plurality of objects; means for determining a score value for each of said plurality of objects, the score value based on matching attributes of the object with said specified attributes of the search query; and means for extracting a particular object having a score value indicating relevance to the search query from the relevant document.
39 . The system of claim 38 , wherein said system is implemented in a Web browser application.
40 . The system of claim 38 , wherein said plurality of objects comprise selected ones of headlines, text, articles, graphic objects, images, bylines, and buttons.
41 . The system of claim 38 , wherein said means for extracting includes means for aggregating said particular object with objects extracted from other relevant documents for display in a single page.
42 . The system of claim 41 , further comprising:
means for transmitting the single page to various devices for display.Join the waitlist — get patent alerts
Track US2006259462A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.