US2008259084A1PendingUtilityA1

Method and apparatus for organizing data sources

Assignee: IBMPriority: Aug 14, 2006Filed: Jun 27, 2008Published: Oct 23, 2008
Est. expiryAug 14, 2026(~0.1 yrs left)· nominal 20-yr term from priority
G06F 16/35Y10S707/99953Y10S707/99933
53
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method and apparatus for organizing deep Web services are provided. In one aspect, the method and apparatus obtains a collection of sources and their associated attributes and/or input modes, for instance, using a crawling algorithm. The method and apparatus uses this information to organize the sources into communities. A mining algorithm such as the hyperclique mining algorithm is used to obtain cliques of highly correlated attributes. A clustering algorithm such as the hierarchical agglomerative clustering algorithm is used to further cluster the cliques of attributes into larger cliques, which in the present disclosure is referred to as signatures. The sources that are associated with each signature form a community and a graph representation of the communities is constructed, where the vertices are communities and the edges are the shared attributes.

Claims

exact text as granted — not AI-modified
1 . A computer-implemented method of organizing data sources, comprising:
 grouping a plurality of items including input attributes, output attributes, or keywords or combination thereof from one or more sources into a plurality of cliques of highly correlated items;   clustering the plurality of cliques into one or more signatures; and   for each of the one or more signatures,
 selecting one or more sources that are associated with a signature; and 
 forming the selected sources into a community. 
   
   
   
       2 . The method of  claim 1 , further including:
 constructing a graph representation of a plurality of communities, the graph representation including at least a plurality of vertices representing the plurality of communities respectively and one or more edges connecting the plurality of vertices, the one or more edges representing one or more input attributes, output attributes, or keywords or combination thereof that are shared between the communities represented in the connecting vertices.   
   
   
       3 . The method of  claim 2 , further including:
 navigating the graph representation including at least one of:   starting at a vertex representing a community, following one or more edges to one or more second vertices representing one or more related communities;   starting at a source, traversing to one or more associated vertices; and   starting from an attribute or a keyword or combination thereof, traversing one or more associated edges and connected vertices.   
   
   
       4 . The method of  claim 1 , wherein the step of grouping is performed using a hyperclique mining algorithm. 
   
   
       5 . The method of  claim 1 , wherein the step of clustering is performed using a hierarchical agglomerative clustering algorithm. 
   
   
       6 . The method of  claim 1 , further including:
 obtaining the plurality of items including input attributes, output attributes, or keywords or combination thereof from one or more sources using a crawling algorithm.   
   
   
       7 . An apparatus for organizing data sources, comprising:
 a means for grouping a plurality of items including input attributes, output attributes, or keywords or combination thereof from one or more sources into a plurality of cliques of highly correlated items; and   a means for clustering the plurality of cliques into one or more signatures,   a means for selecting one or more sources that are associated with a signature and forming the selected sources into a community for each of the one or more signatures; and   a means for constructing a graph representation of a plurality of communities, the graph representation including at least a plurality of vertices representing the plurality of communities respectively and one or more edges connecting the plurality of vertices, the one or more edges representing one or more input attributes, output attributes, or keywords or combination thereof that are shared between the communities represented in the connecting vertices.

Join the waitlist — get patent alerts

Track US2008259084A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.