US2006074900A1PendingUtilityA1

Selecting keywords representative of a document

Individually held — no corporate assignee on recordPriority: Sep 30, 2004Filed: Sep 30, 2004Published: Apr 6, 2006
Est. expirySep 30, 2024(expired)· nominal 20-yr term from priority
G06F 16/3331G06F 16/332G06F 16/367
47
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The method makes use of a given ontology to select keywords representative of a given document. The method finds all the terms in an ontology that occur in a document, and computes their frequency of occurrences in the document. The method then propagates these values from the leaves upwards to the root of the ontology during which it weights them. The method then selects a subset of terms of the ontology structure as keywords representative of the document based on these weights.

Claims

exact text as granted — not AI-modified
1 . A method of selecting keywords representative of a document from an ontology, said method comprising: 
 computing, for each term in the ontology, a value representative of a frequency of occurrence of said term in the document; and    selecting a subset of terms of the ontology as keywords representative of the document based on said value.    
   
   
       2 . A method of selecting keywords representative of a document from an ontology, wherein the ontology comprises terms arranged in a tree-like structure, said method comprising: 
 computing, for each term in the ontology, a first value representative of a frequency of occurrence of said term in the document;    assigning said first value to corresponding vertices in the ontology;    propagating said first value from leaf vertices of the ontology upwards to the one or more root vertices of the ontology by assigning to each vertex a second value, wherein said second value equals a sum of said first value of the vertex plus the second values of immediate descendent vertices of said vertex each multiplied by a corresponding propagation factor; and    selecting k terms of the ontology as keywords representative of the document that have a largest k second value.    
   
   
       3 . A method of selecting keywords representative of a document from an ontology, wherein the ontology comprises terms arranged in a tree-like structure having one or more root vertices, vertices and leaf vertices, said method comprising: 
 computing, for each term in the ontology, a first value representative of a frequency of occurrence of said term in the document;    assigning first values to corresponding vertices in the ontology;    propagating said first values from the leaf vertices of the ontology upwards to the one or more root vertices of the ontology by assigning to each vertex a second value, wherein said second value equals a sum of said first value of the vertex plus the second values of immediate descendent vertices of said vertex each multiplied by a corresponding propagation factor;    generating a sub-structure of the ontology, wherein the sub-structure comprises a unique path for each term so as to disambiguates a context of the terms; and    performing an optimization process, wherein k vertices are selected such that a sum of weighted distances of all the vertices having non-zero second values to associated selected k vertices is minimized, and wherein k terms associated with the selected k vertices are selected as keywords representative of the document.    
   
   
       4 . The method of  claim 3 , wherein the optimization process comprises a greedy facility location process.  
   
   
       5 . The method of  claim 3 , wherein the optimization process comprises a greedy facility location process, wherein the vertices having non-zero second values are clients, the selected k vertices are facilities serving the clients, the weighted distance between a client and a facility is a number of edges of the tree-like structure between the client and the facility multiplied by a sum of the second values of the vertices in a subtree of the facility, wherein facilities can serve only descendent clients and clients can be served by multiple facilities.  
   
   
       6 . The method of  claim 3 , wherein the optimization process comprises an optimal dynamic programming based process.  
   
   
       7 . A method of selecting keywords representative of a document from an ontology, wherein the ontology comprises terms arranged in a tree-like structure having one or more root vertices, vertices and leaf vertices, said method comprising: 
 computing, for each term in the ontology, a first value representative of a frequency of occurrence of said term in the document;    assigning frequency of occurrence values to corresponding vertices in the ontology; and    performing an optimization process, wherein k vertices are selected such that a sum of weighted distances of all the vertices having non-zero first values to associated selected k vertices is minimized, and wherein k terms associated with the selected k vertices are selected as keywords representative of the document.    
   
   
       8 . The method of  claim 7 , wherein the optimization process comprises a greedy facility location process.  
   
   
       9 . The method of  claim 7 , wherein the optimization process comprises a greedy facility location process, wherein the vertices having non-zero second values are clients, the selected k vertices are facilities serving the clients, the weighted distance between a client and a facility is a number of edges of the tree-like structure between the client and the facility multiplied by a sum of the second values of the vertices in a subtree of the facility, wherein facilities can serve only descendent clients and clients can be served by multiple facilities.  
   
   
       10 . The method of  claim 7 , wherein the optimization process comprises an optimal dynamic programming based process.  
   
   
       11 . A computer program product for selecting keywords representative of a document from an ontology, the computer program product comprising computer software recorded on a computer-readable medium for performing a method comprising: 
 computing, for each term in the ontology, a value representative of a frequency of occurrence of said term in the document; and    selecting a subset of terms of the ontology as keywords representative of the document based on said value.    
   
   
       12 . A computer system for selecting keywords representative of a document from an ontology, the computer system comprising computer software recorded on a computer-readable medium for performing a method comprising: 
 computing, for each term in the ontology, a value representative of a frequency of occurrence of said term in the document; and    selecting a subset of terms of the ontology as keywords representative of the document based on said value.    
   
   
       13 . A computer program product for selecting keywords representative of a document from an ontology, the computer program product comprising computer software recorded on a computer-readable medium for performing a method comprising: 
 computing, for each term in the ontology, a first value representative of a frequency of occurrence of said term in the document;    assigning said first value to corresponding vertices in the ontology;    propagating said first value from leaf vertices of the ontology upwards to the one or more root vertices of the ontology by assigning to each vertex a second value, wherein said second value equals a sum of said first value of the vertex plus the second values of immediate descendent vertices of said vertex each multiplied by a corresponding propagation factor; and    selecting k terms of the ontology as keywords representative of the document that have a largest k second value.    
   
   
       14 . A computer system for selecting keywords representative of a document from an ontology, the computer system comprising computer software recorded on a computer-readable medium for performing a method comprising: 
 computing, for each term in the ontology, a first value representative of a frequency of occurrence of said term in the document;    assigning said first value to corresponding vertices in the ontology;    propagating said first value from leaf vertices of the ontology upwards to the one or more root vertices of the ontology by assigning to each vertex a second value, wherein said second value equals a sum of said first value of the vertex plus the second values of immediate descendent vertices of said vertex each multiplied by a corresponding propagation factor; and    selecting k terms of the ontology as keywords representative of the document that have a largest k second value.    
   
   
       15 . A computer program product for selecting keywords representative of a document from an ontology, the computer program product comprising computer software recorded on a computer-readable medium for performing a method comprising: 
 computing, for each term in the ontology, a first value representative of a frequency of occurrence of said term in the document;    assigning first values to corresponding vertices in the ontology;    propagating said first values from the leaf vertices of the ontology upwards to the one or more root vertices of the ontology by assigning to each vertex a second value, wherein said second value equals a sum of said first value of the vertex plus the second values of immediate descendent vertices of said vertex each multiplied by a corresponding propagation factor;    generating a sub-structure of the ontology, wherein the sub-structure comprises a unique path for each term so as to disambiguates a context of the terms; and    performing an optimization process, wherein k vertices are selected such that a sum of weighted distances of all the vertices having non-zero second values to associated selected k vertices is minimized, and wherein k terms associated with the selected k vertices are selected as keywords representative of the document.    
   
   
       16 . A computer system for selecting keywords representative of a document from an ontology, the computer system comprising computer software recorded on a computer-readable medium for performing a method comprising: 
 computing, for each term in the ontology, a first value representative of a frequency of occurrence of said term in the document;    assigning first values to corresponding vertices in the ontology;    propagating said first values from the leaf vertices of the ontology upwards to the one or more root vertices of the ontology by assigning to each vertex a second value, wherein said second value equals a sum of said first value of the vertex plus the second values of immediate descendent vertices of said vertex each multiplied by a corresponding propagation factor;    generating a sub-structure of the ontology, wherein the sub-structure comprises a unique path for each term so as to disambiguates a context of the terms; and    performing an optimization process, wherein k vertices are selected such that a sum of weighted distances of all the vertices having non-zero second values to associated selected k vertices is minimized, and wherein k terms associated with the selected k vertices are selected as keywords representative of the document.    
   
   
       17 . A computer program product for selecting keywords representative of a document from an ontology, the computer program product comprising computer software recorded on a computer-readable medium for performing a method comprising: 
 computing, for each term in the ontology, a first value representative of a frequency of occurrence of said term in the document;    assigning frequency of occurrence values to corresponding vertices in the ontology; and    performing an optimization process, wherein k vertices are selected such that a sum of weighted distances of all the vertices having non-zero first values to associated selected k vertices is minimized, and wherein k terms associated with the selected k vertices are selected as keywords representative of the document.    
   
   
       18 . A computer system for selecting keywords representative of a document from an ontology, the computer system comprising computer software recorded on a computer-readable medium for performing a method comprising: 
 computing, for each term in the ontology, a first value representative of a frequency of occurrence of said term in the document;    assigning frequency of occurrence values to corresponding vertices in the ontology; and    performing an optimization process, wherein k vertices are selected such that a sum of weighted distances of all the vertices having non-zero first values to associated selected k vertices is minimized, and wherein k terms associated with the selected k vertices are selected as keywords representative of the document.

Join the waitlist — get patent alerts

Track US2006074900A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.