US2007067320A1PendingUtilityA1

Detecting relationships in unstructured text

Assignee: IBMPriority: Sep 20, 2005Filed: Sep 20, 2005Published: Mar 22, 2007
Est. expirySep 20, 2025(expired)· nominal 20-yr term from priority
Inventors:Jasmine Novak
G06F 40/295G06F 16/36
47
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Disclosed are embodiments of a system and a method for detecting relationships described in unstructured text-based electronic documents. The system and method incorporate the use of an input file that contains one or more text patterns that represent particular relationships. The text patterns each include regular text expressions that describe the particular relationship and slots for the location of each entity in that relationship. Document(s) are selected by a user and scanned by a proper noun tagger that identifies and tags every occurrence of proper names within the document(s). Then, a pattern matcher scans the document(s) to match text patterns. If a text pattern is matched within a document a relationship detector extracts all pairs of proper names found in the slots for each matched text pattern. The output from the relationship detector includes the names for each entity in the relationship, the type of relationship, and the identity of the document and the location of the sentence describing the relationship in the document.

Claims

exact text as granted — not AI-modified
1 . A computer implemented method of detecting a relationship between a first entity and a second entity, said method comprising: 
 creating a text pattern that represents a type of relationship, wherein said text pattern comprises a first slot for said first entity and a second slot for said second entity;    analyzing a text-based document so as to locate said text pattern within said document;    determining a location for each proper name occurring within said document; and    extracting proper names located within said first slot and said second slot of said text pattern within said document, wherein said proper names located within said first slot and said second slot identify said first entity and said second entity.    
   
   
       2 . The method of  claim 1 , wherein said creating of said text pattern further comprises identifying a keyword describing said relationship and wherein said method further comprises before said analyzing of said document, reviewing said document to determine if said keyword is located in said document.  
   
   
       3 . The method of  claim 1 , wherein said creating of said text pattern further comprises: 
 creating at least one text expression comprising a plurality of words that describe said type of said relationship; and    setting a position of said first slot and said second slot relative to said at least one text expression.    
   
   
       4 . The method of  claim 1 , wherein said determining of said location of each of said proper names comprises: 
 scanning said document to identify each of said proper names occurring within said document based on a set of matching rules;    re-scanning said document to tag said location for each of said proper names identified; and    recording said location for each of said proper names.    
   
   
       5 . The method of  claim 4 , wherein said set of matching rules is based on at least one of word capitalization, sentence structure, sentence boundaries, and excluded words.  
   
   
       6 . The method of  claim 1 , wherein said creating of said text pattern further comprises defining an order of said first entity and said second entity in said relationship based on said locations of said proper names within said first slot and said second slot.  
   
   
       7 . The method of  claim 1 , further comprising storing a record of said relationship comprising at least one of said proper name of said first entity, said proper name of said second entity, said type of relationship between said first entity and said second entity, said order of said first entity and said second entity in said relationship, and an identifier for said document and a location in said document where said relationship is detected.  
   
   
       8 . A system for detecting a relationship between a first entity and a second entity, said system comprising: 
 an input file adapted to store a text pattern that describes a type of relationship, wherein said text pattern comprises a first slot for said first entity and a second slot for said second entity;    a pattern matcher in communication with said input file and adapted to analyze a text-based document so as to locate said text pattern within said document;    a proper noun tagger adapted to locate and record occurrences of proper names within said document; and    a relationship detector in communication with said pattern matcher and said proper noun tagger and adapted to extract said proper names located within said first slot and said second slot of said text pattern within said document so as to identify said first entity and said second entity and, thereby, detect said relationship.    
   
   
       9 . The system of  claim 8 , further comprising a keyword identifier in communication with said input file and adapted to review said document for said keyword and to forward said document to said pattern matcher only if said keyword is located in said document.  
   
   
       10 . The system of  claim 8 , wherein said text pattern further comprises: 
 at least one text expression comprising a plurality of words that describe said relationship; and    positions for said first slot and said second slot relative to said at least one text expression.    
   
   
       11 . The system of  claim 8 , wherein said text pattern further comprises an order of said first entity and said second entity in said relationship based said locations of said proper names within said first slot and said second slot.  
   
   
       12 . The method of  claim 8 , wherein said proper noun tagger is further adapted to scan said document to identify each of said proper names occurring within said document based on a set of matching rules, to re-scan said document to tag said location for each of said proper names, and to record said location for each of said proper names within said document.  
   
   
       13 . The system of  claim 12 , wherein said set of matching rules is based on at least one of word capitalization, sentence structure, sentence boundaries, and excluded words.  
   
   
       14 . The system of  claim 11 , further comprising at least one of a data storage device adapted to store at least one of said proper name of said first entity, said proper name of said second entity, said relationship between said first entity and said second entity, said order of said first entity and said second entity in said relationship, and a record of said document in which said relationship is detected.  
   
   
       15 . A program storage device readable by computer and tangibly embodying a program of instructions executable by said computer to perform a method of detecting a relationship between a first entity and a second entity, said method comprising: 
 creating a text pattern that represents a type of relationship, wherein said text pattern comprises a first slot for said first entity and a second slot for said second entity;    analyzing a text-based document so as to locate said text pattern within said document;    determining a location for each proper name occurring within said document; and    extracting proper names located within said first slot and said second slot of said text pattern within said document, wherein said proper names located within said first slot and said second slot identify said first entity and said second entity    
   
   
       16 . The program storage device of  claim 15 , wherein said creating of said text pattern further comprises identifying a keyword describing said relationship and wherein said method further comprises before said analyzing of said document, reviewing said document to determine if said keyword is located in said document.  
   
   
       17 . The program storage device of  claim 15 , wherein said creating of said text pattern further comprises: 
 creating at least one text expression comprising a plurality of words that describe said type of said relationship; and    setting a position of said first slot and said second slot relative to said at least one text expression.    
   
   
       18 . The program storage device of  claim 15 , wherein said determining of said location for each of said proper names comprises: 
 scanning said document to identify each of said proper names occurring within said document based on a set of matching rules;    re-scanning said document to tag said location for each of said proper names; and    recording said location for each of said proper names.    
   
   
       19 . The program storage device of  claim 18 , wherein said set of matching rules is based on at least one of word capitalization, sentence structure, sentence boundaries, and excluded words.  
   
   
       20 . The program storage device of  claim 15 , wherein said creating of said text pattern further comprises defining an order of said first entity and said second entity in said relationship based on said locations of said proper names within said first slot and said second slot.

Join the waitlist — get patent alerts

Track US2007067320A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.