US2015134574A1PendingUtilityA1

Self-learning methods for automatically generating a summary of a document, knowledge extraction and contextual mapping

Assignee: YASIN SYEDPriority: Feb 3, 2010Filed: Jan 21, 2015Published: May 14, 2015
Est. expiryFeb 3, 2030(~3.5 yrs left)· nominal 20-yr term from priority
Inventors:Syed Yasin
G06F 16/93G06N 7/00G06F 16/345G06F 17/30011G06N 99/005G06N 20/00G06F 16/285
37
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Advance Machine Learning or Unsupervised Machine Learning Techniques are provided that relate to Self-learning processes by which a machine generates a sensible automated summary, extracts knowledge, and extracts contextually related Topics along with the justification that explains “why they are related” automatically without any human intervention or guidance (backed ontology's) during the process. Such processes also relate to generating a 360-Degree Contextual Result (360-DCR) using Auto-summary, Knowledge Extraction and Contextual Mapping.

Claims

exact text as granted — not AI-modified
We claim: 
     
         1 . A self learning method for automatically generating a Summary of a document without human intervention, said method comprising acts of:
 extracting Important Words (IW) of the document based on incremental order of their occurrence;   listing the order of the IW's extracted in the order of highest Word Group (WG), wherein the highest word group is combination of maximum number of words that go together as one word;   for each IW's starting in the order of highest word group, analyzing every sentence in the document to determine presence of the IW and thereafter extracting all the sentences having corresponding IW as important sentences (IS) after eliminating redundancies to generate the auto-summary for the document.   
     
     
         2 . The method as claimed in  claim 1 , wherein identifying WG and IW comprises using word article and punctuation marks from the document. 
     
     
         3 . A self learning method for automatically extracting knowledge of a given set of documents without human intervention, said method comprising acts of
 extracting Important Words (IW) and their corresponding Important Sentences (IS), and Topics (T) of the documents in a predetermined order;   eliminating duplicates of each extracted IW and its corresponding sentences; and   clustering the IS's and Topic's (T) in the list based on the extracted IW's as “Contextual-Topical Cluster” and “Knowledge-cluster” to extract knowledge and related contextual Topics from the set of given documents.   
     
     
         4 . The method as claimed in  claim 3 , wherein defining topic to each documents is done by comparing each IW in the document with its file name and Title name, if any of the IW's matches than that is defined as a Topic. 
     
     
         5 . The method as claimed in  claim 4 , wherein the IW's with highest frequency occurrences in the document is defined as the topic of the document. 
     
     
         6 . The method as claimed in  claim 3 , wherein eliminating the duplicates of each extracted IW and its corresponding sentences using Hashing technique. 
     
     
         7 . A self-learning method for automatically displaying 360-degree Contextual Search Results without human intervention, said method comprising acts of:
 generating Topic (T), Important-Words (IW), Important Sentences (IS) and Auto-Summary (SY) for a given document;   storing the generated Auto-Summary as a field value during indexing along with corresponding Topic and Content of the document in Master-Index;   extracting Topic List by processing the Master-Index and thereafter 360-degree Contextual Mapping (360-DCM) into 360-DCM cluster;   extracting Knowledge from the document into Knowledge Extraction (KE) cluster; and   analyzing user query to identify Topic in the TL and corresponding 360-DCM cluster to return related Topics along with the relationship map, wherein the Master-index returns search results along with auto-summary for each result; and the KE cluster returns relevant knowledge for the search query to display 360-degree Contextual Search Results.   
     
     
         8 . The method as claimed in  claim 7 , wherein the generating auto-summary comprises acts of:
 extracting Important Words (IW) of the document from incremental order of their occurrence;   listing the order of the IW's extracted in the order of highest Word Group (WG);   splitting the document into sentences and storing the sentences in a sequential order as Array of Sentences (AS);   for each IW's starting in the order of highest word group, analyzing every sentence in the AS to determine presence of the IW and thereafter extracting all the sentences having IW as important sentences (IS) to eliminate redundancies; and   removing the extracted sentences from the AS and corresponding IW from the list of IW's to generate the auto-summary for the document.   
     
     
         9 . The method as claimed in  claim 7 , wherein the extracting Knowledge from the document into Knowledge Extraction (KE) cluster comprises acts of extracting Important Words (IW) and their corresponding Important Sentences (IS), and Topics (T) of the documents in a predetermined order; eliminating duplicates of each extracted IW and its corresponding sentences; and clustering the IS's and Topic's (T) in the list based on the extracted IW's or the Topic as “Contextual-Topical Cluster” and “Knowledge-cluster” to extract knowledge and related contextual Topics from the set of given documents. 
     
     
         10 . The method as claimed in  claim 7 , wherein generating 360-degree Contextual Mapping comprises acts of indexing one or more documents; storing the topics identified for each documents in a predetermined order as Topical List (TL) during the indexing process and removing duplicates topics from the TL; extracting predefined number of results for each Topic in the TL by searching one Topic at a time in the index; extracting corresponding Topic and Content for each of the extracted result and storing the extracted Topic and Content in a predetermined order as Result-List (RL) in a temporary storage; analyzing the RL for each topic to extract Related Topic; analyzing corresponding document Content of the Topic to extract “why they are related” phrases from the content; and clustering the resultant “Related Topics” along with their respective sentences to generate 360-degree Contextual Mapping.

Join the waitlist — get patent alerts

Track US2015134574A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.