US2023119422A1PendingUtilityA1

Cluster analysis method, cluster analysis system, and cluster analysis program

Assignee: AIXS INCPriority: May 17, 2019Filed: Nov 28, 2022Published: Apr 20, 2023
Est. expiryMay 17, 2039(~12.8 yrs left)· nominal 20-yr term from priority
G06F 16/358G06V 30/40G06F 16/355G06F 40/216G06F 16/313
55
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A server 4 executes a similarity calculation step (S2) of calculating similarity between content of one document and content of another document, a cluster classification step (S3) of generating a network in which a document is set as a node based on calculated similarity and similar nodes are connected by an edge, and performing classification based on similar documents, a first index calculation step (S4) of calculating a first index indicating centrality of a document in the network, a second index calculation step (S5) of calculating a second index that is different from the first index in the network and indicates importance of a document, and a display data generation step (S6) of generating, regarding a document, first display data indicating the network by an expression of a size of an object of a node according to the first index, an expression of a gauge having a shape corresponding to a shape of the object according to the second index and a length of the gauge, an expression according to a type of the cluster, and an expression according to magnitude of similarity between documents.

Claims

exact text as granted — not AI-modified
1 . A cluster analysis method in which a computer classifies a plurality of documents into clusters according to content of the documents and generates display data indicating a relationship between documents, the cluster analysis method comprising:
 a similarity calculation step of calculating similarity between content of one document and content of another document;   a cluster classification step of generating a network in which a document or a cluster is set as a node based on calculated similarity and similar nodes are connected by an edge, and classifying similar documents into clusters;   a first index calculation step of calculating a first index indicating centrality of a document in the network;   a second index calculation step of calculating a second index that is different from the first index in the network and indicates importance of a document; and   a display data generation step of generating, regarding a document, first display data indicating the network by an expression of a size of an object of a node according to the first index, an expression of a gauge having a shape corresponding to a shape of the object according to the second index and a length of the gauge, an expression according to a type of the cluster, and an expression according to magnitude of similarity between documents,   wherein the display data generation step is capable of expressing an expression according to magnitude of similarity between the documents by thickness of a line connecting documents and displaying the network in enlarged and reduced manners, and generates the first display data by increasing or decreasing number of displayed lines according to the enlarged and reduced display.   
     
     
         2 . The cluster analysis method according to  claim 1 , wherein the display data generation step generates display data in which an object of a first index is represented by a circle, and a gauge of the second index is represented by an arc concentric with the circle of the first index and a length of the arc. 
     
     
         3 - 4 . (canceled) 
     
     
         5 . The cluster analysis method according to  claim 1 , wherein the document is a document published on an academic journal, and the second index is calculated according to citation of the document. 
     
     
         6 . The cluster analysis method according to  claim 1 , wherein the document is a document described on a website acquired by web search up to a predetermined number of items. 
     
     
         7 . The cluster analysis method according to  claim 6 , wherein the second index is calculated according to number of accesses to the website. 
     
     
         8 - 9 . (canceled) 
     
     
         10 . The cluster analysis method according to  claim 1 , further comprising a step of designating a word from those having a high appearance frequency included in the document, excluding the document including the designated word from the target of analysis and performing analysis again. 
     
     
         11 . The cluster analysis method according to  claim 1 , further comprising a step of designating a word from those having a high appearance frequency included in the document and generating first display data for highlighting, on a network, a node indicating a document or a cluster including the designated word. 
     
     
         12 . The cluster analysis method according to  claim 1 , wherein the display data generation step determines arrangement of documents on the network by using a dynamic model so that a plurality of documents are not displayed in an overlapping manner. 
     
     
         13 . (canceled) 
     
     
         14 . A cluster analysis system that classifies a plurality of documents into clusters according to content of the documents and generates display data indicating a relationship between documents, the cluster analysis system comprising:
 a similarity calculation unit that calculates similarity between content of one document and content of another document;   a cluster classification unit that generates a network in which a document is set as a node based on calculated similarity and similar nodes are connected by an edge, and classifies similar documents into clusters;   a first index calculation unit that calculates a first index indicating centrality of a document in the network;   a second index calculation unit that calculates a second index that is different from the first index in the network and indicates importance of a document; and   a display data generation unit that generates, regarding a document, first display data indicating the network by an expression of a size of an object of a node according to the first index, an expression of a gauge having a shape corresponding to a shape of the object according to the second index and a length of the gauge, an expression according to a type of the cluster, and an expression according to magnitude of similarity between documents,   wherein the display data generation step is capable of expressing an expression according to magnitude of similarity between the documents by thickness of a line connecting documents and displaying the network in enlarged and reduced manners, and generates the first display data by increasing or decreasing number of displayed lines according to the enlarged and reduced display.   
     
     
         15 . A cluster analysis program that causes a computer to classify a plurality of documents into clusters according to content of the documents and generate display data indicating a relationship between documents, and to execute:
 a similarity calculation step of calculating similarity between content of one document and content of another document;   a cluster classification step of generating a network in which a document is set as a node based on calculated similarity and similar nodes are connected by an edge, and classifying similar documents into clusters;   a first index calculation step of calculating a first index indicating centrality of a document in the network;   a second index calculation step of calculating a second index that is different from the first index in the network; and   
       a display data generation step of generating, regarding a document, first display data indicating the network by an expression of a size of an object of a node according to the first index, an expression of a gauge having a shape corresponding to a shape of the object according to the second index and a length of the gauge, an expression according to a type of the cluster, and an expression according to magnitude of similarity between documents,
 wherein the display data generation step is capable of expressing an expression according to magnitude of similarity between the documents by thickness of a line connecting documents and displaying the network in enlarged and reduced manners, and generates the first display data by increasing or decreasing number of displayed lines according to the enlarged and reduced display. 
 
     
     
         16 . The cluster analysis method according to  claim 1 , wherein the display data generation step merges a plurality of adjacent nodes having high similarity in accordance with enlarged display and reduced display of the network.

Join the waitlist — get patent alerts

Track US2023119422A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.