Framework for large-scale multi-label classification
Abstract
A framework for large-scale multi-label classification of an electronic document is described. An example multi-label classification system is configured to identify seed labels that represent respective one or more candidate content topics associated with the electronic document and determine additional labels based on the seed labels and label correlation data derived from member profiles maintained by an on-line social network system. The multi-label classification system then constructs a graph comprising nodes that correspond to the seed labels and the additional labels. A clustering algorithm is applied to the constructed graph to produce a labels graph. The labels graph is deemed to include nodes that correspond to topics discussed or referenced in the electronic document.
Claims
exact text as granted — not AI-modified1 . A method comprising:
accessing an electronic document; identifying one or more seed labels, the one or more seed labels representing respective one or more preliminary content topics associated with the electronic document; generating a seed graph, nodes of the seed graph representing the one or more seed labels; based on co-occurrence of labels in member profiles of an on-line social networking system and based on the one or more seed labels, deriving one or more additional labels; using at least one processor, generating an expanded graph comprising a first set of nodes representing the one or more additional labels and a second set of nodes representing the one or more seed labels; applying a clustering algorithm to the expanded graph to generate a labels graph; and identifying nodes of the labels graph, as a set of resolved content topics associated with the electronic document.
2 . The method of claim 1 , wherein an edge in the expanded graph has an edge weight constructed as correlation between labels represented by nodes connected to the edge.
3 . The method of claim 1 , wherein the edge weight is a directed edge weight.
4 . The method of claim 1 , wherein the generating of the expanded graph comprises assigning respective weights to nodes of the expanded graph.
5 . The method of claim 1 , wherein the labels in the member profiles of the on-line social networking system correspond to entries in a dictionary of skills maintained by the on-line networking system, the entries comprising respective words and phrases describing professional skills of respective members of the on-line social networking system.
6 . The method of claim 1 , wherein the one or more weak classifiers comprises a conditional random field based tagger (CRF Tagger).
7 . The method of claim 1 , wherein the one or more weak classifiers comprises a maximum entropy based text classifier.
8 . The method of claim 1 , wherein the clustering algorithm is based on a minimum cut and maximum flow algorithm.
9 . The method of claim 1 , wherein the clustering algorithm is Markov clustering algorithm.
10 . The method of claim 1 , wherein the clustering algorithm is normalized graph cut algorithm.
11 . A computer-implemented system comprising:
an access module, implemented using at least one processor, to access an electronic document; an identifying module, implemented using at least one processor, to identify one or more seed labels, the one or more seed labels representing respective one or more preliminary content topics associated with the electronic document; a graph generator, implemented using at least one processor, to generate a seed graph, nodes of the seed graph representing the one or more seed labels; an expanded nodes detector, implemented using at least one processor, to derive one or more additional labels based on co-occurrence of labels in member profiles of an on-line social networking system and based on the one or more seed labels; an expanded graph generator, implemented using at least one processor, to generate an expanded graph comprising a first set of nodes representing the one or more additional labels and a second set of nodes representing the one or more seed labels; a graph cutting module, implemented using at least one processor, to apply a clustering algorithm to the expanded graph to generate a labels graph, using the at least one processor; and a resolved labels module, implemented using at least one processor, to identify nodes of the labels graph as a set of resolved content topics associated with the electronic document, using the at least one processor.
12 . The system of claim 11 , wherein an edge in the expanded graph has an edge weight constructed as correlation between labels represented by nodes connected to the edge.
13 . The system of claim 11 , wherein the edge weight is a directed edge weight.
14 . The system of claim 11 , wherein the generating of the expanded graph comprises assigning respective weights to nodes of the expanded graph.
15 . The system of claim 11 , wherein the labels in the member profiles of the on-line social networking system correspond to words and phrases describing professional skills of respective members of the on-line social networking system.
16 . The system of claim 11 , wherein the one or more weak classifiers comprises a conditional random field based tagger (CFR Tagger).
17 . The system of claim 11 , wherein the one or more weak classifiers comprises a maximum entropy based text classifier.
18 . The system of claim 11 , wherein the clustering algorithm is based on a minimum cut and maximum flow algorithm.
19 . The system of claim 11 , wherein the clustering algorithm is Markov clustering algorithm.
20 . A machine-readable non-transitory storage medium having instruction data to cause a machine to perform operations comprising:
accessing an electronic document; identifying one or more seed labels, the one or more seed labels representing respective one or more preliminary content topics associated with the electronic document; generating a seed graph, nodes of the seed graph representing the one or more seed labels; based on co-occurrence of labels in member profiles of an on-line social networking system and based on the one or more seed labels, deriving one or more additional labels; generating an expanded graph comprising a first set of nodes representing the one or more additional labels and a second set of nodes representing the one or more seed labels; applying a clustering algorithm to the expanded graph to generate a labels graph; and identifying nodes of the labels graph, as a set of resolved content topics associated with the electronic document.Join the waitlist — get patent alerts
Track US2015039613A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.