US2020160231A1PendingUtilityA1

Method and System for Using a Multi-Factorial Analysis to Identify Optimal Annotators for Building a Supervised Machine Learning Model

Assignee: IBMPriority: Nov 19, 2018Filed: Nov 19, 2018Published: May 21, 2020
Est. expiryNov 19, 2038(~12.3 yrs left)· nominal 20-yr term from priority
G06N 20/00G06N 5/04G06F 16/9024G06F 17/30011G06N 99/005G06F 17/30958G06F 16/35G06N 5/022G06F 16/906
35
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method, system, apparatus, and a computer program product are provided for identifying ground truth annotators by applying statistical analyses to a document corpus and to a plurality of annotator profiles to identify, respectively, corpus complexity attributes for the document corpus and annotator qualification attributes for each candidate annotator which are compared with a matching analysis to identify one or more recommended annotators from the plurality of candidate annotators based on the matching analysis.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer implemented method in a data processing system comprising a processor and a memory, the memory comprising instructions that are executed by the processor to cause the processor to implement a system for identifying one or more annotators for annotating a document corpus to create a ground truth based on which a model can be trained, the method comprising:
 receiving, by the processor, a document corpus, wherein the document corpus comprises a plurality of documents related to a particular domain;   receiving, by the processor, annotator profiles for a plurality of candidate annotators, wherein each annotator profile comprises profile data selected from a group consisting of prior annotation history, writing style, technical domain expertise, publicly expressed area of interests, and personality insights for each candidate annotator;   applying, by the processor, a first plurality of statistical analyses to the document corpus to identify corpus complexity attributes for the document corpus:   applying, by the processor, a second plurality of statistical analyses to the annotator profiles to identify annotator qualification attributes for each candidate annotator; and   identifying, by the processor, one or more recommended annotators from the plurality of candidate annotators based on a matching analysis of the corpus complexity attributes for the document corpus with the annotator qualification attributes for each candidate annotator.   
     
     
         2 . The method as recited in  claim 1 , where receiving the document corpus comprises uploading, by the processor, the document corpus from a knowledge database. 
     
     
         3 . The method as recited in  claim 1 , where receiving the annotator profiles comprises uploading, by the processor, annotator registration information selected from the group consisting of annotator age, gender, location, languages, expertise, profession, past annotation work, IAA score, and statistics based on the domain. 
     
     
         4 . The method as recited in  claim 1 , where applying the first plurality of statistical analyses comprises:
 identifying, by the processor, a plurality of frequently occurring words from the document corpus;   extracting, by the processor, a conceptual text for each frequently occurring word from a structured information database;   performing, by the processor, a cluster analysis on each conceptual text to identify a plurality of possible entity types;   performing, by the processor, a frequency analysis on the plurality of possible entity types to select at least one entity type;   identifying, by the processor, a relation between two entities in the document corpus that are related;   identifying, by the processor, a relation type between the two entities in the document corpus that are related based on two entity types of the two entities and a plurality of words appearing within, before, between, or after each pair of entity mentions; and   generating, by the processor, the type system including at least one entity type.   
     
     
         5 . The method as recited in  claim 1 , where applying the second plurality of statistical analyses comprises:
 identifying, by the processor, a plurality of frequently occurring words from each annotator profile;   extracting, by the processor, a conceptual text for each frequently occurring word from a structured information database;   performing, by the processor, a cluster analysis on each conceptual text to identify a plurality of possible entity types;   performing, by the processor, a frequency analysis on the plurality of possible entity types to select at least one entity type;   identifying, by the processor, a relation between two entities in the annotator profiles that are related;   identifying, by the processor, at least one relation type between the two entities in the annotator profiles that are related based on the entity types of the two entities and a plurality of words appearing before, between, or after instances of the two entities; and   generating, by the processor, the type system including at least one entity type.   
     
     
         6 . The method as recited in  claim 1 , where identifying one or more recommended annotators comprises matching the annotator profiles to the document corpus by applying weighted scores to the annotator qualification attributes based on how closely the annotator qualification attributes match the corpus complexity attributes. 
     
     
         7 . The method of  claim 1 , further comprising:
 constructing, by the processor, a hierarchical knowledge graph of the particular domain for the document corpus based on high-level concepts extracted from the document corpus; and   determining, by the processor, a frequencies of terms at levels of the hierarchical knowledge graph to evaluate a complexity measure for the terms and density of occurrence of the terms in the document corpus.   
     
     
         8 . An information handling system comprising:
 one or more processors;   a memory coupled to at least one of the processors;   a set of instructions stored in the memory and executed by at least one of the processors to identifying one or more annotators for annotating a document corpus, wherein the set of instructions are executable to perform actions of:   receiving, by the system, a document corpus, wherein the document corpus comprises a plurality of documents related to a particular domain;   receiving, by the system, annotator profiles for a plurality of candidate annotators, wherein each annotator profile comprises profile data selected from a group consisting of prior annotation history, writing style, technical domain expertise, publicly expressed area of interests, and personality insights for each candidate annotator;   applying, by the system, a first plurality of statistical analyses to the document corpus to identify corpus complexity attributes for the document corpus;   applying, by the system, a second plurality of statistical analyses to the annotator profiles to identify annotator qualification attributes for each candidate annotator; and   identifying, by the system, one or more recommended annotators from the plurality of candidate annotators based on a matching analysis of the corpus complexity attributes for the document corpus with the annotator qualification attributes for each candidate annotator.   
     
     
         9 . The information handling system of  claim 8 , wherein the set of instructions are executable to receive the document corpus by uploading the document corpus from a knowledge database. 
     
     
         10 . The information handling system of  claim 8 , wherein the set of instructions are executable to receive the annotator profiles by uploading annotator registration information selected from the group consisting of annotator age, gender, location, languages, expertise, profession, past annotation work, IAA score, and statistics based on the domain. 
     
     
         11 . The information handling system of  claim 8 , wherein the set of instructions are executable to apply the first plurality of statistical analyses by:
 identifying, by the system, a plurality of frequently occurring words from the document corpus;   extracting, by the system, a conceptual text for each frequently occurring word from a structured information database;   performing, by the system, a cluster analysis on each conceptual text to identify a plurality of possible entity types;   performing, by the system, a frequency analysis on the plurality of possible entity types to select at least one entity type;   identifying, by the system, a relation between two entities in the document corpus that are related;   identifying, by the system, a relation type between the two entities in the document corpus that are related based on entity types of the two entities and a plurality of words appearing before, between, or after instances of the two entities; and   generating, by the system, the type system including at least one entity type.   
     
     
         12 . The information handling system of  claim 8 , wherein the set of instructions are executable to apply the second plurality of statistical analyses by:
 identifying, by the system, a plurality of frequently occurring words from each annotator profile;   extracting, by the system, a conceptual text for each frequently occurring word from a structured information database;   performing, by the system, a cluster analysis on each conceptual text to identify a plurality of possible entity types;   performing, by the system, a frequency analysis on the plurality of possible entity types to select at least one entity type;   identifying, by the system, a relation between two entities in the annotator profiles that are related;   identifying, by the system, at least one relation type between the two entities in the annotator profiles that are related based on the entity types of the two entities and a plurality of words appearing before, between, or after instances of the two entities; and   generating, by the system, the type system including at least one entity type.   
     
     
         13 . The information handling system of  claim 8 , wherein the set of instructions are executable to identify one or more recommended annotators by matching the annotator profiles to the document corpus by applying weighted scores to the annotator qualification attributes based on how closely the annotator qualification attributes match the corpus complexity attributes. 
     
     
         14 . The information handling system of  claim 8 , wherein the set of instructions are executable to:
 construct, by the system, a hierarchical knowledge graph of the particular domain for the document corpus based on high-level concepts extracted from the document corpus; and   determine, by the system, a frequencies of terms at levels of the hierarchical knowledge graph to evaluate a complexity measure for the terms and density of occurrence of the terms in the document corpus.   
     
     
         15 . A computer program product stored in a computer readable storage medium, comprising computer instructions that, when executed by a processor at an information handling system, causes the system to identify one or more annotators for annotating a document corpus by:
 receiving, by the processor, a document corpus, wherein the document corpus comprises a plurality of documents related to a particular domain;   receiving, by the processor, annotator profiles for a plurality of candidate annotators, wherein each annotator profile comprises profile data selected from a group consisting of prior annotation history, writing style, technical domain expertise, publicly expressed area of interests, and personality insights for each candidate annotator;   applying, by the processor, a first plurality of statistical analyses to the document corpus to identify corpus complexity attributes for the document corpus;   applying, by the processor, a second plurality of statistical analyses to the annotator profiles to identify annotator qualification attributes for each candidate annotator; and   identifying, by the processor, one or more recommended annotators from the plurality of candidate annotators based on a matching analysis of the corpus complexity attributes for the document corpus with the annotator qualification attributes for each candidate annotator.   
     
     
         16 . The computer program product of  claim 15 , further comprising computer instructions that, when executed by the system, causes the system to receive the annotator profiles by uploading annotator registration information selected from the group consisting of annotator age, gender, location, languages, expertise, profession, past annotation work, IAA score, and statistics based on the domain. 
     
     
         17 . The computer program product of  claim 15 , further comprising computer instructions that, when executed by the system, causes the system to apply the first plurality of statistical analyses by:
 identifying, by the processor, a plurality of frequently occurring words from the document corpus;   extracting, by the processor, a conceptual text for each frequently occurring word from a structured information database;   performing, by the processor, a cluster analysis on each conceptual text to identify a plurality of possible entity types;   performing, by the processor, a frequency analysis on the plurality of possible entity types to select at least one entity type;   identifying, by the processor, a relation between two entities in the document corpus that are related;   identifying, by the processor, a relation type between the two entities in the document corpus that are related based on the entity types of the two entities and a plurality of words appearing before, between, or after instances of the two entities; and   generating, by the processor, the type system including at least one entity type.   
     
     
         18 . The computer program product of  claim 15 , further comprising computer instructions that, when executed by the system, causes the system to apply the second plurality of statistical analyses by:
 identifying, by the processor, a plurality of frequently occurring words from each annotator profile;   extracting, by the processor, a conceptual text for each frequently occurring word from a structured information database;   performing, by the processor, a cluster analysis on each conceptual text to identify a plurality of possible entity types;   performing, by the processor, a frequency analysis on the plurality of possible entity types to select at least one entity type;   identifying, by the processor, a relation between two entities in the annotator profiles that are related;   identifying, by the processor, at least one relation type between the two entities in the annotator profiles that are related based on the entity types of the two entities and a plurality of words appearing before, between, or after instances of the two entities; and   generating, by the processor, the type system including at least one entity type.   
     
     
         19 . The computer program product of  claim 15 , further comprising computer instructions that, when executed by the system, causes the system to identify one or more recommended annotators by matching the annotator profiles to the document corpus by applying weighted scores to the annotator qualification attributes based on how closely the annotator qualification attributes match the corpus complexity attributes. 
     
     
         20 . The computer program product of  claim 15 , further comprising computer instructions that, when executed by the system, causes the system to:
 construct, by the processor, a hierarchical knowledge graph of the particular domain for the document corpus based on high-level concepts extracted from the document corpus; and   determine, by the processor, a frequencies of terms at levels of the hierarchical knowledge graph to evaluate a complexity measure for the terms and density of occurrence of the terms in the document corpus.

Join the waitlist — get patent alerts

Track US2020160231A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.