US2024331815A1PendingUtilityA1

Named-entity recognition of protected health information

Assignee: PALO ALTO NETWORKS INCPriority: Mar 28, 2023Filed: Mar 28, 2023Published: Oct 3, 2024
Est. expiryMar 28, 2043(~16.7 yrs left)· nominal 20-yr term from priority
G06N 5/01G06N 20/20G06F 40/295G06N 3/045G06N 3/09G16H 10/60G06F 40/279
47
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A named-entity recognition (NER) model detects named entities with types that correspond to protected health information (PHI) in potentially sensitive documents. The NER model is trained to detect named entities corresponding to both personally identifiable information (PII) and medical terms. Output of the NER model is preprocessed as input to a random forest classifier that outputs a verdict that documents comprise sensitive data. The verdict is interpretable via high confidence named entities detected by the NER model that led to the verdict.

Claims

exact text as granted — not AI-modified
1 . A method comprising:
 determining a plurality of named entities in a document and a plurality of confidence values indicating confidence of types for each of the plurality of named entities with a named-entity recognition model, wherein the document comprises potentially sensitive data, wherein the named-entity recognition model was at least partially trained on named entities known to correspond to sensitive data;   based, at least in part, on the plurality of named entities and the plurality of confidence values, determining whether the document comprises sensitive data with a classifier; and   based on determining that the document comprises sensitive data based on output of the classifier, indicating one or more of the plurality of named entities as comprising sensitive data in the document.   
     
     
         2 . The method of  claim 1 , wherein the named entities known to correspond to sensitive data comprise entities corresponding to personally identifiable information and medical terms, wherein the potentially sensitive data in the document comprise protected health information. 
     
     
         3 . The method of  claim 1 , wherein the named-entity recognition model comprises a Bidirectional Encoder Representations from Transformers model, wherein the classifier comprises a random forest classifier. 
     
     
         4 . The method of  claim 1 , further comprising identifying high confidence named entities in the plurality of named entities that determined whether the document comprises sensitive data. 
     
     
         5 . The method of  claim 4 , wherein identifying high confidence named entities in the plurality of named entities comprises identifying, for each of a subset of types of named entities indicated as high importance by the classifier for determining that the document comprises sensitive data, a named entity in the plurality of named entities among named entities with a same type having a highest confidence value in the plurality of confidence values. 
     
     
         6 . The method of  claim 1 , wherein determining whether the document comprises sensitive data with a random forest classifier comprises,
 vectorizing the plurality of confidence values according to corresponding ones of the plurality of named entities to generate a vector of confidence values; and   inputting the vector of confidence values into the random forest classifier to output an indication of whether the document comprises sensitive data.   
     
     
         7 . The method of  claim 6 , wherein vectorizing the plurality of confidence values according to corresponding ones of the plurality of named entities comprises,
 identifying a maximal confidence value and a mean confidence value based on the plurality of confidence values for each type of entity in the plurality of named entities; and   concatenating the maximal confidence value and the mean confidence value for each type of entity to generate the vector of confidence values.   
     
     
         8 . A non-transitory machine-readable medium having program code stored thereon, the program code comprising instructions to:
 input a plurality of tokens of a document into a named-entity recognition model;   obtain, from output of the named-entity recognition model, a plurality of named entities and a plurality of confidence values indicating confidence in types of each of the plurality of named entities, wherein the document comprises potentially sensitive data, wherein the plurality of named entities correspond to a plurality of subsets of the document;   determine whether the document comprises sensitive data based on input of a plurality of feature values generated from the plurality of named entities and the plurality of confidence values into a random forest classifier; and   based on a determination that the document comprises sensitive data indicated in output of the random forest classifier, identify one or more of the plurality of named entities corresponding to subsets of the document that comprise sensitive data.   
     
     
         9 . The non-transitory machine-readable medium of  claim 8 , wherein the plurality of named entities comprise named entities corresponding to personally identifiable information and medical terms, wherein the potentially sensitive data in the document comprise protected health information. 
     
     
         10 . The non-transitory machine-readable medium of  claim 8 , wherein the named-entity recognition model comprises a Bidirectional Encoder Representations from Transformers model. 
     
     
         11 . The non-transitory machine-readable medium of  claim 8 , wherein the program code further comprises instructions to generate the plurality of feature values, wherein the instructions to generate the plurality of feature values comprise instructions to vectorize the plurality of confidence values according to corresponding ones of the plurality of named entities to generate a vector of confidence values as the plurality of feature values. 
     
     
         12 . The non-transitory machine-readable medium of  claim 8 , wherein the instructions to identify the one or more of the plurality of named entities in the document that comprise sensitive data comprise instructions to identify the one or more of the plurality of named entities as corresponding to highest confidence values in the plurality of confidence values for named entities of a same type of named entity. 
     
     
         13 . The non-transitory machine-readable medium of  claim 12 , wherein the program code to identify the one or more of the plurality of named entities further comprises instructions to identify the one or more of the plurality of named entities as named entities with types indicated as highest importance by the random forest classifier. 
     
     
         14 . An apparatus comprising:
 a processor; and   a machine-readable medium having instructions stored thereon that are executable by the processor to cause the apparatus to,   identify a plurality of named entities in a document with a plurality of confidence values indicating confidence of types of each of the plurality of named entities based on input of a plurality of tokens of the document into a named-entity recognition model;   input a plurality of feature values of the plurality of named entities and plurality of confidence values into a random forest classifier to determine whether the document comprises sensitive data; and   based on a determination that the document comprises sensitive data, identify one or more of the plurality of named entities in the document that comprise sensitive data.   
     
     
         15 . The apparatus of  claim 14 , wherein the plurality of named entities comprise entities that correspond to personally identifiable information and medical terms, wherein the determination that the document comprises sensitive data comprises a determination that the document comprises protected health information. 
     
     
         16 . The apparatus of  claim 14 , wherein the named-entity recognition model comprises a Bidirectional Encoder Representations from Transformers model. 
     
     
         17 . The apparatus of  claim 14 , wherein the machine-readable medium further has stored thereon instructions executable by the processor to cause the apparatus to generate the plurality of feature values, wherein the instructions executable by the processor to cause the apparatus to generate the plurality of feature values comprise instructions to vectorize the plurality of confidence values according to corresponding ones of the plurality of named entities to generate a vector of confidence values as the plurality of feature values. 
     
     
         18 . The apparatus of  claim 14 , wherein the instructions executable by the processor to cause the apparatus to identify the one or more of the plurality of named entities in the document that comprise sensitive data comprise instructions to identify the one or more of the plurality of named entities as corresponding to highest confidence values in the plurality of confidence values for those of the plurality of named entities with a same type. 
     
     
         19 . The apparatus of  claim 14 , wherein the machine-readable medium further has stored thereon instructions executable by the processor to cause the apparatus to indicate decisions at a plurality of nodes in the random forest classifier that determined whether the document comprises sensitive data. 
     
     
         20 . The apparatus of  claim 14 , wherein the instructions executable by the processor to cause the apparatus to identify the plurality of named entities and the plurality of confidence values of the named entities in the document further comprise instructions to:
 subdivide the document into a plurality of sub-documents; and   tokenize the plurality of sub-documents to extract the plurality of tokens.

Join the waitlist — get patent alerts

Track US2024331815A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.