US2022293256A1PendingUtilityA1

Machine learning analysis of databases

Assignee: PALANTIR TECHNOLOGIES INCPriority: Dec 22, 2016Filed: May 26, 2022Published: Sep 15, 2022
Est. expiryDec 22, 2036(~10.4 yrs left)· nominal 20-yr term from priority
Inventors:Logan Kendall
G06Q 10/10G06Q 10/105G16H 40/20G06Q 10/00G16H 70/20G06Q 40/08G16H 10/60
50
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods include one or more processors, and a memory storing instructions that, when executed by the one or more processors, in conjunction with a particular machine learning model for a subset of the instructions, cause the system to perform automatically obtaining data of entities from databases based on a frequency at which the data changes, storing the obtained data in a repository, and using the particular machine learning model, performing analysis within the databases.

Claims

exact text as granted — not AI-modified
1 . A system comprising:
 one or more processors; and   a memory storing instructions that, when executed by the one or more processors, in conjunction with a particular machine learning model for a subset of the instructions, cause the system to perform:
 obtaining data of entities from databases based on a frequency at which the data changes; 
 storing the obtained data in a repository; 
 using the particular machine learning model, detecting misuse among entities, wherein training of the particular machine learning model comprises:
 obtaining a first training dataset from among known outcomes of previous analyses based on first sources verified to have been associated with misuse; and 
 obtaining a second training dataset from among known outcomes of previous analyses based on second sources verified to have been nonassociated with misuse; and 
 
 outputting an indication of the detected misuse. 
   
     
     
         2 . The system of  claim 1 , wherein the obtaining of the first training dataset and the second training dataset is further based on a rate of convergence of the particular machine learning model resulting from training using the first sources and the second sources. 
     
     
         3 . The system of  claim 1 , wherein the first sources are associated with highest rates of convergence of the particular machine learning model compared to other sources verified to have been associated with misuse, and the second sources are associated with highest rates of convergence of the particular machine learning model compared to other sources verified to have been nonassociated with misuse. 
     
     
         4 . The system of  claim 1 , wherein the first sources are associated with highest uncertainties compared to other sources verified to have been associated with misuse, and the second sources are associated with highest uncertainties compared to other sources verified to have been nonassociated with misuse. 
     
     
         5 . The system of  claim 1 , wherein the instructions further cause the system to perform translating indicators of misuse within the first training dataset and the second training dataset into particular metrics, metric values, or weights, wherein the particular metrics, metric values, or weights are used to iteratively train the particular machine learning model. 
     
     
         6 . The system of  claim 5 , wherein the iterative training comprises modifying weights assigned to signals of the particular machine learning model. 
     
     
         7 . The system of  claim 1 , wherein the particular machine learning model comprises a nearest neighbor model. 
     
     
         8 . The system of  claim 1 , wherein the first training dataset and the second training dataset are obtained from a different model. 
     
     
         9 . The system of  claim 1 , wherein the instructions further cause the system to perform obtaining a third training dataset from among previous analyses by selecting previous third sources that were indeterminate regarding an association with misuse. 
     
     
         10 . The system of  claim 1 , wherein the instructions further cause the system to perform appending, to an interface, a natural language explanation of the detected misuse and a correlation between the detected misuse and a previous instance of misuse. 
     
     
         11 . A method implemented by a computing system including one or more processors and storage media storing machine-readable instructions, wherein the method is performed using the one or more processors, in conjunction with a particular machine learning model, the method comprising:
 obtaining data of entities from databases based on a frequency at which the data changes;   storing the obtained data in a repository;   using the particular machine learning model, detecting misuse among entities, wherein training of the particular machine learning model comprises:
 obtaining a first training dataset from among known outcomes of previous analyses based on first sources verified to have been associated with misuse; and 
 obtaining a second training dataset from among known outcomes of previous analyses based on second sources verified to have been nonassociated with misuse; and 
   outputting an indication of the detected misuse.   
     
     
         12 . The method of  claim 11 , wherein the obtaining of the first training dataset and the second training dataset is further based on a rate of convergence of the particular machine learning model resulting from training using the first sources and the second sources. 
     
     
         13 . The method of  claim 11 , wherein the first sources are associated with highest rates of convergence of the particular machine learning model compared to other sources verified to have been associated with misuse, and the second sources are associated with highest rates of convergence of the particular machine learning model compared to other sources verified to have been nonassociated with misuse. 
     
     
         14 . The method of  claim 11 , wherein the first sources are associated with highest uncertainties compared to other sources verified to have been associated with misuse, and the second sources are associated with highest uncertainties compared to other sources verified to have been nonassociated with misuse. 
     
     
         15 . The method of  claim 11 , further comprising translating indicators of misuse within the first training dataset and the second training dataset into particular metrics, metric values, or weights, wherein the particular metrics, metric values, or weights are used to iteratively train the particular machine learning model. 
     
     
         16 . The method of  claim 15 , wherein the iterative training comprises modifying weights assigned to signals of the particular machine learning model. 
     
     
         17 . The method of  claim 11 , wherein the particular machine learning model comprises a nearest neighbor model. 
     
     
         18 . The method of  claim 11 , wherein the first training dataset and the second training dataset are obtained from a different model. 
     
     
         19 . The method of  claim 11 , further comprising obtaining a third training dataset from among previous analyses by selecting previous third sources that were indeterminate regarding an association with misuse. 
     
     
         20 . The method of  claim 11 , further comprising appending, to an interface, a natural language explanation of the detected misuse and a correlation between the detected misuse and a previous instance of misuse.

Join the waitlist — get patent alerts

Track US2022293256A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.