US2022004878A1PendingUtilityA1

Systems and methods for synthetic document and data generation

Assignee: CAPITAL ONE SERVICES LLCPriority: Oct 17, 2018Filed: Sep 21, 2021Published: Jan 6, 2022
Est. expiryOct 17, 2038(~12.2 yrs left)· nominal 20-yr term from priority
G06N 3/044G06N 3/045G06N 3/09G06N 3/096G06N 3/0464G06F 40/205G06F 16/355G06F 16/35G06N 3/084
67
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present disclosure relates to systems and methods for determining synthetic information for documents. In one implementation, a system for determining synthetic information for documents may include at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising: receiving a plurality of documents; determining a distribution of values for pixels of the documents; identifying at least one input field based on the determined distribution; extracting information from the at least one input field; calculate at least one statistic associated with the extracted information; generating a template having the at least one input field; and inserting data based on the calculated at least one statistic into the at least one input field of the template to generate a synthetic document.

Claims

exact text as granted — not AI-modified
1 - 20 . (canceled) 
     
     
         21 . A system for determining synthetic information for documents, comprising:
 at least one processor; and   at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:
 receiving a plurality of documents; 
 determining a distribution of values for pixels of the documents; 
 identifying at least one input field based on the determined distribution; 
 extracting information from the at least one input field; 
 calculate at least one statistic associated with the extracted information; 
 generating a template having the at least one input field; and 
 inserting data based on the calculated at least one statistic into the at least one input field of the template to generate a synthetic document. 
   
     
     
         22 . The system of  claim 21 , wherein determining the distribution of values for pixels of the documents comprises generating the distribution. 
     
     
         23 . The system of  claim 21 , wherein the documents are of a same document type. 
     
     
         24 . The system of  claim 21 , the operations further comprising determining a statistical value of the distribution of values for pixels of the documents, wherein the at least one input field is identified based on the statistical value. 
     
     
         25 . The system of  claim 24 , wherein the statistical value is associated with pixel values of pixels having a same position across multiple documents of the plurality. 
     
     
         26 . The system of  claim 24 , wherein the statistical value is a standard deviation. 
     
     
         27 . The system of  claim 26 , wherein identifying the at least one input field is based on the standard deviation being greater than or equal to a threshold value. 
     
     
         28 . The system of  claim 26 , the operations further comprising identifying, based on the standard deviation being below a threshold value, a common feature of multiple documents of the plurality. 
     
     
         29 . The system of  claim 21 , the operations further comprising identifying a common feature of documents of the plurality, wherein the template is generated to include the common feature. 
     
     
         30 . The system of  claim 21 , the operations further comprising extracting data from a document using the generated template. 
     
     
         31 . The system of  claim 30 , wherein extracting data from the document using the generated template comprises performing a visual characterization process to input fields of the document. 
     
     
         32 . The system of  claim 31 , wherein the visual characterization process comprises optical character recognition. 
     
     
         33 . The system of  claim 21 , wherein at least a portion of the inserted data is retrieved from a database prior to insertion. 
     
     
         34 . The system of  claim 33 , wherein the database associates synthetic documents with respective metadata. 
     
     
         35 . The system of  claim 21 , wherein the inserted data comprises synthetic data and actual data. 
     
     
         36 . The system of  claim 21 , wherein the inserted data comprises only synthetic data. 
     
     
         37 . The system of  claim 21 , the operations further comprising:
 extracting metadata from the documents; and   storing the metadata in a database.   
     
     
         38 . The system of  claim 21 , the operations further comprising training, using the generated synthetic document, a program to:
 identify handwritten information;   identify typed information;   identify an expected data type; or   identify a document type.   
     
     
         39 . A method for determining synthetic information for documents, comprising:
 receiving a plurality of documents;   determining a distribution of values for pixels of the documents;   identifying at least one input field based on the determined distribution;   extracting information from the at least one input field;   calculating at least one statistic associated with the extracted information;   generating a template having the at least one input field; and   inserting data based on the calculated at least one statistic into the at least one input field of the template to generate a synthetic document.   
     
     
         40 . A system for determining synthetic information for documents, comprising:
 at least one processor; and   at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:
 receiving a plurality of documents associated with respective individuals; 
 identifying input fields of the plurality of documents; 
   extracting data from the identified input fields;
 clustering the extracted data into clusters based on characteristics of the individuals; 
 determine expected values for the identified input fields based on the clustered extracted data; and 
 performing at least one of:
 generating synthetic data based on the expected values and training a model learning model using the generated synthetic data; or 
 generating a document template based on the identified input fields and populating the generated document template using the expected values.

Join the waitlist — get patent alerts

Track US2022004878A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.