Systems and methods for synthetic document and data generation
Abstract
The present disclosure relates to systems and methods for determining synthetic information for documents. In one implementation, a system for determining synthetic information for documents may include at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising: receiving a plurality of documents; determining a distribution of values for pixels of the documents; identifying at least one input field based on the determined distribution; extracting information from the at least one input field; calculate at least one statistic associated with the extracted information; generating a template having the at least one input field; and inserting data based on the calculated at least one statistic into the at least one input field of the template to generate a synthetic document.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A system for determining synthetic information for documents, comprising:
at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:
receiving a plurality of documents;
determining a distribution of values for pixels of the documents;
identifying at least one input field based on the determined distribution;
extracting information from the at least one input field;
calculate at least one statistic associated with the extracted information;
generating a template having the at least one input field; and
inserting data based on the calculated at least one statistic into the at least one input field of the template to generate a synthetic document.
22 . The system of claim 21 , wherein determining the distribution of values for pixels of the documents comprises generating the distribution.
23 . The system of claim 21 , wherein the documents are of a same document type.
24 . The system of claim 21 , the operations further comprising determining a statistical value of the distribution of values for pixels of the documents, wherein the at least one input field is identified based on the statistical value.
25 . The system of claim 24 , wherein the statistical value is associated with pixel values of pixels having a same position across multiple documents of the plurality.
26 . The system of claim 24 , wherein the statistical value is a standard deviation.
27 . The system of claim 26 , wherein identifying the at least one input field is based on the standard deviation being greater than or equal to a threshold value.
28 . The system of claim 26 , the operations further comprising identifying, based on the standard deviation being below a threshold value, a common feature of multiple documents of the plurality.
29 . The system of claim 21 , the operations further comprising identifying a common feature of documents of the plurality, wherein the template is generated to include the common feature.
30 . The system of claim 21 , the operations further comprising extracting data from a document using the generated template.
31 . The system of claim 30 , wherein extracting data from the document using the generated template comprises performing a visual characterization process to input fields of the document.
32 . The system of claim 31 , wherein the visual characterization process comprises optical character recognition.
33 . The system of claim 21 , wherein at least a portion of the inserted data is retrieved from a database prior to insertion.
34 . The system of claim 33 , wherein the database associates synthetic documents with respective metadata.
35 . The system of claim 21 , wherein the inserted data comprises synthetic data and actual data.
36 . The system of claim 21 , wherein the inserted data comprises only synthetic data.
37 . The system of claim 21 , the operations further comprising:
extracting metadata from the documents; and storing the metadata in a database.
38 . The system of claim 21 , the operations further comprising training, using the generated synthetic document, a program to:
identify handwritten information; identify typed information; identify an expected data type; or identify a document type.
39 . A method for determining synthetic information for documents, comprising:
receiving a plurality of documents; determining a distribution of values for pixels of the documents; identifying at least one input field based on the determined distribution; extracting information from the at least one input field; calculating at least one statistic associated with the extracted information; generating a template having the at least one input field; and inserting data based on the calculated at least one statistic into the at least one input field of the template to generate a synthetic document.
40 . A system for determining synthetic information for documents, comprising:
at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:
receiving a plurality of documents associated with respective individuals;
identifying input fields of the plurality of documents;
extracting data from the identified input fields;
clustering the extracted data into clusters based on characteristics of the individuals;
determine expected values for the identified input fields based on the clustered extracted data; and
performing at least one of:
generating synthetic data based on the expected values and training a model learning model using the generated synthetic data; or
generating a document template based on the identified input fields and populating the generated document template using the expected values.Join the waitlist — get patent alerts
Track US2022004878A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.