US11620451B2ActiveUtilityA1

Iterative training for text-image-layout transformer

Assignee: APPLICA SP Z O OPriority: Feb 17, 2021Filed: Jun 16, 2022Granted: Apr 4, 2023
Est. expiryFeb 17, 2041(~14.6 yrs left)· nominal 20-yr term from priority
G06N 3/0895G06N 3/0464G06N 3/09G06N 3/0455G06F 40/30G06F 40/295G06F 40/106G06V 10/82G06N 3/08G06V 10/454G06V 30/40G06T 11/60G06F 18/24133
92
PatentIndex Score
7
Cited by
92
References
28
Claims

Abstract

Disclosed herein is a system and method for Natural Language Processing (NLP) of real world documents. the system and method combines various models not previously combined and overcomes the challenges of this combination. Models include an encoder-decoder model, a spatial model, and a multi-modal model. An iterative training process receives documents and generates outputs, wherein the iterative training process comprises enabling information retrieval from documents without training data.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
       1. A system for Natural Language Processing (NLP) of real world documents, the system comprising:
 one or more processors on which an NLP process is executed, the process comprising executable instructions that when executed by the one or more processors, perform a method, the method comprising,
 receiving data comprising at least text data, layout data, and image data; and 
 operating on the received data to generate a useful output that relates to analysis of the received data; and 
 an iterative training process that receives documents and generates outputs, wherein the iterative training process comprises enabling information retrieval from documents without training data. 
 
 
     
     
       2. The system of  claim 1 , wherein the method further comprises executing one or more models selected from a group comprising:
 an encoder-decoder model; 
 a spatial model, and; 
 a multi-modal model. 
 
     
     
       3. The system of  claim 1 , wherein the method further comprises receiving one or more questions regarding the received data. 
     
     
       4. The system of  claim 3 , wherein the useful output comprises answers to the one or more questions. 
     
     
       5. The system of  claim 3 , wherein the useful output comprises key information. 
     
     
       6. The system of  claim 3 , wherein the useful output comprises document classification. 
     
     
       7. The system of  claim 2 , wherein the spatial model comprises a spatial-aware transformer that employs self-attention and a word-centric masking method that concerns both images and text. 
     
     
       8. The system of  claim 2 , wherein the multi-modal model comprises adding visual features to word embeddings that are contextualized on multiple resolution levels of an image. 
     
     
       9. The system of  claim 8 , wherein the multi-modal model further comprises reliance on relative attention biases. 
     
     
       10. The system of  claim 1 , wherein the method further comprises extending a T5 architectural approach to spatial dimensions. 
     
     
       11. The system of  claim 1 , wherein the method further comprises generating contextualized image embeddings. 
     
     
       12. The system of  claim 1 , wherein the method further comprises spatial bias augmentation. 
     
     
       13. A method for Natural Language Processing (NLP) of real world documents, the system comprising:
 one or more processors on which an NLP process is executed, the process comprising executable instructions that when executed by the one or more processors, perform a method, the method comprising,
 receiving data comprising at least text data, layout data, and image data; and 
 operating on the received data to generate a useful output that relates to analysis of the received data; and 
 an iterative training process that receives documents and generates outputs, wherein the iterative training process comprises enabling information retrieval from documents without training data. 
 
 
     
     
       14. The method of  claim 13 , wherein the method further comprises executing one or more models selected from a group comprising:
 an encoder-decoder model; 
 a spatial model, and; 
 a multi-modal model. 
 
     
     
       15. The method of  claim 13 , wherein the method further comprises receiving one or more questions regarding the received data. 
     
     
       16. The method of  claim 15 , wherein the useful output comprises answers to the one or more questions. 
     
     
       17. The method of  claim 15 , wherein the useful output comprises key information. 
     
     
       18. The method of  claim 15 , wherein the useful output comprises document classification. 
     
     
       19. The method of  claim 14 , wherein the spatial model comprises a spatial-aware transformer that employs self-attention and a word-centric masking method that concerns both images and text. 
     
     
       20. The method of  claim 14  wherein the multi-modal model comprises adding visual features to word embeddings that are contextualized on multiple resolution levels of an image. 
     
     
       21. The method of  claim 20 , wherein the multi-modal model further comprises reliance on relative attention biases. 
     
     
       22. The method of  claim 13 , wherein the method further comprises extending a T5 architectural approach to spatial dimensions. 
     
     
       23. The method of  claim 13 , wherein the method further comprises generating contextualized image embeddings. 
     
     
       24. The method of  claim 13 , wherein the method further comprises spatial bias augmentation. 
     
     
       25. A non-transient computer medium having stored thereon instructions, which when executed by a processor perform a method, the method comprising:
 receiving data comprising at least text data, layout data, and image data; and 
 operating on the received data to generate a useful output that relates to analysis of the received data; and 
 an iterative training process that receives documents and generates outputs, wherein the iterative training process comprises enabling information retrieval from documents without training data. 
 
     
     
       26. The non-transient computer medium of  claim 25 , wherein the method further comprises executing one or more models selected from a group comprising:
 an encoder-decoder model; 
 a spatial model, and; 
 a multi-modal model. 
 
     
     
       27. The non-transient computer medium of  claim 25 , wherein the method further comprises receiving one or more questions regarding the received data. 
     
     
       28. The non-transient computer medium of  claim 27 , wherein the useful output comprises answers to the one or more questions.

Join the waitlist — get patent alerts

Track US11620451B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.