US11620451B2ActiveUtilityA1
Iterative training for text-image-layout transformer
Est. expiryFeb 17, 2041(~14.6 yrs left)· nominal 20-yr term from priority
G06N 3/0895G06N 3/0464G06N 3/09G06N 3/0455G06F 40/30G06F 40/295G06F 40/106G06V 10/82G06N 3/08G06V 10/454G06V 30/40G06T 11/60G06F 18/24133
92
PatentIndex Score
7
Cited by
92
References
28
Claims
Abstract
Disclosed herein is a system and method for Natural Language Processing (NLP) of real world documents. the system and method combines various models not previously combined and overcomes the challenges of this combination. Models include an encoder-decoder model, a spatial model, and a multi-modal model. An iterative training process receives documents and generates outputs, wherein the iterative training process comprises enabling information retrieval from documents without training data.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A system for Natural Language Processing (NLP) of real world documents, the system comprising:
one or more processors on which an NLP process is executed, the process comprising executable instructions that when executed by the one or more processors, perform a method, the method comprising,
receiving data comprising at least text data, layout data, and image data; and
operating on the received data to generate a useful output that relates to analysis of the received data; and
an iterative training process that receives documents and generates outputs, wherein the iterative training process comprises enabling information retrieval from documents without training data.
2. The system of claim 1 , wherein the method further comprises executing one or more models selected from a group comprising:
an encoder-decoder model;
a spatial model, and;
a multi-modal model.
3. The system of claim 1 , wherein the method further comprises receiving one or more questions regarding the received data.
4. The system of claim 3 , wherein the useful output comprises answers to the one or more questions.
5. The system of claim 3 , wherein the useful output comprises key information.
6. The system of claim 3 , wherein the useful output comprises document classification.
7. The system of claim 2 , wherein the spatial model comprises a spatial-aware transformer that employs self-attention and a word-centric masking method that concerns both images and text.
8. The system of claim 2 , wherein the multi-modal model comprises adding visual features to word embeddings that are contextualized on multiple resolution levels of an image.
9. The system of claim 8 , wherein the multi-modal model further comprises reliance on relative attention biases.
10. The system of claim 1 , wherein the method further comprises extending a T5 architectural approach to spatial dimensions.
11. The system of claim 1 , wherein the method further comprises generating contextualized image embeddings.
12. The system of claim 1 , wherein the method further comprises spatial bias augmentation.
13. A method for Natural Language Processing (NLP) of real world documents, the system comprising:
one or more processors on which an NLP process is executed, the process comprising executable instructions that when executed by the one or more processors, perform a method, the method comprising,
receiving data comprising at least text data, layout data, and image data; and
operating on the received data to generate a useful output that relates to analysis of the received data; and
an iterative training process that receives documents and generates outputs, wherein the iterative training process comprises enabling information retrieval from documents without training data.
14. The method of claim 13 , wherein the method further comprises executing one or more models selected from a group comprising:
an encoder-decoder model;
a spatial model, and;
a multi-modal model.
15. The method of claim 13 , wherein the method further comprises receiving one or more questions regarding the received data.
16. The method of claim 15 , wherein the useful output comprises answers to the one or more questions.
17. The method of claim 15 , wherein the useful output comprises key information.
18. The method of claim 15 , wherein the useful output comprises document classification.
19. The method of claim 14 , wherein the spatial model comprises a spatial-aware transformer that employs self-attention and a word-centric masking method that concerns both images and text.
20. The method of claim 14 wherein the multi-modal model comprises adding visual features to word embeddings that are contextualized on multiple resolution levels of an image.
21. The method of claim 20 , wherein the multi-modal model further comprises reliance on relative attention biases.
22. The method of claim 13 , wherein the method further comprises extending a T5 architectural approach to spatial dimensions.
23. The method of claim 13 , wherein the method further comprises generating contextualized image embeddings.
24. The method of claim 13 , wherein the method further comprises spatial bias augmentation.
25. A non-transient computer medium having stored thereon instructions, which when executed by a processor perform a method, the method comprising:
receiving data comprising at least text data, layout data, and image data; and
operating on the received data to generate a useful output that relates to analysis of the received data; and
an iterative training process that receives documents and generates outputs, wherein the iterative training process comprises enabling information retrieval from documents without training data.
26. The non-transient computer medium of claim 25 , wherein the method further comprises executing one or more models selected from a group comprising:
an encoder-decoder model;
a spatial model, and;
a multi-modal model.
27. The non-transient computer medium of claim 25 , wherein the method further comprises receiving one or more questions regarding the received data.
28. The non-transient computer medium of claim 27 , wherein the useful output comprises answers to the one or more questions.Join the waitlist — get patent alerts
Track US11620451B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.