Method and system for automated generation of text captions from medical images
Abstract
Computer implemented method for generating captions for medical images and/or clinical reports are provided. The methods comprise obtaining one or more medical images; using an image processing component to process the one or more images, wherein the image processing component comprises a deep learning model that takes as input the one or more medical images and produces as an output an image feature tensor; and using a natural language processing component to generate a caption for the one or more medical images, wherein the natural language processing component comprises a transformer-based model that takes as input the image feature tensor from the image processing component and produces as output a probability for each word in a vocabulary. Related systems and products are also described.
Claims
exact text as granted — not AI-modified1 . A computer implemented method for generating captions for medical images, the method comprising:
obtaining one or more medical images; using an image processing component to process the one or more medical images, wherein the image processing component comprises a deep learning model that takes as input the one or more medical images and produces as an output an image feature tensor; using a natural language processing component to generate a caption for the one or more medical images, wherein the natural language processing component comprises a transformer-based model that takes as input the image feature tensor from the image processing component and produces as output a probability for each word in a vocabulary.
2 . The method of claim 1 , wherein using a the natural language processing component to generate the caption for the one or more medical images comprises a first step which comprises using the transformer-based model to predict a probability for each word in the vocabulary and a second step which comprises sampling one or more words using the probabilities from the first step.
3 . (canceled)
4 . The method of claim 1 , wherein the transformer-based model further takes as input an input tensor derived from a set of one or more words.
5 - 6 . (canceled)
7 . The method of claim 4 , further comprising obtaining the input tensor by tokenising and embedding the set of one or more words.
8 - 10 . (canceled)
11 . The method of claim 1 , wherein the transformer-based model is obtained by training a pre-trained GPT-2 model, a pre-trained BERT model or a pre-trained T5 model.
12 . (canceled)
13 . The method of claim 1 , wherein the image processing component and the natural language processing component have been trained jointly to minimise at least one of the cross entropy loss and the perplexity of the predictions of the transformer-based model over a set of data.
14 . The method of claim 1 , further comprising receiving training data from a user and at least partially re-training the deep learning models in the image processing component and the transformer-based model in the natural language processing component using the training data.
15 . The method of any claim 1 , wherein the one or more medical images comprise multiple medical images and the method comprises generating a caption for the multiple medical images jointly, and wherein the multiple medical images are related to each other by sharing one or more features selected from: being associated with the same subject, being acquired using the same modality, showing the same pathology, showing the same organ or body part.
16 . The method of claim 1 , wherein the image processing component and the natural language processing component have been trained using training data comprising images that share one or more features with the one or more medical images, the one or more features being selected from: being associated with the same subject, being acquired using the same modality, showing the same pathology, showing the same organ or body part.
17 . The method of claim 1 , further comprising pre-processing the one or more medical images, by performing one or more steps selected from: randomly re-ordering the one or more medical images, normalising pixel values across the one or more medical images, changing the aspect ratio of one or more of the one or more medical images, scaling one or more of the one or more medical images, re-sizing one or more of the one or more medical images.
18 . The method of claim 1 , wherein the caption comprises free text.
19 . The method of claim 1 , wherein the one or more medical images are associated with a patient and the caption is a clinical report for the patient.
20 . The method of claim 1 , wherein the one or more medical images are selected from: histopathology images, radiography images, magnetic resonance images, ultrasound images, endoscopy images, positron emission tomography (PET) images, single-photon emission computed tomography (SPECT) images, and gross pathology images.
21 . The method of claim 1 , wherein the natural language processing component comprises a transformer-based model with a single stack architecture.
22 . The method of claim 1 , wherein the transformer-based model uses an attention mask that, is configured to forbid elements in an input tensor derived from a set of one or more words from attending to one another.
23 . The method of claim 1 , wherein the transformer-based model comprises one or more encoder and decoder blocks, each comprising a multi-head attention layer.
24 . The method of claim 1 , wherein the transformer-based model further takes as input a vector comprising information about a relative position of elements in an input tensor derived from a set of one or more words, wherein the relative position of the elements corresponds to an order of the one or more words in the set of one or more words from which the input tensor was derived.
25 . (canceled)
26 . The method of claim 7 , wherein the input tensor has a size K×M, wherein M is the size of the embedding used by the transformer-based model and K is a number of tokens derived from the set of one or more words by tokenisation.
27 . The method of claim 4 , wherein the transformer-based model takes as input a tensor that comprises the image feature tensor pre-pended to the input tensor.
28 - 62 . (canceled)
63 . A system comprising:
at least one processor; and
at least one non-transitory computer readable medium containing instructions that, when executed by the at least one processor, cause the at least one processor to:
obtain one or more medical images; use an image processing component to process the one or more medical images, wherein the image processing component comprises a deep learning model that takes as input the one or more medical images and produces as an output an image feature tensor; use a natural language processing component to generate a caption for the one or more medical images, wherein the natural language processing component comprises a transformer-based model that takes as input the image feature tensor from the image processing component and produces as output a probability for each word in a vocabulary.
64 - 68 . (canceled)Join the waitlist — get patent alerts
Track US2023274420A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.