Line of Therapy Identification from Clinical Documents
Abstract
A method includes receiving input data including unstructured text representing one or more sequences of terms. For each respective sequence of terms, the method includes generating a corresponding line of therapy (LoT) pseudo-label indicating whether the respective sequence of terms includes LoT information, generating a corresponding LoT indicator predicting whether the respective sequence of terms includes LoT information, and determining a corresponding LoT indication loss based on the corresponding LoT pseudo-label and the corresponding LoT indicator. The method also includes fine-tuning a pre-trained transformer model based on the LoT indication losses determined for the one or more sequences of terms.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations comprising:
receiving input data comprising unstructured text representing one or more sequences of terms; for each respective sequence of terms:
generating, using regular expression rules, a corresponding line of therapy (LoT) pseudo-label indicating whether the respective sequence of terms comprises LoT information;
generating, using a pre-trained transformer model, a corresponding LoT indicator predicting whether the respective sequence of terms comprises LoT information; and
determining a corresponding LoT indication loss based on the corresponding LoT pseudo-label and the corresponding LoT indicator; and
fine-tuning the pre-trained transformer model based on the LoT indication losses determined for the one or more sequences of terms.
2 . The computer-implemented method of claim 1 , wherein the operations further comprise:
for each respective sequence of terms:
generating, using the regular expression rules, a corresponding classification pseudo-label indicating a classification of the respective sequence of terms;
generating, using the pre-trained transformer model, a corresponding LoT classification predicting the classification of the respective sequence of terms; and
determining a corresponding LoT classification loss based on the corresponding classification pseudo-label and the corresponding LoT classification; and
fine-tuning the pre-trained transformer model based on the LoT classification losses determined for the one or more sequences of terms.
3 . The computer-implemented method of claim 2 , wherein the classification of the respective sequence of terms indicates a particular LoT step from a series of LoT steps associated with the respective sequence of terms.
4 . The computer-implemented method of claim 1 , wherein the input data comprises a plurality of clinical trial documents comprising the unstructured text.
5 . The computer-implemented method of claim 1 , wherein the operations further comprise updating the regular expression rules based on the corresponding LoT indication loss.
6 . The computer-implemented method of claim 1 , wherein the pre-trained transformer model comprises a Bidirectional Encoder from Transformers for Biomedical Text Mining (BioBERT) model pre-trained on a corpus of biomedical text data.
7 . The computer-implemented method of claim 6 , wherein the pre-trained BioBERT model comprises a stack of multi-headed self-attention layers.
8 . The computer-implemented method of claim 1 , wherein the operations further comprise fine-tuning the pre-trained transformer model using an aggregation of training data including human annotated labels, labels generated using the regular expression rules, and labels generated by the pre-trained transformer model.
9 . The computer-implemented method of claim 1 , wherein the operations further comprise storing the fine-tuned transformer model in memory hardware in communication with the data processing hardware.
10 . The computer-implemented method of claim 1 , wherein the operations further comprising transmitting, via a network, the fine-tuned transformer model to one or more computing devices.
11 . A system comprising:
data processing hardware; and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
receiving input data comprising unstructured text representing one or more sequences of terms;
for each respective sequence of terms:
generating, using regular expression rules, a corresponding line of therapy (LoT) pseudo-label indicating whether the respective sequence of terms comprises LoT information;
generating, using a pre-trained transformer model, a corresponding LoT indicator predicting whether the respective sequence of terms comprises LoT information; and
determining a corresponding LoT indication loss based on the corresponding LoT pseudo-label and the corresponding LoT indicator; and
fine-tuning the pre-trained transformer model based on the LoT indication losses determined for the one or more sequences of terms.
12 . The system of claim 11 , wherein the operations further comprise:
for each respective sequence of terms:
generating, using the regular expression rules, a corresponding classification pseudo-label indicating a classification of the respective sequence of terms;
generating, using the pre-trained transformer model, a corresponding LoT classification predicting the classification of the respective sequence of terms; and
determining a corresponding LoT classification loss based on the corresponding classification pseudo-label and the corresponding LoT classification; and
fine-tuning the pre-trained transformer model based on the LoT classification losses determined for the one or more sequences of terms.
13 . The system of claim 12 , wherein the classification of the respective sequence of terms indicates a particular LoT step from a series of LoT steps associated with the respective sequence of terms.
14 . The system of claim 11 , wherein the input data comprises a plurality of clinical trial documents comprising the unstructured text.
15 . The system of claim 11 , wherein the operations further comprise wherein the operations further comprise updating the regular expression rules based on the corresponding LoT indication loss.
16 . The system of claim 11 , wherein the pre-trained transformer model comprises a Bidirectional Encoder from Transformers for Biomedical Text Mining (BioBERT) model pre-trained on a corpus of biomedical text data.
17 . The system of claim 16 , wherein the pre-trained BioBERT model comprises a stack of multi-headed self-attention layers.
18 . The system of claim 11 , wherein the operations further comprise fine-tuning the pre-trained transformer model using an aggregation of training data including human annotated labels, labels generated using the regular expression rules, and labels generated by the pre-trained transformer model.
19 . The system of claim 11 , wherein the operations further comprise storing the fine-tuned transformer model in memory hardware in communication with the data processing hardware.
20 . The system of claim 11 , wherein the operations further comprising transmitting, via a network, the fine-tuned transformer model to one or more computing devices.Join the waitlist — get patent alerts
Track US2024185972A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.