Methods for forecasting clinical course of diffuse large b-cell lymphoma using rna-based biomarkers and machine learning algorithms
Abstract
A novel classification strategy is described for forecasting clinical outcomes of Diffuse Large B-cell Lymphoma using targeted RNA sequencing combined with machine learning algorithms. The novel method classifies subjects with DLBCL into subgroups based on the clinical course of their disease and expected survival, rather than on Cell of Origin. To focus on survival, the methods first deploy machine learning and divide the subjects into subgroups based on their overall survival. A modified Bayesian classifier is then used to select genes that can forecast various survival groups, followed by validation of these biomarkers using an independent set of clinical cases. This novel approach for stratifying subjects with DLBCL based on the clinical outcome of rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone (R-CHOP) chemotherapy can be used to select high responders and low responders to R-CHOP. Low responders may be offered additional or alternative therapies to improve their survival.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for treating a subject with a heterogeneous disease, wherein the heterogeneous disease is defined as a group of biologically diverse conditions affecting same cells or tissues and causing same or similar symptoms, the method comprising:
a. providing a mathematical algorithm for forecasting clinical course of the subject with the heterogeneous disease by classifying the subject into one of several predetermined survival groups based on response to a known therapy, wherein the mathematical algorithm is trained using machine learning by analyzing a plurality of RNA-based biomarkers from a training set of subjects with the same heterogenous disease treated by the known therapy, each subject is characterized by their respective known plurality of individual RNA-based biomarkers and known survival time, and wherein the mathematical algorithm is further trained to divide all subjects from the training set of subjects into predetermined survival groups based on survival time, the mathematical algorithm is further trained to define a subset of RNA-based biomarkers corresponding thereto; b. obtaining the subset of individual RNA-based biomarkers defined in step (a) for the subject; c. forecasting clinical course for the subject using the subset of individual RNA-based biomarkers obtained from the subject; and d. treating the subject forecasted in step (c) with the known therapy.
2 . The method as in claim 1 , wherein in step (a) the mathematical algorithm is further trained to divide all training set subjects into a first group of high responders to the known therapy, and a second group of low responders to the known therapy, wherein the first group of high responders is characterized by survival time longer than average survival time for the entire training set of subjects, the second group of low responders is characterized by survival time shorter than average survival time for the entire training set of subjects.
3 . The method as in claim 2 , wherein the mathematical algorithm is further trained to define a first subset of RNA-based biomarkers corresponding to dividing all training set subjects into the first group of high responders and the second group of low responders.
4 . The method as in claim 3 , wherein a presence of a TP53 mutation is a predictor for a second group of low responders.
5 . The method as in claim 2 , wherein the mathematical algorithm is further trained to subdivide the first group of high responders into a third group of high responders and a fourth group of high responders, wherein the third group of high responders is characterized by survival time longer than average survival time for the entire first group of high responders, the fourth group of high responders is characterized by survival time shorter than average survival time for the entire first group of high responders.
6 . The method as in claim 5 , wherein the mathematical algorithm is further trained to define a second subset of RNA-based biomarkers corresponding to dividing all subjects of the first group of high responders into the third group of high responders and the fourth group of high responders.
7 . The method as in claim 6 , wherein the second subset of RNA-based biomarkers is different from the first subset of RNA-based biomarkers.
8 . The method as in claim 7 , wherein the mathematical algorithm is further trained to subdivide the second group of low responders into a fifth group of low responders and a sixth group of low responders, wherein the fifth group of low responders is characterized by survival time longer than average survival time for the entire second group of low responders, the sixth group of low responders is characterized by survival time shorter than average survival time for the entire second group of low responders.
9 . The method as in claim 8 , wherein the mathematical algorithm is further trained to define a third subset of RNA-based biomarkers corresponding to dividing all subjects of the second group of low responders into the fifth group of low responders and the sixth group of low responders.
10 . The method as in claim 9 , wherein the third subset of RNA-based biomarkers is different from the first subset of RNA-based biomarkers.
11 . The method as in claim 2 , wherein treating the subject in step (d) comprises:
a step of treating the subject forecasted in step (c) as a high responder with the known therapy; a step of treating the subject forecasted in step (c) as a low responder with a further therapy or an additional therapy; or a combination thereof.
12 . The method as in claim 1 , wherein the mathematical algorithm is based on a naïve Bayesian classifier that is a generalized naïve Bayesian classifier defined by applying a geometric mean to a likelihood product.
13 . The method as in claim 12 , wherein the naïve Bayesian classifier is trained to rank individual RNA-based biomarkers from initial set of available RNA-based biomarkers that includes at least 500 individual genes.
14 . The method as in claim 13 , wherein at least some of the individual RNA-based biomarkers are cross-validated by subdividing the training set of subjects into a plurality of subsets, constructing a naïve Bayesian classifier for the individual RNA-based biomarker for one of the subsets and verifying the same RNA-based biomarker for at least some of the remaining subsets thereby reducing noise and overfitting.
15 . The method as in claim 14 , wherein:
after cross-validation the number of ranked RNA-based biomarkers is between 50 and 70 for each of the subdividing step of the first group and the second group, the third group and the fourth group, and the fifth group and the sixth group of the training set of subjects; the set of individual RNA-based biomarkers for dividing the entire training set of subjects into the first group and the second group is different from the respective set of individual RNA-based biomarkers for subdividing the first group of high responders into the third group and the fourth group; and the set of individual RNA-based biomarkers for dividing the entire training set of subjects into the first group and the second group is different from the respective set of individual RNA-based biomarkers for subdividing the second group of low responders into the fifth group and the sixth group.
16 . The method as in claim 15 , wherein:
the set of RNA-based biomarkers for dividing the training set into the first group and the second group is selected from a group consisting of PPP2R1B, GOLGA5, LINGO2, HMGA1, SIN3A, ARID1A, BCL7A, CDK5RAP2, MAGED1, CREB3L1, AMER1, DLL1, GSTT1, GPR34, DNM2, CCNB1IP1, MUTYH, RET, CDH1, POFUT1, XRCC6, KIT, RALGDS, SS18, CD22, BRCA2, HDAC3, LHX4, FAM19A2, PRG2, PRCC, TBL1XR1, HIF1A, EDIL3, ROS1, DKK4, CDC25A, WNT7B, MYBL1, MLLT10, SLCO1B3, TACC2, CANT1, NCAM1, FGF3, FGF19, PPP3R2, CRADD, ETV6, SPP1, SDHB, FGF2, SUZ12, MB21D2, MYC, BAX, CEP57, ITGA5, ABCC3, and HECW1; the set of RNA-based biomarkers for dividing the first group of the training set into the third group of high responders and the fourth group of high responders is selected from a group consisting of DUSP22, CTNNA1, DUX2, SSX1, SSX2, CTNNB1, DCLK2, FH, DUSP9, FCGR2B, STAT5B, ESR1, CD274, TERF1, AKAP9, DGKI, HMGA1, ARNT, MAFB, PPP3CC, COL3A1, NUTM2A, CIT, MGMT, CDK6, SORT1, RCSD1, CDK5RAP2, SIN3A, RABEP1, MB21D2, KDR, SS18L1, SSBP2, SH2D5, ASXL1, AMER1, AFF1, PRKCD, 2-Sep, TPM4, FIGF, NODAL, GRM3, STAT6, GAB1, RPL22, BDNF, SNX29, MELK, ARRDC4, FGF10, MMP9, YY1AP1, HAS2, DLEC1, DEK, TLL2, BCL2L2, and ID3; the set of RNA-based biomarkers for dividing the second group of low responders of the training set into the fifth group and the sixth group is selected from a group consisting of AHI1, EPHA5, DUSP22, DUSP26, DUSP9, DUX2, MGMT, MIB1, MIPOL1, MIR1260B, MIR4321, MIR4683, MIR4758, MIR6515, MIR6752, MIR6765, BIVM-ERCC5, SSX1, SSX2, LTBP1, MAFB, TLR4, CTNNB1, ETV5, CHEK2, FUS, SS18L1, SSBP2, DGKI, CIT, TFE3, FGF19, TRIM33, CTCF, LAMA1, TBL1XR1, TOP1, RB1, OLR1, DOCK1, ARID1A, RABEP1, EP400, STK11, ETS1, MAPK1, CDC14A, LMO7, SS18, ICK, FLI1, POU5F1, RCSD1, HRAS, BACH2, CDK7, GAS5, CARS, SRSF2, and MAP3K6; or combinations thereof.
17 . A method for identifying one or more individual RNA-based biomarkers for forecasting clinical course of a subject with a heterogeneous disease, wherein the heterogeneous disease is defined as a group of biologically diverse conditions affecting same cells or tissues and causing same or similar symptoms, the method comprising the following steps:
a. providing a training set of subjects with the heterogenous disease with known plurality of individual RNA-based biomarkers and known survival time; b. based on survival time, dividing all subjects from the training set into a first group of high responders and a second group of low responders, and c. using machine learning, identifying a first subset of one or more individual RNA-based biomarkers from a plurality of individual RNA-based biomarkers, wherein the first subset of one or more individual RNA-based biomarkers is identified as correlating to dividing the subjects into the first group and the second group.
18 . The method as in claim 17 further comprising a step (d) of dividing the first group of high responders into a third group of high responders and a fourth group of high responders, wherein the third group of high responders is characterized by survival time longer than average survival time for the entire first group of high responders, the fourth group of high responders is characterized by survival time shorter than average survival time for the entire first group of high responders.
19 . A method for treating a subject with diffuse large B-cell lymphoma, comprising a step of using a Bayesian classifier to define the subject as a high responder or a low responder to chemotherapy using one or more of individual RNA-based biomarkers selected from a group consisting of PPP2R1B, GOLGA5, LINGO2, HMGA1, SIN3A, ARID1A, BCL7A, CDK5RAP2, MAGED1, CREB3L1, AMER1, DLL1, GSTT1, GPR34, DNM2, CCNB1IP1, MUTYH, RET, CDH1, POFUT1, XRCC6, KIT, RALGDS, SS18, CD22, BRCA2, HDAC3, LHX4, FAM19A2, PRG2, PRCC, TBL1XR1, HIF1A, EDIL3, ROS1, DKK4, CDC25A, WNT7B, MYBL1, MLLT10, SLCO1B3, TACC2, CANT1, NCAM1, FGF3, FGF19, PPP3R2, CRADD, ETV6, SPP1, SDHB, FGF2, SUZ12, MB21D2, MYC, BAX, CEP57, ITGA5, ABCC3, and HECW1.
20 . The method as in claim 76 , wherein treating the subject in step (d) comprises:
a step of treating the subject forecasted as a high responder with the known therapy; a step of treating the subject forecasted as a low responder with a further therapy or an additional therapy; or a combination thereof.Join the waitlist — get patent alerts
Track US2022415448A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.