Proteogenomic-based method for identifying tumor-specific antigens
Abstract
T cells, notably CD8 T cells, are known to be essential players in tumor eradication as the presence of tumor-infiltrating lymphocytes (TILs) in several cancers positively correlates with a good prognosis. To eliminate tumor cells, CD8 T cells recognize tumor antigens, which are MHC I-associated peptides present at the surface of tumor cells, with no or very low expression on normal cells. Described herein a proteogenomic approach using RNA-sequencing data from cancer and normal-matched mTEChi samples in order to identify non-tolerogenic tumor-specific antigens derived from (i) coding and non-coding regions of the genome, (ii) non-synonymous single-base mutations or short insertion/deletions and more complex rearrangements as well as (iii) endogenous retroelements, which works regardless of the sample's mutational load or complexity.
Claims
exact text as granted — not AI-modified1 - 75 . (canceled)
76 . A method for identifying a tumor antigen candidate in a tumor cell sample, the method comprising:
(a) generating a tumor-specific proteome database by:
(i) extracting a set of subsequences (k-mers) comprising at least 33 base pairs from tumor RNA-sequences;
(ii) comparing the set of tumor subsequences of (i) to a set of corresponding control subsequences comprising at least 33 base pairs extracted from RNA-sequences from normal cells;
(iii) extracting the tumor subsequences that are absent in the corresponding control subsequences, thereby obtaining tumor-specific subsequences; and
(iv) in silico translating the tumor-specific subsequences, thereby obtaining the tumor-specific proteome database;
(b) generating a personalized tumor proteome database by:
(i) comparing the tumor RNA-sequences to a reference genome sequence to identify single-base mutations in said tumor RNA-sequences;
(ii) inserting the single-base mutations identified in (i) in the reference genome sequence, thereby creating a personalized tumor genome sequence;
(iii) in silico translating the expressed protein-coding transcripts from said personalized tumor genome sequence, thereby obtaining the personalized tumor proteome database;
(c) comparing the sequences of major histocompatibility complex (MHC)-associated peptides (MAPs) from said tumor with the sequences of the tumor-specific proteome database of (a) and the personalized tumor proteome database of (b) to identify the MAPs; and (d) identifying a tumor antigen candidate among the MAPs identified in (c), wherein a tumor antigen candidate is a peptide whose sequence and/or encoding sequence is overexpressed or overrepresented in tumor cells relative to normal cells.
77 . The method of claim 76 , wherein the above-noted method further comprises (1) isolating and sequencing major histocompatibility complex (MHC)-associated peptides (MAPs) from the tumor cell sample, and/or (2) performing whole transcriptome sequencing on the tumor cell sample, to obtain the tumor RNA-sequences.
78 . The method of claim 77 , wherein said isolating MAPs comprises (i) releasing said MAPs from said cell sample by mild acid treatment; and (ii) subjecting the released MAPs to chromatography.
79 . The method of claim 78 , wherein said method further comprises filtering the released peptides with a size exclusion column prior to said chromatography.
80 . The method of claim 79 , wherein said size exclusion column has a cut-off of about 3000 Da.
81 . The method of claim 76 , wherein said subsequences comprises from 33 to 54 base pairs.
82 . The method of claim 76 , further comprising assembling overlapping tumor-specific subsequences into longer tumor subsequences (contigs).
83 . The method of claim 76 , wherein said sequencing of MAPs comprises subjecting the isolated MAPs to mass spectrometry (MS) sequencing analysis.
84 . The method of claim 76 , wherein said method further comprises generating a personalized normal proteome database using corresponding normal cells, and wherein said identifying in (d) comprises excluding said MAP if its sequence is detected in the normal personalized proteome database.
85 . The method of claim 76 , wherein the method further comprises generating 24- or 39-nucleotide k-mer databases from said tumor RNA-sequences and from RNA-sequences from normal cells to obtain a tumor k-mer database and a normal k-mer database; and comparing the tumor k-mer database and a normal k-mer database to 24- or 39-nucleotide k-mer derived from the MAP encoding sequence, wherein an overexpression or overrepresentation of the k-mer derived from the MAP encoding sequence in said tumor k-mer database relative to said normal k-mer database is indicative that the corresponding MAP is a tumor antigen candidate.
86 . The method of claim 85 , wherein the k-mer derived from the MAP encoding sequence is overexpressed or overrepresented by at least 10-fold in said tumor k-mer database relative to said normal k-mer database.
87 . The method of claim 85 , wherein the k-mer derived from the MAP encoding sequence is absent from said normal k-mer database.
88 . The method of claim 76 , wherein said method comprises:
(a) isolating and sequencing MAPs in a tumor cell sample; (b) performing whole transcriptome sequencing on said tumor cell sample, thereby obtaining tumor RNA-sequences; (c) generating a tumor-specific proteome database by:
(i) extracting a set of subsequences comprising at least 33 nucleotides from said tumor RNA-sequences;
(ii) comparing the set of tumor subsequences of (i) to a set of corresponding control subsequences comprising at least 33 nucleotides extracted from RNA-sequences from normal cells;
(iii) extracting the tumor subsequences that are absent, or underexpressed by at least 4-fold, in the corresponding control subsequences, thereby obtaining tumor-specific subsequences; and
(iv) in silico translating the tumor-specific subsequences, thereby obtaining the tumor-specific proteome database;
(d) generating a personalized tumor proteome database by:
(i) comparing the tumor RNA-sequences to a reference genome sequence to identify single-base mutations in said tumor RNA-sequences;
(ii) inserting the single-base mutations identified in (i) in the reference genome sequence, thereby creating a personalized tumor genome sequence;
(iii) in silico translating the expressed protein-coding transcripts from said personalized tumor genome sequence, thereby obtaining the personalized tumor proteome database;
(e) generating a personalized normal proteome database by:
(i) comparing RNA-sequences from normal cells to a reference genome sequence to identify single-base mutations in said normal RNA-sequences;
(ii) inserting the single-base mutations identified in (i) in the reference genome sequence, thereby creating a personalized normal genome sequence;
(iii) in silico translating the expressed protein-coding transcripts from said personalized normal genome sequence, thereby obtaining the personalized normal proteome database;
(f) generating a normal and a tumor k-mer database by (i) extracting a set of subsequences comprising at least 24 nucleotides from said RNA-sequences from normal cells and said tumor RNA-sequences; (g) comparing the sequences of the MAPs obtained in (a) with the sequences of the tumor-specific proteome database of (c) and the personalized tumor proteome database of (d) to identify the MAPs; and (h) identifying a tumor antigen candidate among the MAPs identified in (f), wherein a tumor antigen candidate corresponds to a MAP (1) whose sequence is not present in the personalized normal proteome database; and (2) (i) whose sequence is present in the personalized tumor proteome database; and/or (ii) whose encoding sequence is overexpressed or overrepresented in said tumor k-mer database relative to said normal k-mer database.
89 . The method of claim 76 , wherein said method further comprises selecting MAPs having a length of 8 to 11 amino acids.
90 . The method of claim 76 , further comprising comparing the coding sequence of said tumor antigen candidate to sequences from normal tissues.
91 . The method of claim 76 , further comprising assessing the binding of the tumor antigen candidate to an MHC molecule.
92 . The method of claim 91 , wherein said binding is assessed using an MHC binding prediction algorithm.
93 . The method of claim 76 , further comprising assessing the frequency of T cells recognizing the tumor antigen candidate in a cell population.
93 . The method of claim 76 , further comprising assessing the ability of the tumor antigen candidate to induce T cell activation.
94 . The method of claim 93 , wherein the ability of the tumor antigen candidate to induce T cell activation is assessed by measuring cytokine production by T cells contacted with cells having said tumor antigen candidate bound to MHC class I molecules at their cell surface.
95 . The method of claim 76 , further comprising assessing the ability of said tumor antigen candidate to induce T-cell-mediated tumor cell killing and/or to inhibit tumor growthJoin the waitlist — get patent alerts
Track US2023242583A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.