Method to Map Protein Landscapes
Abstract
In shotgun proteomics, generally only a fraction of peptides from a parent protein are actually detected. Because a large portion of the protein sequence is not detected, it is often impossible to determine whether the expressed protein is present in a modified, spliced, or truncated form. Provided herein are methods and systems for analyzing polypeptides which allow for the increase of the mean sequence coverage of a protein concomitant with bioinformatics analysis in order to distinguish putative proteoforms with improved amino acid resolution. Aspects of the invention include (1) a deep sequencing strategy to provide more protein sequence coverage than is typically achieved, and (2) a computational approach to view protein expression across its full length and identify regions of the protein that are potentially subject to such regulation. This technology has global utility in proteomics and will be of particular use for the analysis of biosimilar protein drug therapeutics.
Claims
exact text as granted — not AI-modified1 . A method for analyzing a polypeptide having an amino acid sequence comprising the steps of:
a) digesting a first sample of the polypeptide with a first protease or chemical agent; b) digesting a second sample of the polypeptide with a second protease or chemical agent; c) generating tandem mass spectrometry data on each digested polypeptide sample comprising the steps of: generating a distribution of precursor ions during MS 1 stage ionization, fragmenting precursor ions having a mass-to-charge ratio (m/z) within a selected target m/z range during MS 2 stage fragmentation, thereby generating a plurality of product ions wherein the product ions correspond to portions of the amino acid sequence of the polypeptide, and measuring m/z and intensity of the product ions, thereby generating mass spectrometry data for each digested polypeptide sample; and d) combining mass spectrometry data from each digested polypeptide sample to generate comprehensive mass spectrometry data on the polypeptide, wherein the comprehensive mass spectrometry data provides sequence coverage for at least 20% of the full length amino acid sequence of the polypeptide.
2 . The method of claim 1 further comprising generating at least a partial consensus amino acid sequence for the polypeptide from the comprehensive mass spectrometry data.
3 . The method of claim 1 further comprising calculating abundances of amino acids for one or more selected portions of the polypeptide from the comprehensive mass spectrometry data, wherein an isotopic or chemical label is not attached to the polypeptide in order to calculate the abundances of amino acids.
4 . The method of claim 3 wherein the one or more selected portions comprises the N-terminus of the polypeptide.
5 . The method of claim 1 further comprising normalizing the measured intensities of the product ions from each digested polypeptide sample during generation of the comprehensive mass spectrometry data, and identifying a portion of the product ions as corresponding to one or more known amino acid sequence fragments of said polypeptide.
6 . The method of claim 5 further comprising performing k-means clustering analysis on normalized intensity data during generation of the comprehensive mass spectrometry data.
7 . The method of claim 1 further comprising digesting one or more additional samples of the polypeptide with one or more additional proteases or chemical agents, wherein the protease or chemical agent used for each sample is a different protease or chemical agent used to digest any other sample.
8 . The method of claim 7 further comprising digesting a third sample of the polypeptide with a third protease, digesting a fourth sample of the polypeptide with a fourth protease, digesting a fifth sample of the polypeptide with a fifth protease, and digesting a sixth sample of the polypeptide with a sixth protease.
9 . The method of claim 7 wherein the proteases or chemical agents are proteases selected from the group consisting of trypsin, Lys-N, Lys-C, Glu-C, chymotrypsin, Asp-N, and combinations thereof, wherein each selected protease is different for each sample.
10 . The method of claim 1 wherein the comprehensive mass spectrometry data provides sequence coverage for at least 50% of the full length amino acid sequence of the polypeptide.
11 . The method of claim 1 wherein the comprehensive mass spectrometry data provides sequence coverage for at least 80% of the full length amino acid sequence of the polypeptide.
12 . The method of claim 1 wherein the polypeptide is an antibody, antibody-drug conjugate, or a therapeutic protein.
13 . A method for analyzing two or more polypeptides comprising the steps of:
a) independently digesting a first sample of a first polypeptide and a first sample of a second polypeptide with a first protease or chemical agent; b) independently digesting a second sample of the first polypeptide and a second sample of the second polypeptide with a second protease or chemical agent; c) generating tandem mass spectrometry data on each digested polypeptide sample comprising the steps of: generating a distribution of precursor ions during MS 1 stage ionization, fragmenting precursor ions having a mass-to-charge ratio (m/z) within a selected target m/z range during MS 2 fragmentation, thereby generating a plurality of product ions wherein the product ions correspond to amino acid sequences of the polypeptides, and measuring m/z and intensity of the product ions, thereby generating mass spectrometry data for each digested polypeptide sample; and d) for each polypeptide, combining mass spectrometry data from each digested polypeptide sample of that polypeptide to generate comprehensive mass spectrometry data, wherein the comprehensive mass spectrometry data provides sequence coverage for at least 20% of the full length amino acid sequence for that polypeptide; and e) generating at least a partial consensus amino acid sequence for each polypeptide from the comprehensive mass spectrometry data or calculating abundances of amino acids for selected portions of each polypeptide from the comprehensive mass spectrometry data.
14 . The method of claim 13 where the first polypeptide is a control polypeptide and the second polypeptide is the same polypeptide which has undergone a suspected modification, splice, truncation, polymorphism or mutation.
15 . The method of claim 14 wherein the modification is a post-translational modification or the result of a single nucleotide polymorphism.
16 . The method of claim 13 where the first polypeptide is a control polypeptide and the second polypeptide is produced by a cell which has been administered a treatment.
17 . The method of claim 13 where the first polypeptide is a control therapeutic polypeptide, antibody, or antibody-drug conjugate, and the second polypeptide is a production therapeutic polypeptide, antibody, or antibody-drug conjugate made during a biochemical process or manufacturing process.
18 . The method of claim 13 further comprising independently digesting one or more additional samples of the first polypeptide and one or more additional samples of the second polypeptide proteases with one or more additional proteases or chemical agents, wherein the one or more additional proteases or chemical agents used for each additional sample is a different protease or chemical agent.
19 . The method of claim 18 wherein the proteases or chemical agents are proteases selected from the group consisting of trypsin, Lys-N, Lys-C, Glu-C, chymotrypsin, Asp-N, and combinations thereof, wherein each selected protease is different.
20 . The method of claim 13 further comprising independently digesting a third sample of the first polypeptide and a third sample of the second polypeptide with a third protease, digesting a fourth sample of the first polypeptide and a fourth sample of the second polypeptide with a fourth protease, digesting a fifth sample of the first polypeptide and a fifth sample of the second polypeptide with a fifth protease, and digesting a sixth sample of the first polypeptide and a sixth sample of the second polypeptide with a sixth protease.
21 . The method of claim 13 wherein the comprehensive mass spectrometry data provides sequence coverage for at least 50% of the amino acid sequence of the polypeptide.
22 . The method of claim 13 further comprising normalizing the measured intensities of the product ions from each digested polypeptide sample during generation of the comprehensive mass spectrometry data, and identifying a portion of the product ions as corresponding to one or more known amino acid sequence fragments for each polypeptide.
23 . The method of claim 22 further comprising performing k-means clustering analysis on normalized intensity data during generation of the comprehensive mass spectrometry data.
24 . The method of claim 13 further comprising comparing the consensus amino acid sequences of each polypeptide or the abundances of amino acids of each polypeptide, and identifying differences in amino acid sequence or amino acid abundance between the polypeptides.
25 . A system for analyzing a polypeptide having an amino acid sequence comprising:
a) an ion source for generating ions from a plurality of digested samples of the polypeptide; b) ion fragmentation optics in communication with the ion source for generating product ions; c) an ion detector in communication with the ion fragmentation optics for detecting ions according to their mass-to-charge ratios; d) a mass analyzer in communication with the ion detector, wherein the mass analyzer comprises a software program enabling the mass analyzer to: i) measure mass-to-charge ratios and intensity of the detected ions, thereby generating mass spectrometry data for each digested polypeptide sample; ii) normalize the measured intensities of the product ions from each digested polypeptide sample; and iii) combine mass spectrometry data from each digested polypeptide sample to generate comprehensive mass spectrometry data on the polypeptide, wherein the comprehensive mass spectrometry data provides sequence coverage for at least 20% of the full length amino acid sequence of the polypeptide.Join the waitlist — get patent alerts
Track US2018340941A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.