Device and method for detecting tumor mutation burden (tmb) based on capture sequencing
Abstract
A device and a method for detecting tumor mutation burden (TMB) based on capture sequencing are disclosed. The device includes: a panel design module configured to uniformly add population single-nucleotide polymorphism (SNP) sites to a genome and screen out gene regions that show the highest consistency with whole exome sequencing (WES); a data acquisition module configured to acquire tissue and plasma samples of a target object and acquire sequencing data of the samples; an alignment module configured to align the sequencing data with a reference genome to acquire mutation data results; a somatic mutation analysis module configured to perform somatic analysis on the mutation data results to obtain somatic mutation results; a filtering module configured to remove unreal mutation sites from the somatic mutation results; and a calculation module configured to calculate the TMB.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A device for detecting tumor mutation burden (TMB) based on a capture sequencing, comprising:
a panel design module, wherein the panel design module is configured to uniformly add population single-nucleotide polymorphism (SNP) sites to a genome and screen out genome regions, and the genome regions show a highest consistency with whole exome sequencing (WES); a data acquisition module, wherein the data acquisition module is configured to acquire tissue and plasma samples of a target object and acquire sequencing data of the tissue and plasma samples based on the genome regions screened out by the panel design module; an alignment module, wherein the alignment module is configured to align the sequencing data acquired by the data acquisition module with a reference genome to acquire mutation data results; a somatic mutation analysis module, wherein the somatic mutation analysis module is configured to perform a somatic analysis on the mutation data results obtained by the alignment module to obtain somatic mutation results; a filtering module, wherein the filtering module is configured to remove unreal mutation sites from the somatic mutation results obtained by the somatic mutation analysis module to obtain real somatic mutation sites; and a calculation module, wherein the calculation module is configured to calculate the TMB according to a number of the real somatic mutation sites obtained by the filtering module.
2 . The device according to claim 1 , wherein
the panel design module comprises a uniform site design unit and a region screening unit, wherein, the uniform site design unit is configured to screen out the genome regions for designing probes according to a first preset rule and then uniformly add the population SNP sites, and the population SNP sites are screened out according to a second preset rule, and the region screening unit is configured to screen out the genome regions, and the genome regions show the highest consistency with the WES by machine learning on exons; the first preset rule comprises: removing gaps and regions with a mappability lower than 40 in the genome; and/or after the genome is divided according to a preset window and a step size, removing regions with a GC content higher than 30% and lower than 60%; and/or removing regions with a corresponding preset length, wherein the corresponding preset length comprises a preset number of sites with an Asian population heterozygosity greater than a preset threshold; and the second preset rule comprises: SNP sites with the Asian population heterozygosity greater than the preset threshold; SNP sites meeting Hardy-Weinberg equilibrium; and/or extending each of the SNP sites to a preset size on both sides to obtain a region and aligning the region with the reference genome, counting a number of positions in the region, wherein the positions are aligned to the reference genome, and removing regions with a number greater than the preset threshold.
3 . The device according to claim 1 , wherein
the data acquisition module comprises an acquisition unit and a quality control unit, wherein, the acquisition unit is configured to acquire raw data of the tissue and plasma samples of the target object, and the quality control unit is configured to perform a quality control on the raw data of the tissue and plasma samples separately to obtain the sequencing data; and/or the alignment module comprises a first alignment unit and a second alignment unit, wherein, the first alignment unit is configured to align the sequencing data with the reference genome to obtain an alignment result file, and the second alignment unit is configured to subject the alignment result file to a de-redundancy and a re-alignment in terms of InDel regions to obtain the mutation data results.
4 . The device according to claim 1 , wherein the device further comprises a specific baseline building module, and the specific baseline building module is configured to build different sequencing depth baselines and tumor fraction baselines for different sequencing depth intervals, sample types, and tumor fraction intervals.
5 . The device according to claim 4 , wherein
the somatic mutation analysis module is configured to perform the somatic analysis on the mutation data results obtained by the alignment module with VarDict or MuTect2 to obtain the somatic mutation results; or the somatic mutation analysis module is configured to select a corresponding sequencing depth baseline according to a sequencing depth and a sample type of the tissue and plasma samples and acquire the somatic mutation results based on an in silico germline subtraction algorithm.
6 . The device according to claim 4 , wherein
the filtering module is configured to filter according to annotation results of the somatic mutation results obtained by the somatic mutation analysis module to remove the unreal mutation sites and obtain the real somatic mutation sites; and a filtering rule of the filtering module comprises: removing in silico germline mutations according to sample types; filtering out sites with an annotated frequency less than 5% and an occurrence frequency more than 0.2% in a population database; filtering out known tumor-driven gene mutations; filtering out mutation sites manifested as non-germline sites with a predetermined population frequency; filtering out repeat regions or false positive sites generated from an alignment of homologous regions according to a pre-built noise baseline of FFPE sample feature sequence-specific error (SSE); filtering out panel of normal (PoN) sites with a frequency less than a sum of a mean value and 5-fold standard deviation for PoN sites; filtering out preset black-listed sites, wherein the preset black-listed sites have an occurrence frequency greater than 30% in populations or the preset black-listed sites have a population frequency greater than 20% in two sample types of FFPE samples, plasma samples, and blood cell samples; and/or screening out mutations meeting depth requirements according to a sequencing depth baseline of the different sequencing depth baselines and screening out mutations meeting a tumor fraction according to a tumor fraction baseline of the different tumor fraction baselines.
7 . A method for detecting TMB based on a capture sequencing, comprising the following steps:
uniformly adding population SNP sites to a genome and screening out genome regions, wherein the genome regions show a highest consistency with WES; acquiring tissue and plasma samples of a target object and acquiring sequencing data of the tissue and plasma samples based on the genome regions screened out; aligning the sequencing data with a reference genome to acquire mutation data results; performing a somatic analysis on the mutation data results to obtain somatic mutation results; removing unreal mutation sites from the somatic mutation results to obtain real somatic mutation sites; and calculating the TMB according to a number of the real somatic mutation sites.
8 . The method according to claim 7 , wherein
the step of uniformly adding population SNP sites to the genome and screening out the genome regions comprises: after the genome regions for designing probes are screened out according to a first preset rule, uniformly adding the population SNP sites screened out according to a second preset rule; the first preset rule comprises: removing gaps and regions with a mappability lower than 40 in the genome; and/or after the genome is divided according to a preset window and a step size, removing regions with a GC content higher than 30% and lower than 60%; and/or removing regions with a corresponding preset length, wherein the corresponding preset length comprises a preset number of sites with an Asian population heterozygosity greater than a preset threshold; and the second preset rule comprises: SNP sites with the Asian population heterozygosity greater than the preset threshold; SNP sites meeting Hardy-Weinberg equilibrium; and/or extending each of the SNP sites to a preset size on both sides to obtain a region and aligning the region with the reference genome, counting a number of positions in the region, wherein the positions are aligned to the reference genome, and removing regions with a number greater than the preset threshold.
9 . The method according to claim 7 , wherein
the step of uniformly adding population SNP sites to the genome and screening out the genome regions further comprises: counting a number of mutations in exons of the genome of each sample, selecting the exons according to a TMB value on the WES of the each sample to obtain selected exons, and ranking the selected exons based on importance; starting from a most important exon, adding a marked exon in sequence according to the ranking, and calculating a TMB value of an exon set after each addition and a correlation of the TMB value with a corresponding TMB value obtained from the WES to obtain a calculated correlation; and according to the calculated correlation, screening out the genome regions with the highest consistency with the WES.
10 . The method according to claim 7 , wherein
the step of acquiring the tissue and plasma samples of the target object and acquiring the sequencing data of the tissue and plasma samples based on the genome regions screened out comprises: acquiring raw data of the tissue and plasma samples of the target object, and performing a quality control on the raw data of the tissue and plasma samples separately to obtain sequencing data; and/or the step of aligning the sequencing data with the reference genome to acquire the mutation data results comprises: aligning the sequencing data with the reference genome to obtain an alignment result file, and subjecting the alignment result file to a de-redundancy and a re-alignment in terms of InDel regions to obtain the mutation data results.
11 . The method according to claim 7 , wherein the method for detecting the TMB further comprises a step of building different sequencing depth baselines and tumor fraction baselines for different sequencing depth intervals, sample types, and tumor fraction intervals.
12 . The method according to claim 11 , wherein
the step of performing the somatic analysis on the mutation data results to obtain the somatic mutation results comprises: performing the somatic analysis on the mutation data results obtained by the alignment module with VarDict or MuTect2 to obtain the somatic mutation results; or the step of performing the somatic analysis on the mutation data results to obtain the somatic mutation results comprises: selecting a corresponding sequencing depth baseline according to a sequencing depth and a sample type of the tissue and plasma samples; and acquiring the somatic mutation results based on an in silico germline subtraction algorithm.
13 . The method according to claim 11 , wherein
the step of removing the unreal mutation sites from the somatic mutation results to obtain the real somatic mutation sites comprises: filtering according to annotation results of the somatic mutation results obtained by the somatic mutation analysis module to remove the unreal mutation sites and obtain the real somatic mutation sites; and a filtering rule of the filtering module comprises: removing in silico germline mutations according to sample types; filtering out sites with an annotated frequency less than 5% and an occurrence frequency more than 0.2% in a population database; filtering out known tumor-driven gene mutations; filtering out mutation sites manifested as non-germline sites with a predetermined population frequency; filtering out repeat regions or false positive sites generated from alignment of homologous regions according to a pre-built noise baseline of FFPE sample feature sequence-specific error (SSE); filtering out panel of normal (PoN) sites with a frequency less than a sum of a mean value and 5-fold standard deviation for PoN sites; filtering out preset black-listed sites, wherein the preset black-listed sites have an occurrence frequency greater than 30% in populations or the preset black-listed sites have a population frequency greater than 20% in two sample types of FFPE samples, plasma samples, and blood cell samples; and/or screening out mutations meeting depth requirements according to a sequencing depth baseline of the different sequencing depth baselines and screening out mutations that meet a tumor fraction according to a tumor fraction baseline of the different tumor fraction baselines.
14 . A terminal device, comprising a memory, a processor, and computer programs, wherein the computer programs are stored in the memory and are running on the processor, wherein, when the computer programs are running on the processor, the steps of the method for detecting the TMB based on the capture sequencing according to claim 7 are implemented.
15 . A computer-readable storage medium, wherein the computer-readable storage medium stores computer programs, wherein, when the computer programs are executed by a processor, the steps of the method for detecting the TMB based on the capture sequencing according to claim 7 are implemented.
16 . The device according to claim 2 , wherein the device further comprises a specific baseline building module, and the specific baseline building module is configured to build different sequencing depth baselines and tumor fraction baselines for different sequencing depth intervals, sample types, and tumor fraction intervals.
17 . The device according to claim 3 , wherein the device further comprises a specific baseline building module, and the specific baseline building module is configured to build different sequencing depth baselines and tumor fraction baselines for different sequencing depth intervals, sample types, and tumor fraction intervals.
18 . The method according to claim 8 , wherein
the step of uniformly adding population SNP sites to the genome and screening out the genome regions further comprises: counting a number of mutations in exons of the genome of each sample, selecting the exons according to a TMB value on the WES of the each sample to obtain selected exons, and ranking the selected exons based on importance; starting from a most important exon, adding a marked exon in sequence according to the ranking, and calculating a TMB value of an exon set after each addition and a correlation of the TMB value with a corresponding TMB value obtained from the WES to obtain a calculated correlation; and according to the calculated correlation, screening out the genome regions with the highest consistency with the WES.
19 . The method according to claim 8 , wherein the method for detecting the TMB further comprises a step of building different sequencing depth baselines and tumor fraction baselines for different sequencing depth intervals, sample types, and tumor fraction intervals.
20 . The method according to claim 10 , wherein the method for detecting the TMB further comprises a step of building different sequencing depth baselines and tumor fraction baselines for different sequencing depth intervals, sample types, and tumor fraction intervals.Join the waitlist — get patent alerts
Track US2022072553A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.