Method and system for identifying genomic regions with condition sensitive occupancy/positioning of nucleosomes and/or chromatin
Abstract
Aspects of the present invention relate at least in part to identification of regions within the genome that are sensitive to a condition. Particularly, although not exclusively, embodiments of the present invention relate to a method and system for identifying regions of the genome where DNA protection from digestion changes in response to a condition, e.g. the nucleosomal organisation is different as compared to a genomic region in a subject without the condition. The condition may be a pathological disorder e.g. cancer, or a variation of a healthy state e.g. depending on person's lifestyle or age. In certain embodiments, the method and systems may identify regions within the genome which differ in patients with sub-sets of the same condition. Aspects of the present invention comprise identification, stratification and monitoring of subjects suffering from a condition by sequencing predetermined regions of a genome.
Claims
exact text as granted — not AI-modified1 . A method for identifying genomic regions with condition-sensitive occupancy of nucleosomes and/or chromatin macromolecules, the method comprising:
(a) comparing, to a reference genome sequence, at least a portion of:
(i) a plurality of first nucleic acid sequence datasets, each first nucleic acid sequence dataset being obtained from a plurality of digestion-protected regions of a plurality of nucleic acid molecules obtained from a first subject with a first condition, wherein the plurality of first nucleic acid sequence datasets each comprise a plurality of first nucleic acid fragments;
wherein a genomic location of each of the plurality of first nucleic acid fragments is identified;
(b) comparing, to the reference genome sequence, at least a portion of:
(i) a plurality of second nucleic acid sequence datasets, each second nucleic acid sequence dataset being obtained from a plurality of digestion-protected regions of a plurality of nucleic acid molecules obtained from a second subject with a second condition, wherein the plurality of second nucleic acid sequence datasets each comprise a plurality of second nucleic acid fragments;
wherein a genomic location of each of the plurality of second nucleic acid fragments is identified;
(c) (i) determining an average normalised occupancy of digestion-protected regions of nucleic acid fragments per genomic region (O 1 ) of each of the first subjects with the first condition and (ii) determining an average normalised occupancy of digestion-protected regions of nucleic acid fragments per genomic region (O 2 ) of each of the second subjects with the second condition; wherein the method optionally comprises repeating step (c) for one or more other conditions (N) to determine average normalised occupancy of digestion-protected regions of nucleic acid fragments per genomic region (O N ) of subjects with condition N; (d) determining one or more stable-nucleosome regions of the genome in which the variation of the occupancy of digestion-protected regions of nucleic acid fragments in each of the first subjects with the first condition is below a first set threshold value; (e) determining one or more stable-nucleosome regions of the genome in which the variation of the occupancy of digested-protected regions of nucleic acid fragments in each of the second subjects with the second condition is below a second set threshold value; (f) comparing (i) the one or more stable-nucleosome regions of the genome of the first subjects with the first condition and (ii) the one or more stable-nucleosome regions of the genome of the second subjects with the second condition; and (g) identifying one or more regions of the genome which have stable-nucleosome regions in the first condition and stable-nucleosome regions in the second condition, and which have a difference between the average normalised occupancy of digestion-protected regions of nucleic acid fragments in the first condition and the average normalised occupancy of digestion-protected regions of nucleic acid fragments in the second condition that is larger or smaller than a set threshold value, to thereby identify one or more condition-sensitive-genomic regions.
2 . The method according to claim 1 , wherein the stable nucleosome region is a stable-nucleosome-occupancy region.
3 . The method according to claim 1 , wherein step (d) comprises (ii) determining one or more fuzzy-nucleosome regions of the genome in which the variation of the occupancy of the nucleic acid fragments in digestion-protected regions in each of the first subjects with the first condition is above a first set threshold value.
4 . (canceled)
5 . (canceled)
6 . A method for identifying genomic regions with condition-sensitive positioning of nucleosomes and/or chromatin macromolecules, the method comprising:
(a) comparing, to a reference genome sequence, at least a portion of:
(i) a plurality of first nucleic acid sequence datasets, each first nucleic acid sequence dataset being obtained from a plurality of digestion-protected regions of a plurality of nucleic acid molecules obtained from a first subject with a first condition, wherein the plurality of first nucleic acid sequence datasets each comprise a plurality of first nucleic acid fragments;
wherein a genomic location of each of the plurality of first nucleic acid fragments is identified;
(b) comparing, to the reference genome sequence, at least a portion of:
(i) a plurality of second nucleic acid sequence datasets, each second nucleic acid sequence dataset being obtained from a plurality of digestion-protected regions of a plurality of nucleic acid molecules obtained from a second subject with a second condition, wherein the plurality of second nucleic acid sequence datasets each comprise a plurality of second nucleic acid fragments;
wherein a genomic location of each of the plurality of second nucleic acid fragments is identified;
(c) (i) determining the genomic locations defined by the region start and region end coordinates of digestion-protected regions of nucleic acid fragments of each of the first subjects with the first condition and (ii) determining the genomic locations of digestion-protected regions of nucleic acid fragments of each of the second subjects with the second condition; wherein the method optionally comprises repeating step (c) for one or more other conditions (N) to determine the genomic locations of digestion-protected regions of nucleic acid fragments of subjects with condition N; (d) determining one or more stable-nucleosome-positioning regions of the genome in which the variation of the genomic locations defined by the start and end coordinates or the center coordinates of digestion-protected regions of nucleic acid fragments in each of the first subjects with the first condition is below a first set threshold value; (e) determining one or more stable-nucleosome-positioning regions of the genome in which the variation of the genomic locations defined by the start and end coordinates or the center coordinates of digested-protected regions of nucleic acid fragments in each of the second subjects with the second condition is below a second set threshold value; (f) comparing (i) the one or more stable-nucleosome-positioning regions of the genome of the first subjects with the first condition and (ii) the one or more stable-nucleosome-positioning regions of the genome of the second subjects with the second condition; (g) identifying one or more condition-sensitive regions of the genome which have stable-nucleosome-positioning regions in the first condition and which changed their genomic locations in the second condition by a value larger or smaller than a set threshold value, to thereby identify one or more condition-sensitive genomic regions where the nucleosome locations shifted to form the dataset of shifted nucleosomes; (h) identifying one or more condition-sensitive regions of the genome which have stable-nucleosome-positioning regions in the first condition and which do not overlap with stable-nucleosome-positioning regions in the second condition, to thereby identify one or more condition-sensitive genomic regions that preferentially contain a nucleosome in the first condition and preferentially lose nucleosomes in the second condition to form the dataset of lost nucleosomes; and (i) identifying one or more condition-sensitive regions of the genome which have stable-nucleosome-positioning regions in the second condition and which do not overlap with stable-nucleosome-positioning regions in the first condition, to thereby identify one or more condition-sensitive genomic regions that preferentially do not contain a nucleosome in the first condition and which gained nucleosomes in the second condition to form the dataset of gained nucleosomes.
7 . The method according to claim 1 , which further comprises:
identifying one or more regions of the genome which comprise condition-sensitive regions for combinations of different conditions by determining intersections, unions and/or exclusions of condition-sensitive regions, wherein:
the condition-sensitive regions comprise regions with changed DNA protection by nucleosomes and/or other chromatin complexes according to claim 1 or regions with changed nucleosome positioning according to claim 6 ;
intersections define regions sensitive to each of a plurality of conditions,
unions are composed of condition-sensitive regions defined for more than two pairs of conditions of interest; wherein unions define regions sensitive to at least one of a plurality of conditions; and
exclusions define regions sensitive to a set of conditions but not sensitive to a differing set of conditions; and
refining the set of condition-sensitive-nucleosome genomic regions by including or excluding condition-sensitive-genomic regions defined for comorbidities such as ageing.
8 .- 11 . (canceled)
12 . The method according to claim 1 , wherein step (c) comprises splitting the reference genomic sequence into regions of a predetermined length and determining average normalised occupancy of protected regions of nucleic acid fragments within each region.
13 . The method according to claim 12 , wherein the sizes of the genomic regions for the calculation of normalised occupancy are between 10 base pairs (bp) and 100000 bp in length, optionally wherein the regions are 50-150 bp in length, wherein optionally the sizes of the genomic regions for the calculation of normalised occupancy are between 10 base pairs (bp) and 10000 bp in length.
14 . (canceled)
15 . (canceled)
16 . The method according to claim 1 , which comprises applying a pairwise dissimilarity threshold value defining a minimal acceptable value of the relative difference of an average normalised occupancy of protected regions of nucleic acid fragments in the first condition (O 1 ) and the second condition (O 2 ), wherein the relative difference is defined as (O 2 −O 1 )/(01+Oz).
17 . The method according to claim 1 , wherein the condition-sensitive region of the genome comprises a difference between the subjects with the first condition and the subjects with the second condition in one or more of the following, or in a combination thereof:
(i) average profile of the occupancy of protected regions of nucleic acid fragments; (ii) genomic location of the center of nucleosome; (iii) genomic locations of the start and end of the nucleosome; (iv) size of linker DNA between nucleosomes; (v) stability of nucleosomes against digestion by MNase or another nuclease; (vi) stability of the nucleosome against partial DNA unwrapping; (vii) stability of the nucleosome against partial disassembly of the histone octamer; (viii) accessibility of DNA as measured by ATAC-seq or/and DNase-seq; and/or (ix) protein binding as measured by ChIP-seq or CUT&RUN or CUT&Tag.
18 . The method according to claim 1 , which further comprises, prior to step (a) and/or step (b):
(i) obtaining first nucleic acid sequence data from the digestion-protected regions of the nucleic acid molecules from a plurality of subjects with the first condition, wherein the first nucleic acid sequence data comprises a plurality of first nucleic acid fragments; and/or (ii) obtaining second nucleic acid sequence data obtained from digestion-protected regions of nucleic acid molecules from a plurality of subjects with the second condition, wherein the second nucleic acid sequence data comprises a plurality of second nucleic acid fragments.
19 .- 29 . (canceled)
30 . The method according to claim 1 , wherein the first condition is a pathological disorder selected from a cancer, a sub-type of cancer, a viral infection, a bacterial infection, an inflammatory disorder, sepsis, cardiovascular disorder, acute cellular rejection, benign kidney disease, benign liver disease, hepatitis B, inflammatory bowel disease, lupus, diabetes, Crohn's disease, myocarditis, pericarditis, multiple sclerosis, psoriasis and a neurological disease.
31 . (canceled)
32 . (canceled)
33 . The method according to claim 1 , wherein the second condition is the absence of a pathological disorder.
34 . (canceled)
35 . The method according to claim 1 , wherein the first condition is a pathological disorder and the second condition is the absence of the pathological disorder.
36 .- 40 . (canceled)
41 . A system for identifying condition-sensitive regions in cell-free DNA, the system comprising a computer program configured to;
(a) compare, to a reference genome sequence, at least a portion of:
(i) a plurality of first nucleic acid sequence datasets, each first nucleic acid sequence dataset being obtained from a plurality of digestion-protected regions of a plurality of nucleic acid molecules obtained from a first subject with a first condition, wherein the plurality of first nucleic acid sequence datasets each comprise a plurality of first nucleic acid fragments;
wherein a genomic location of each of the plurality of first nucleic acid fragments is identified;
(b) compare, to the reference genome sequence, at least a portion of:
(i) a plurality of second nucleic acid sequence datasets, each second nucleic acid sequence dataset being obtained from a plurality of digestion-protected regions of a plurality of nucleic acid molecules obtained from a second subject with a second condition, wherein the plurality of second nucleic acid sequence datasets each comprise a plurality of second nucleic acid fragments;
wherein a genomic location of each of the plurality of second nucleic acid fragments is identified;
(c) (i) determine an average normalised occupancy of digestion-protected regions of nucleic acid fragments per genomic region (O 1 ) of the first subjects with the first condition and (ii) determine an average occupancy of digestion-protected regions of nucleic acid fragments per genomic region (O 2 ) of the second subjects with the second condition; optionally repeat step (c) for any other condition N to determine average normalised occupancy of protected regions of nucleic acid fragments per genomic region (O N ) of subjects with condition N; (d) determine one or more stable-nucleosome regions of the genome in which the variation of the occupancy of protected regions of nucleic acid fragments in each of the first subjects with the first condition is below a first set threshold value; (e) determine one or more stable-nucleosome regions of the genome in which the variation of the occupancy of protected regions of nucleic acid fragments in each of the second subjects with the second condition is below a second set threshold value; (f) compare (i) the one or more stable-nucleosome regions of the genome of the first subjects with the first condition and (ii) the one or more stable-nucleosome regions of the genome of the second subjects with the second condition; and (g) identify one or more regions of the genome which have stable-nucleosome regions in the first condition and stable-nucleosome regions in the second condition, and which have the difference between the average occupancy of protected regions of nucleic acid fragments in the first condition and the average occupancy of protected regions of nucleic acid fragments in the second condition larger or smaller than set threshold values, to thereby identify one or more condition-sensitive regions.
42 . The system according to claim 41 , wherein the stable nucleosome region is a stable-nucleosome-occupancy region.
43 . The system according to claim 41 , wherein step (d) comprises (ii) determining one or more fuzzy-nucleosome regions of the genome in which the variation of the occupancy of the nucleic acid fragments in digestion-protected regions in each of the first subjects with the first condition is above a first set threshold value.
44 . (canceled)
45 . (canceled)
46 . A system for identifying condition-sensitive regions in cell-free DNA, the system comprising a computer program configured to:
(a) comparing, to a reference genome sequence, at least a portion of:
(i) a plurality of first nucleic acid sequence datasets, each first nucleic acid sequence dataset being obtained from a plurality of digestion-protected regions of a plurality of nucleic acid molecules obtained from a first subject with a first condition, wherein the plurality of first nucleic acid sequence datasets each comprise a plurality of first nucleic acid fragments;
wherein a genomic location of each of the plurality of first nucleic acid fragments is identified;
(b) comparing, to the reference genome sequence, at least a portion of:
(i) a plurality of second nucleic acid sequence datasets, each second nucleic acid sequence dataset being obtained from a plurality of digestion-protected regions of a plurality of nucleic acid molecules obtained from a second subject with a second condition, wherein the plurality of second nucleic acid sequence datasets each comprise a plurality of second nucleic acid fragments;
wherein a genomic location of each of the plurality of second nucleic acid fragments is identified;
(c) (i) determining the genomic locations defined by the region start and region end coordinates of digestion-protected regions of nucleic acid fragments of each of the first subjects with the first condition and (ii) determining the genomic locations of digestion-protected regions of nucleic acid fragments of each of the second subjects with the second condition; wherein the method optionally comprises repeating step (c) for one or more other conditions (N) to determine the genomic locations of digestion-protected regions of nucleic acid fragments of subjects with condition N; (d) determining one or more stable-nucleosome-positioning regions of the genome in which the variation of the genomic locations defined by the start and end coordinates of digestion-protected regions of nucleic acid fragments in each of the first subjects with the first condition is below a first set threshold value; (e) determining one or more stable-nucleosome-positioning regions of the genome in which the variation of the genomic locations defined by the start and end coordinates of digested-protected regions of nucleic acid fragments in each of the second subjects with the second condition is below a second set threshold value; (f) comparing (i) the one or more stable-nucleosome-positioning regions of the genome of the first subjects with the first condition and (ii) the one or more stable-nucleosome-positioning regions of the genome of the second subjects with the second condition; (g) identifying one or more condition-sensitive regions of the genome which have stable-nucleosome-positioning regions in the first condition and which changed their genomic locations in the second condition by a value larger or smaller than a set threshold value, to thereby identify one or more condition-sensitive genomic regions where the nucleosome locations shifted (“shifted nucleosomes”); (h) identifying one or more condition-sensitive regions of the genome which have stable-nucleosome-positioning regions in the first condition and which do not overlap with stable-nucleosome-positioning regions in the second condition, to thereby identify one or more condition-sensitive genomic regions that preferentially contain a nucleosome in the first condition and preferentially lost nucleosome in the second condition (“lost nucleosomes”); and (i) identifying one or more condition-sensitive regions of the genome which have stable-nucleosome-positioning regions in the second condition and which do not overlap with stable-nucleosome-positioning regions in the first condition, to thereby identify one or more condition-sensitive genomic regions that preferentially do not contain a nucleosome in the first condition and gained nucleosome in the second condition (“gained nucleosomes”).
47 . The system according to claim 41 , which is further configured to:
(h) identify one or more regions of the genome which comprise condition-sensitive regions for combinations of different conditions by determining intersections, unions or exclusions of condition-sensitive regions, where intersections define regions sensitive to each of several conditions of interest, unions define regions sensitive to at least one of several conditions of interest and exclusions define regions sensitive to some conditions but not sensitive to other specified conditions (for example, sensitive to cancer but not sensitive to ageing); and refine the set of condition-sensitive genomic regions by including or excluding condition-sensitive regions defined for comorbidities.
48 . A method of identifying a condition in a subject, the method comprising:
(a) defining one or more characteristics for a set of condition-specific regions; (b) defining the set of condition-specific regions by performing a method for identifying genomic regions with condition-sensitive occupancy or positioning of nucleosomes and/or chromatin macromolecules as claimed in claim 1 ; (c) obtaining nucleic acid sequence data from at least a portion of cell free DNA (cfDNA) isolated from a sample derived from the subject, wherein the subject is a first subject in which a condition is to be determined; (d) performing an alignment of sequenced data to a reference genome to define the genomic coordinates of sequenced reads; (e) calculating a normalised occupancy of cfDNA per genomic region, separately for each sample; (f) creating a reference set of samples, each of which are known to be obtained from a subject having a predetermined condition; (g) calculating an average normalised occupancy of cfDNA, separately for each sample in the reference set of step (f) for each condition-specific region; (h) performing dimensionality reduction analysis on (1) the sample obtained from the first subject in which the condition needs to be determined and (2) the samples from the reference set of samples; and (i) performing a classification of the sample from the first subject based on the similarity of the average normalised cfDNA occupancy in condition-sensitive regions to clusters formed by the samples from the reference set.
49 . (canceled)
50 . The method according to claim 48 , wherein the normalisation is performed by dividing the number of protected regions of nucleic acid fragments in a predetermined genomic region by the average occupancy for a predetermined genomic region in a predetermined sample in a predetermined condition.
51 .- 54 . (canceled)Join the waitlist — get patent alerts
Track US2024352514A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.