Method for predicting regulatory elements in repetitive sequences using transcription factor binding sites
Abstract
Repeat sequences are the most abundant in the extragenic region of genomes, while a large number of regulatory elements are found in this region. The invention attempts to mine rules on how combinations of individual binding sites are distributed in repeat sequences. These mined association rules would facilitate identifying gene classes regulated by similar mechanisms and accurately predicting regulatory elements. Herein, the combinations of transcription factor binding sites in the repeat sequences are obtained, and data mining techniques are applied to mine the association rules from the combinations of binding sites. In addition, the associations are further pruned to remove insignificant associations and obtain a set of discovered associations. The discovered association rules are used to partially classify the repeat sequences in the repeat sequence database.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for predicting regulatory elements in repetitive sequences using transcription factor binding sites, comprising:
preprocessing the transcription factor binding sites in a transcription factor binding site database and mapping the transcription factor binding sites to transcription factor names; mapping the transcription factor binding sites in the transcription factor binding site database to repeat sequences in a repeat sequence database in order to find combinations of transcription factors in the repeat sequences; applying a data mining approach to generate association rules; pruning a portion of the generated association rules by using a significance test; classifying the remained association rules to cover and non-cover sets after pruning; and using the remained association rules to classify the repeat sequences in the repeat sequence database.
2 . The method as claimed in claim 1 , wherein the transcription factor binding site database comprises a TRANSFAC database.
3 . The method as claimed in claim 1 , wherein the significance test comprises a Chi-square test.
4 . The method as claimed in claim 1 , wherein the step of applying the data mining approach comprises the following steps:
inputting a set of data-sequences, wherein each data-sequence is a list of transactions and each transaction is a set of items; providing a plurality of sequential patterns, wherein each sequential pattern consists of a list of sets of items; and finding the sequential patterns with a user-specified minimum support in the data-sequences, where the support of a sequential pattern is a percentage of data-sequences that contain the pattern.
5 . A method for mining association rules from combinations of transcription factor binding sites in repeat sequences, comprising:
preprocessing the transcription factor binding sites in a transcription factor binding site database and mapping the transcription factor binding sites to transcription factor names; mapping the transcription factor binding sites in the transcription factor binding site database to repeat sequences in a repeat sequence database in order to find the combinations of transcription factors in the repeat sequences; applying a data mining approach to generate association rules; using a significance test to prune a portion of the association rules; and classifying the remained association rules to cover and non-cover sets.
6 . The method as claimed in claim 5 , wherein the transcription factor binding site database comprises a TRANSFAC database.
7 . The method as claimed in claim 5 , wherein the significance test comprises a Chi-square test.
8 . The method as claimed in claim 5 , wherein the step of applying the data mining approach comprises the following steps:
inputting a set of data-sequences, wherein each data-sequence is a list of transactions and each transaction is a set of items; providing a plurality of sequential patterns, wherein each sequential pattern consists of a list of sets of items; and finding the sequential patterns with a user-specified minimum support in the data-sequences, where the support of a sequential pattern is a percentage of data-sequences that contain the pattern..
9 . A computerized system for predicting regulatory elements in repetitive sequences using transcription factor binding sites, wherein the system can assess the transcription factor binding site database and a repeat sequence database, the system comprising:
means for inputting commands from a user; means for storing; means for preprocessing the transcription factor binding sites in the transcription factor binding site database and mapping the transcription factor binding sites to transcription factor names; means for mapping the transcription factor binding sites in the transcription factor binding site database to repeat sequences in the repeat sequence database, in order to find the combinations of transcription factors in the repeat sequences; means for generating association rules by applying a data mining approach; means for pruning a portion of the mined association rules using a significance test; means for classifying the remained association rules to cover and non-cover sets; means for classifying the repeat sequences in the repeat sequence database using the mined association rules; and means for outputting.
10 . The system as claimed in claim 10 , wherein the transcription factor binding site database comprises a TRANSFAC database.
11 . The method as claimed in claim 10 , wherein the significance test comprises a Chi-square test.
12 . The method as claimed in claim 10 , wherein the data mining approach comprises the following steps:
inputting a set of data-sequences, wherein each data-sequence is a list of transactions and each transaction is a set of items; providing a plurality of sequential patterns, wherein each sequential pattern consists of a list of sets of items; and finding the sequential patterns with a user-specified minimum support in the data-sequences, where the support of a sequential pattern is a percentage of data-sequences that contain the pattern.
13 . A storage system comprising an operating program for predicting regulatory elements in repetitive sequences using transcription factor binding sites, wherein the program comprises instructions for causing the system to:
preprocess the transcription factor binding sites in a transcription factor binding site database and mapping the transcription factor binding sites to transcription factor names; map the transcription factor binding sites in the transcription factor binding site database to repeat sequences in a repeat sequence database in order to find combinations of transcription factors in the repeat sequences; apply a data mining approach to generate association rules; prune a portion of the generated association rules by using a significance test; classify the remained association rules to cover and non-cover sets after pruning; and classify the repeat sequences in the repeat sequence database using the remained association rules.
14 . The system as claimed in claim 13 , wherein the transcription factor binding site database comprises a TRANSFAC database.
15 . The method as claimed in claim 13 , wherein the significance test comprises a Chi-square test.
16 . The method as claimed in claim 13 , wherein the application of the data mining approach comprises the following steps:
inputting a set of data-sequences, wherein each data-sequence is a list of transactions and each transaction is a set of items; providing a plurality of sequential patterns, wherein each sequential pattern consists of a list of sets of items; and finding the sequential patterns with a user-specified minimum support in the data-sequences, where the support of a sequential pattern is a percentage of data-sequences that contain the pattern.
17 . A storage system comprising an operating program for mining association rules from combinations of transcription factor binding sites in repeat sequences, wherein the program comprises instructions for causing the system to:
preprocess the transcription factor binding sites in a transcription factor binding site database and mapping the transcription factor binding sites to transcription factor names; map the transcription factor binding sites in the transcription factor binding site database to repeat sequences in a repeat sequence database in order to find combinations of transcription factors in the repeat sequences; apply a data mining approach to generate association rules; use a significance test to prune a portion of the generated association rules; and classify the remained association rules to cover and non-cover sets.
18 . The system as claimed in claim 17 , wherein the transcription factor binding site database comprises a TRANSFAC database.
19 . The method as claimed in claim 17 , wherein the significance test comprises a Chi-square test.
20 . The method as claimed in claim 17 , wherein the application of the data mining approach comprises the following steps:
inputting a set of data-sequences, wherein each data-sequence is a list of transactions and each transaction is a set of items; providing a plurality of sequential patterns, wherein each sequential pattern consists of a list of sets of items; and finding the sequential patterns with a user-specified minimum support in the data-sequences, where the support of a sequential pattern is a percentage of data-sequences that contain the pattern.Join the waitlist — get patent alerts
Track US2003068617A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.