US2004171051A1PendingUtilityA1

Method and system for detecting near identities in large DNA databases

Assignee: ZYMOGENETICS INCPriority: Jan 31, 2000Filed: Jan 21, 2004Published: Sep 2, 2004
Est. expiryJan 31, 2020(expired)· nominal 20-yr term from priority
Inventors:James Holloway
G16B 30/10G16B 30/20G16B 50/00G16B 30/00
67
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present invention provides a solution to the needs described above through a system and method for efficiently detecting near identities in large DNA databases. The system and method disclosed herein make use of an algorithm used to construct and maintain unique DNA databases wherein the unique database contains no two DNA sequences such that one is nearly identical to a region of the other. The system and method are applicable to problems such as an all against all comparison of all the DNA sequences in a large DNA database, clustering and assembling ESTs into the cDNAs that generated the ESTs, mapping assembled ESTs onto genomic sequence, mapping cDNAs onto genomic sequences and locating alternately spliced cDNAs.

Claims

exact text as granted — not AI-modified
1 . A method for creating a unique DNA genome database, comprising the steps of: 
 providing available genomic sequence data in a first database    enumerating regions of identity between genomic sequences in the first data base and other genomic sequences in the first database as if the first database was also a query database, the enumerating being done on a computer having a processor, memory, input/output mechanisms; and    removing from the first database, genomic sequences that are nearly identical to a region of a longer genomic sequence, whereby a unique DNA genome database is created.    
     
     
         2 . A system for creating a unique DNA genome database, comprising; 
 means for providing available genomic sequence data in a first database    means for enumerating regions of identity between genomic sequences in the first data base and other genomic sequences in the first database as if the first database was also a query database, the enumerating being done on a computer having a processor, memory, input/output mechanisms; and    means for removing from the first database, genomic sequences that are nearly identical to a region of a longer genomic sequence, whereby a unique DNA genome database is created.    
     
     
         3 . A computer program for finding near identities in a DNA sequence database, comprising: 
 a first code mechanism for comparing a DNA sequence on a query database to DNA sequences on a data database, wherein a tag array (designated as Qtags) is generated for each of the DNA sequences on the query database and wherein a tag array (designated Dtags) is generated for each of the DNA sequences on the data database; and    a second code mechanism for comparing each Qtag to each Dtag, using a comparison model, wherein near identities of sequences in the two databases are identified.

Join the waitlist — get patent alerts

Track US2004171051A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.