US2022207018A1PendingUtilityA1

Large data set negative information storage model

Assignee: H LEE MOFFITT CANCER CENTER & RES INSTITUTE INCPriority: Aug 13, 2015Filed: Dec 13, 2021Published: Jun 30, 2022
Est. expiryAug 13, 2035(~9 yrs left)· nominal 20-yr term from priority
G06F 16/215G16B 50/00G16B 30/00G06F 16/2365G06F 16/22G16B 50/50G16B 30/10
38
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for storing large data sets, such as genetic sequence information. Within a “targeted subset” of positions with information, the system stores, both variant states and missing states at each position. Reference states are not stored, but are inferred within the targeted subset when neither a variant nor a missing state is stored at a given position. The absence of a variant state at a given position is assumed to be a reference state. The criteria for missing data are defined in pre-processing and are customizable based on the use case. For example, each data point may represent the genetic information of a sample at a position in the genome. The targeted subset may represent those positions that were included in a sequencing test.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . A method of querying and decompressing a negative storage model stored in a database, comprising:
 receiving a query associated with a position stored in the database in the negative storage model;   determining if the position is in a targeted subset of a full dataset, and if so;
 determining if the position is a variant and returning the variant to the full dataset; 
 determining if the position is missing, and if so, returning a value of missing to the full dataset; 
 determining if the position is neither variant nor missing, and if so, inferring that the position is a reference value and returning the reference value to the full dataset; and 
 determining if the position is not in the targeted subset of the full dataset, and if so, inferring that the position is the reference value and returning the reference value to the full dataset. 
   
     
     
         3 . The method of  claim 2 , wherein the full dataset is reconstituted from the subset, variant, and missing states only. 
     
     
         4 . The method of  claim 2 , wherein the negative storage model stored in the database has a size that is smaller than a size of the full dataset. 
     
     
         5 . The method of  claim 2 , further comprising retrieving additional information from the database related to the variation in the database. 
     
     
         6 . The method of  claim 5 , wherein the additional information comprises the actual value of the variation. 
     
     
         7 . The method of  claim 2 , wherein a missing value is user defined for the targeted subset. 
     
     
         8 . The method of  claim 7 , further comprising determining plural missing values. 
     
     
         9 . A database storage apparatus for querying a compressed negative storage model, comprising:
 a processor;   a memory that contains computer executable instructions that when executed by the processor causes the database storage apparatus to:   receive a query associated with a position stored in the database in the negative storage model;   determine if the position is in a targeted subset of a full dataset, and if so;
 determine if the position is a variant and returning the variant to the full dataset; 
 determine if the position is missing, and if so, returning a value of missing to the full dataset; 
 determine if the position is neither variant nor missing, and if so, inferring that the position is a reference value and returning the reference value to the full dataset; and 
 determine if the position is not in the targeted subset of the full dataset, and if so, inferring that the position is the reference value and returning the reference value to the full dataset. 
   
     
     
         10 . The database storage apparatus of  claim 2 , wherein the full dataset is reconstituted from the subset, variant, and missing states only. 
     
     
         11 . The database storage apparatus of  claim 2 , wherein the negative storage model stored in the database has a size that is smaller than a size of the full dataset. 
     
     
         12 . The database storage apparatus of  claim 2 , wherein additional information is retrievable from the database that is related to the variation in the database. 
     
     
         13 . The database storage apparatus of  claim 5 , wherein the additional information comprises the actual value of the variation. 
     
     
         14 . The database storage apparatus of  claim 2 , wherein a missing value is user defined for the targeted subset. 
     
     
         15 . The database storage apparatus of  claim 7 , wherein plural missing values are retrievable from the database.

Join the waitlist — get patent alerts

Track US2022207018A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.