US2016232224A1PendingUtilityA1

Categorization and filtering of scientific data

Assignee: NEXTBIOPriority: Dec 16, 2005Filed: Aug 17, 2015Published: Aug 11, 2016
Est. expiryDec 16, 2025(expired)· nominal 20-yr term from priority
G06F 17/3053G06F 17/30327G06F 17/30598G06F 19/28G06N 20/00G16B 20/20G16B 20/10G16B 50/10G16B 50/00G06F 16/285G16B 20/00G06F 16/24578G06F 16/2246G16B 30/00
44
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present invention relates to methods, systems and apparatus for capturing, integrating, organizing, navigating and querying large-scale data from high-throughput biological and chemical assay platforms. It provides a highly efficient meta-analysis infrastructure for performing research queries across a large number of studies and experiments from different biological and chemical assays, data types and organisms, as well as systems to build and add to such an infrastructure. According to various embodiments, methods, systems and interfaces for associating experimental data, features and groups of data related by structure and/or function with chemical, medical and/or biological terms in an ontology or taxonomy are provided. According to various embodiments, methods, systems and interfaces for filtering data by data source information are provided, allowing dynamic navigation through large amounts of data to find the most relevant results for a particular query.

Claims

exact text as granted — not AI-modified
1 . A computer-implemented method of correlating chemical or biological categories with other information in a database, said method comprising:
 obtaining from the database a taxonomy of biological or chemical categories arranged in a hierarchical structure comprising at least one top-level category having at least one child category;   obtaining from the database a plurality of feature sets or feature groups, each feature set comprising two or more features of chemical or biological entities and statistical information associated with each of the two or more features and each feature group comprising a list of inter-related features of chemical or biological entities, wherein at least some features of the feature sets or feature groups have different names and are associated with each other and the feature sets or feature groups are obtained from across different experiments, platforms, or organisms;   obtaining from the database a plurality of globally unique mapping identifiers;   identifying, for each globally unique mapping identifier, one or more features associated with the globally unique mapping identifier;   mapping, for each globally unique mapping identifier, the identified one or more features to the globally unique mapping identifier, wherein at least some features having different names and being associated with each other are mapped to a same globally unique mapping identifier;   identifying, for each of a plurality of the categories in the taxonomy, contributing feature sets that contribute to scoring a category under consideration by identifying feature sets among the plurality of feature sets that are associated with the category under consideration or its child categories;   using the contributing feature sets of the category under consideration and globally unique mapping identifiers mapped to features of the contributing feature sets to calculate a correlation score indicating a correlation between the category under consideration and a feature, a feature set, or a feature group in the database; and   storing the correlation score in the database.   
     
     
         2 . The method of  claim 1 , wherein identifying contributing feature sets that contribute to scoring the category under consideration further comprises filtering the identified contributing feature sets to remove some feature sets. 
     
     
         3 . The method of  claim 1 , wherein the correlation score is calculated using pre-computed correlation scores indicating the correlation between the contributing feature sets and at least some of the other feature sets in the database. 
     
     
         4 . The method of  claim 1 , wherein the correlation score is calculated using pre-computed correlation scores indicating the correlation between the contributing feature sets and at least some of the feature groups in the database. 
     
     
         5 . The method of  claim 1 , wherein the correlation score is calculated using normalized ranks of at least some of the features in the contributing feature sets. 
     
     
         6 . The method of  claim 1 , wherein the correlation score is calculated using pre-computed correlation scores indicating the correlation between the contributing feature sets of the category under consideration and contributing feature sets of at least some of the other categories in the database. 
     
     
         7 . The computer implemented method of  claim 1  comprising generating one or more feature sets from raw data from one or more samples, wherein the raw data includes information on one or more features with indications of one or more of: differential expression of said features, abundance of said features, responses of said features to a treatment or stimulus, and effects of said features on biological systems. 
     
     
         8 . The method of  claim 1  further comprising importing the one or more generated feature sets into the database. 
     
     
         9 . (canceled) 
     
     
         10 . A computer program product comprising a non-transitory machine readable medium on which is provided program instructions for correlating chemical or biological categories with other information in a database, said program instructions comprising:
 code for obtaining from the database a taxonomy of biological or chemical categories arranged in a hierarchical structure comprising at least one top-level category having at least one child category;   code for obtaining from the database a plurality of feature sets or feature groups, each feature set comprising two or more features of chemical or biological entities and statistical information associated with each of the two or more features and each feature group comprising a list of inter-related features of chemical or biological entities, wherein at least some features of the feature sets or feature groups have different names and are associated with each other and the feature sets or feature groups are obtained from across different experiments, platforms, or organisms;   code for obtaining from the database a plurality of globally unique mapping identifiers;   code for identifying, for each globally unique mapping identifier, one or more features associated with the globally unique mapping identifier;   code for mapping, for each globally unique mapping identifier, the identified one or more features to the globally unique mapping identifier, wherein at least some features having different names and being associated with each other are mapped to a same globally unique mapping identifier;   code for identifying, for each of a plurality of the categories in the taxonomy, contributing feature sets that contribute to scoring a category under consideration by identifying feature sets among the plurality of feature sets that are associated with the category under consideration or its child categories;   code for using the contributing feature sets of the category under consideration and globally unique mapping identifiers mapped to features of the contributing feature sets to calculate a correlation score indicating a correlation between the category under consideration and a feature, a feature set, or a feature group in the database; and   code for storing the correlation score in the database.   
     
     
         11 . A system for correlating chemical or biological categories with other information in a database, said system comprising:
 a memory;   one or more processors in communication with the memory; and   one or more computer-readable storage media having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the system to:
 obtain from the database a taxonomy of biological or chemical categories arranged in a hierarchical structure comprising at least one top-level category having at least one child category; 
 obtain from the database a plurality of feature sets or feature groups, each feature set comprising two or more features of chemical or biological entities and statistical information associated with each of the two or more features and each feature group comprising a list of inter-related features of chemical or biological entities, wherein at least some features of the feature sets or feature groups have different names and are associated with each other and the feature sets or feature groups are obtained from across different experiments, platforms, or organisms; 
 obtain from the database a plurality of globally unique mapping identifiers; 
 identify, for each globally unique mapping identifier, one or more features associated with the globally unique mapping identifier; 
 map, for each globally unique mapping identifier, the identified one or more features to the globally unique mapping identifier, wherein at least some features having different names and being associated with each other are mapped to a same globally unique mapping identifier; 
 identify, for each of a plurality of the categories in the taxonomy, contributing feature sets that contribute to scoring a category under consideration by identifying feature sets among the plurality of feature sets that are associated with the category under consideration and/or its child categories; 
 use the contributing feature sets of the category under consideration and globally unique mapping identifiers mapped to features of the contributing feature sets to calculate a correlation score indicating a correlation between the category under consideration and a feature, a feature set, or a feature group in the database; and 
 store the correlation score in the database. 
   
     
     
         12 - 36 . (canceled) 
     
     
         37 . The method of  claim 1 , wherein the correlation score indicates a correlation between the category under consideration and a feature set under consideration. 
     
     
         38 . The method of  claim 37 , wherein calculating the correlation score comprises:
 computing contributing correlation scores between the contributing feature sets and the feature set under consideration using associated statistical information of the contributing feature sets and the globally unique mapping identifiers mapped to the features of the contributing feature sets; and   combining the contributing correlation scores to obtain the correlation score indicating a correlation between the category under consideration and the feature set under consideration.   
     
     
         39 . The method of  claim 38 , wherein computing the contributing correlation scores comprises associating features of a contributing feature set with features of the feature set under consideration, wherein two associated features are mapped to a same globally unique mapping identifier. 
     
     
         40 . The method of  claim 1 , wherein the correlation score indicates a correlation between the category under consideration and a feature group under consideration. 
     
     
         41 . The method of  claim 40 , wherein calculating the correlation score comprises:
 computing contributing correlation scores between the contributing feature sets and the feature group under consideration using the globally unique mapping identifiers mapped to the features of the contributing feature sets; and   combining the contributing correlation scores to obtain the correlation score indicating a correlation between the category under consideration and the feature group under consideration.   
     
     
         42 . The method of  claim 41 , wherein computing the contributing correlation scores comprises associating features of a contributing feature set with features of the feature group under consideration, wherein two associated features are mapped to a same globally unique mapping identifier. 
     
     
         43 . The method of  claim 1 , wherein calculating the correlation score comprises:
 obtaining feature scores of all features in the contributing feature sets that are mapped to the globally unique mapping identifier under consideration;   identifying a feature mapped to a globally unique mapping identifier under consideration as the feature under consideration; and   obtaining the correlation score from the feature scores, wherein the correlation score indicates a correlation between the category under consideration and the feature under consideration.   
     
     
         44 . The method of  claim 43 , wherein the feature scores are feature ranks, each feature rank indicating the importance of a feature in an experiment. 
     
     
         45 . The method of  claim 1 , wherein the database comprises a plurality of collections of data stored on a plurality of storage devices. 
     
     
         46 . The method of  claim 2 , wherein the category under consideration comprises a disease, and the removed feature sets are obtained from cell lines that are not affected by the disease.

Join the waitlist — get patent alerts

Track US2016232224A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.