US2024029825A1PendingUtilityA1
System and method for genomic association
Est. expiryMar 8, 2042(~15.6 yrs left)· nominal 20-yr term from priority
G16B 40/30G16B 40/20G16B 20/00G06N 20/00G06N 5/022G06N 3/088G06N 3/0455G06N 3/0464G06N 7/01G06N 20/10
63
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
In variants, a method for genomic association can include: determining observed variable values and observed phenotype values for each organism in a population, removing information from variables of interest, determining a phenotype-variable association model, identifying causal variables associated with a phenotype, and/or any other suitable steps.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A method, comprising:
a) for each of a set of organisms:
determining a trait value for a trait; and
determining variable values for a set of variables;
b) selecting a subset of variables from the set of variables; and c) determining a first model configured to predict values for variables of interest in the set of variables based on values for the subset of variables; d) determining test variable values for the variables of interest using the first model; e) using the test variable values, determining a second model comprising a relationship between the set of variables and the trait; and f) identifying causal variables from the set of variables based on the second model.
2 . The method of claim 1 , further comprising training the first model using values for a second set of variables, different from the subset of variables.
3 . The method of claim 2 , wherein the first model comprises an autoencoder.
4 . The method of claim 1 , wherein determining the first model comprises a fitting a linear regression based on the variable values for the set of variables.
5 . The method of claim 1 , further comprising clustering the set of variables based on autocorrelation analysis of the variable values, wherein the subset of variables and variables of interest are from a shared cluster.
6 . The method of claim 5 , wherein variables in the set of variables comprise k-mers.
7 . The method of claim 1 , wherein (b)-(c) are iteratively repeated until a model fit metric for the first model rises above a threshold.
8 . The method of claim 7 , further comprising, in a first iteration, segmenting the subset of variables into high importance variables and low importance variables based on the first model, wherein, in a second iteration, selecting the subset of variables comprises replacing the low importance variables.
9 . The method of claim 1 , wherein a size of the subset of variables is less than a size of the set of organisms.
10 . The method of claim 1 , further comprising:
determining target values for the causal variables based on a target trait; and based on the target values, breeding organisms in the set of organisms to generate a new organism with the target trait.
11 . The method of claim 1 , wherein variables in the set of variables comprise variables for at least one of: loci, gene expression, protein expression, methylation, environmental parameters, or protein binding affinity.
12 . A method, comprising:
for each of a set of organisms:
determining a trait value for a trait; and
determining variable values associated with the trait value for a set of variables;
transforming values for a subset of variables to a reduced dimension space; determining transformed test values for a variable of interest based on the transformed values for the subset of variables; using the transformed values for the subset of variables and the transformed test values, determining a second model comprising a relationship between transformed variables and the trait; identifying transformed causal variables based on the second model; and decoding the transformed causal variables to determine causal variables.
13 . The method of claim 12 , wherein the reduced dimension space comprises a latent space, wherein transforming values for the subset of variables comprises training an autoencoder to encode variable values to the latent space, wherein the trained autoencoder is used to decode the transformed causal variables.
14 . The method of claim 13 , wherein the autoencoder is trained using values for a second subset of variables, different from the subset of variables.
15 . The method of claim 12 , further comprising determining a distribution of the subset of transformed variables, wherein determining transformed test values for the variable of interest comprises selecting transformed test values from the distribution.
16 . The method of claim 12 , wherein transforming values for a subset of variables comprises:
training a neural network to predict a trait value based on the respective variable values; and using a first layer of the trained neural network to transform the values for the subset of variables.
17 . The method of claim 12 , further comprising selecting organisms from the set of organisms for cross-breeding based on the causal variables and a target trait value.
18 . The method of claim 12 , wherein the set of variables comprise genomic variables, wherein the subset of variables in the reduced dimension space comprises a set of features, wherein the transformed values comprise feature values, wherein the variable of interest comprises a feature of interest, wherein the transformed test values comprise test feature values for the feature of interest, wherein the second model comprises a relationship between the feature of interest, the set of features, and the trait, wherein the transformed causal variables comprise causal features of interest, and wherein decoding the transformed causal variables comprises decoding the causal features of interest into causal genomic variables.Join the waitlist — get patent alerts
Track US2024029825A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.