Methods And Systems For Identifying Progenies For Use In Plant Breeding
Abstract
Exemplary methods for identifying progenies for use in plant breeding are disclosed. One exemplary computer-implemented method includes accessing a data structure including data representative of a pool of progenies and determining a prediction score for at least a portion of the pool of progenies based on the data included in the data structure. The prediction score indicates a probability of selection of the progeny based on historical data. The method further includes selecting a group of progenies from the pool of progenies based on the prediction score, identifying a set of progenies, from the group of progenies, based on at least one of an expected performance of the group of progenies and at least one factor associated with the set of progenies, the pool of progenies and/or the group of progenies, and directing the set of progenies into a validation phase of a breeding pipeline.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for identifying progeny for use in a plant breeding pipeline, the method comprising:
accessing a data structure including phenotypic data representative of a pool of progenies; determining, by at least one computing device, using a prediction model, a prediction score for at least a portion of the pool of progenies based on the phenotypic data in the data structure, the prediction score indicative of a probability of selection of the progeny based on historical selection data for the pool of progenies; selecting, by the at least one computing device, a group of progenies from the pool of progenies based on the prediction score; identifying, by the at least one computing device, a set of progenies from the group of progenies based on an expected performance of the set of progenies, risk of failure of the group of progenies, genetic diversity of the set of progenies, and trait(s) of the set of progenies; and directing the set of progenies to a next phase of the plant breeding pipeline.
2 . The method of claim 1 , further comprising generating, by the at least one computing device, the prediction model based on historical phenotypic data included in the data structure, the historical phenotypic data associated with plant material of a type consistent with a plant type of the pool of progenies; and
wherein the prediction model includes a random forest prediction model.
3 . The method of claim 1 , wherein selecting the group of progenies includes selecting one or more progenies from the pool when the prediction score of the selected progeny satisfies one or more thresholds; and
wherein identifying the set of progenies is further based on probability of success of base origins, base pedigrees, and/or heterotic groups for the group of progenies.
4 . The method of claim 1 , wherein identifying the set of progenies is based on the following set identification algorithm:
x
opt
=
argmax
x
∈
{
0
,
1
}
n
N
(
λ
p
∑
i
=
1
n
N
x
i
p
i
-
λ
r
∑
i
=
1
nN
x
i
r
i
-
λ
d
1
1
T
θ
-
λ
d
2
1
T
φ
-
λ
d
3
1
T
γ
)
;
wherein λ p Σ i=1 nN x i p i is associated with performance of the group of progenies; λ r Σ i=1 nN x i r i is associated with risk; λ d 1 1 T θ, λ d 2 1 T φ, and λ d 3 1 T γ are associated with deviations from one or more performance profiles; and p i , and r i are associated with performance and risk scores, respectively, for the group of progenies.
5 . The method of claim 4 , wherein the set identification algorithm is subject to at least one of the following algorithms:
∑
i
=
1
n
N
X
M
(
i
)
*
x
i
≥
α
M
·
r
;
∑
i
=
1
n
N
X
F
(
i
)
*
x
i
≥
α
F
·
r
;
and
α
T
k
l
(
i
)
≤
∑
j
=
1
N
M
T
k
(
i
,
j
)
*
x
j
≤
α
T
k
h
(
i
)
wherein X F and X M are vectors of female and male gender of the group of progenies; T k is a trait to be included in the group of progenies; matrix M indicates the presence or absence of said trait in the group of progenies; and α T k l (i) and α T k u (i) are lower and upper bounds, respectively.
6 . The method of claim 5 , wherein the set identification algorithm is subject to at least one of the following algorithms:
-
θ
i
≤
∑
j
=
1
n
N
M
l
(
i
,
j
)
*
x
j
-
o
j
≤
θ
i
;
-
φ
k
≤
∑
j
=
1
N
M
o
(
k
,
j
)
(
∑
j
=
1
n
N
M
l
(
i
,
j
)
*
x
j
)
-
b
k
≤
φ
k
;
and
-
γ
i
≤
∑
j
=
1
n
N
M
H
(
i
,
j
)
*
x
j
-
h
j
≤
γ
i
wherein θ i , φ k , γ i include auxiliary variables; o i is a performance profile for origins of said group of progenies; −θ i and θ i are bounds of a deviation defined by o i .
7 . The method of claim 1 , wherein directing the set of progenies to a testing and cultivation phase of a breeding pipeline includes including one or more plants in a growing space of the breeding pipeline, the one or more plants derived from the identified set of progenies.
8 . A system for identifying progeny for use in plant breeding, the system comprising:
a data storage device including phenotypic data related to a pool of progenies, each of the progenies based on one or more origins; and a computing device coupled in communication with the data storage device and configured, by executable instructions, to:
access the phenotypic data in the data structure related to the pool of progenies;
determine a prediction score, using a prediction model, for each of the progenies in the pool of progenies based on the accessed phenotypic data, the prediction score indicative of a probability of selection of the progeny based on historical selection data for the pool of progenies;
select a group of progenies from the pool of progenies based on the prediction score for each of the progenies in the pool of progenies;
identify a set of progenies, from the group of progenies, based on an expected performance of the set of progenies, risk of failure of the group of progenies, genetic diversity of the set of progenies, and trait(s) of the set of progenies; and
direct the set of progenies to a next phase of a breeding pipeline for commercialization.
9 . The system of claim 8 , wherein the computing device is further configured to identify the set of progenies based on the following algorithm:
x
opt
=
argmax
x
∈
{
0
,
1
}
n
N
(
λ
p
∑
i
=
1
n
N
x
i
p
i
-
λ
r
∑
i
=
1
nN
x
i
r
i
-
λ
d
1
1
T
θ
-
λ
d
2
1
T
φ
-
λ
d
3
1
T
γ
)
wherein λ p Σ i=1 nN x i p i is associated with performance of the group of progenies; λ r Σ n=1 nN x i r i is associated with risk; λ d 1 1 T θ, λ d 2 1 T φ, and λ d 3 1 T γ are associated with deviations from one or more performance profiles; and p i , and r i are associated with performance and risk scores, respectively, for the group of progenies.
10 . The system of claim 8 , further comprising the breeding pipeline coupled in communication with the computing device, the breeding pipeline including the next phase;
wherein a plant derived from at least one of the set of progenies is planted in a growing space of the next phase of the breeding pipeline, after the set of progenies are directed to the breeding pipeline.
11 . The system of claim 8 , wherein the computing device is further configured to identify, based on a user input, the pool of progenies, prior to accessing the phenotypic data in the data structure related to the pool of progenies.
12 . The system of claim 8 , wherein the computing device is configured to identify the set of progenies further based on a value associated with the probability of success of the set of progenies less a value associated with a deviation of the set of progenies from a desired profile.
13 . The system of claim 8 , further comprising a growing space of the breeding pipeline; and
wherein the growing space includes, as part of the next phase, at least one plant included in the growing space, the at least one plant derived from at least one of the progenies of the identified set of progenies.
14 . A non-transitory computer readable storage media including executable instructions for identifying progeny for use in plant breeding, which, when executed by at least one processor, cause the at least one processor to:
access a data structure including phenotypic data representative of a pool of progenies; determine, using a prediction model, a prediction score for at least a portion of the pool of progenies based on the phenotypic data in the data structure, the prediction score indicative of a probability of selection of the progeny based on historical selection data for the pool of progenies; select a group of progenies from the pool of progenies based on the prediction score; identify a set of progenies, from the group of progenies, based on: an expected performance of the set of progenies, risk of failure of the group of progenies, genetic diversity of the set of progenies, and trait(s) of the set of progenies; and direct the set of progenies to a next phase of a breeding pipeline.
15 . The non-transitory computer readable storage media of claim 14 , wherein the executable instructions, when executed by the at least one processor, further cause the at least one processor to, prior to determining the prediction score for the at least a portion of the pool of progenies, train the prediction model based on historical phenotypic data included in the data structure, the historical phenotypic data associated with plant material of a type consistent with a plant type of the pool of progenies; and
wherein the prediction model includes one of: a random forest, support vector machine, logistic regression, tree-based, naïve Bayes, linear/logistic regression, deep learning, and/or Gaussian process regression model.
16 . The non-transitory computer readable storage media of claim 15 , wherein the executable instructions, when executed by the at least one processor in connection with selecting the group of progenies, cause the at least one processor to select one or more progenies from the pool in response to the prediction score of the selected progeny satisfies one or more thresholds.
17 . The non-transitory computer readable storage media of claim 15 , wherein the executable instructions, when executed by the at least one processor in connection with identifying the set of progeny, cause the at least one processor to identify the set of progeny based on the following set identification algorithm:
x
opt
=
argmax
x
∈
{
0
,
1
}
n
N
(
λ
p
∑
i
=
1
n
N
x
i
p
i
-
λ
r
∑
i
=
1
n
N
x
i
r
i
-
λ
d
1
1
T
θ
-
λ
d
2
1
T
φ
-
λ
d
3
1
T
γ
)
;
and
wherein the set identification algorithm is subject to at least one of the following algorithms:
∑
i
=
1
n
N
X
M
(
i
)
*
x
i
≥
α
M
·
r
;
∑
i
=
1
n
N
X
F
(
i
)
*
x
i
≥
α
F
·
r
;
α
T
k
l
(
i
)
≤
∑
j
=
1
N
M
T
k
(
i
,
j
)
*
x
j
≤
α
T
k
h
(
i
)
;
-
θ
i
≤
∑
i
=
1
n
N
M
l
(
i
,
j
)
*
x
j
-
o
j
≤
θ
i
;
-
φ
k
≤
∑
j
=
1
N
M
o
(
k
,
j
)
(
∑
j
=
1
n
N
M
l
(
i
,
j
)
*
x
j
)
-
b
k
≤
φ
k
;
and
-
γ
i
≤
∑
j
=
1
n
N
M
H
(
i
,
j
)
*
x
j
-
h
j
≤
γ
i
wherein λ p Σ i=1 nN x i p i is associated with performance of the group of progenies; λ r Σ n=1 nN x i r i is associated with risk; λ d 1 1 T θ, λ d 2 1 T φ, and λ d 3 1 T γ are associated with deviations from one or more performance profiles; and p i , and r i are associated with performance and risk scores, respectively, for the group of progenies;
wherein X F and X M are vectors of female and male gender of the group of progenies; T k is a trait to be included in the group of progenies; matrix M indicates the presence or absence of said trait in the group of progenies; and α T k l (i) and α T k u (i) are lower and upper bounds, respectively; and
wherein θ i , φ k , γ i include auxiliary variables; o i is a performance profile for origins of said group of progenies; −θ i and θ i are bounds of a deviation defined by o i .
18 . The non-transitory computer readable storage media of claim 15 , wherein the executable instructions, when executed by the at least one processor, cause the at least one processor to identify the set of progenies further based on one or more of: a deviation of the set of progenies from at least one profile, risk, probability of success of base origins, probability of success of base pedigrees, probability of success of heterotic groups, one or more trait profiles, market segmentation, production cost, and/or trait integration.Join the waitlist — get patent alerts
Track US2023386609A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.