US2022310206A1PendingUtilityA1

Specific Nuclear-Anchored Independent Labeling System

Assignee: UNIV CARNEGIE MELLONPriority: Jun 18, 2019Filed: Jun 18, 2020Published: Sep 29, 2022
Est. expiryJun 18, 2039(~12.9 yrs left)· nominal 20-yr term from priority
G06N 3/047G06N 3/045G06N 3/09G06N 3/094G06N 3/0464G06N 3/0475G16B 30/10G06N 20/10C07K 2319/60C07K 14/47G16B 40/00G06N 3/08C12N 15/86C07K 2319/03C12N 2750/14143G06N 3/0472G06N 3/0454
49
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Materials for generating data representing a synthetic genetic sequence configured for labeling at least one cell type by causing expression of a marker in the at least one cell type are provided herein. Also provided herein are methods and materials for labeling and isolating particular cell types from mixed cell populations are provided herein.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for generating, by a data processing system, data representing a synthetic genetic sequence configured for labeling at least one cell type by causing expression of a marker in the at least one cell type, the method comprising:
 receiving training data associating at least one feature of a genetic sequence of the at least one cell type with expression of a hallmark of the at least one cell type;   training, based on the training data, a model configured to generate data representing a synthetic genetic sequence;   receiving input data comprising the at least one feature; and   generating the synthetic genetic sequence in response to receiving the input data comprising the at least one feature.   
     
     
         2 . The method of  claim 1 , wherein the model comprises a generative adversarial network (GAN). 
     
     
         3 . The method of  claim 2 , wherein the GAN is a deep convolutional GAN (DCGAN). 
     
     
         4 . The method of  claim 2 , wherein the GAN comprises:
 a generator network configured to receive latent random noise data as the input data and generate the synthetic genetic sequence; and   a discriminator network configured to generate a probability value representing whether the input sequence is drawn from the synthetic genetic sequence from the generator network or from a distribution of natural genetic sequences.   
     
     
         5 . The method of  claim 4 , comprising:
 training the generator network and the discriminator network of the GAN by alternating the input data between synthetic sequence derived from the latent noise data input to the generator network and natural genetic sequence data.   
     
     
         6 . The method of  claim 5 , comprising:
 receiving input data representing endogenous genomic sequences underlying previously profiled open chromatin regions of at least two cell types; and updating the discriminator network for distinguishing between the at least two cell types.   
     
     
         7 . The method of  claim 6 , wherein the at least two cell types include parvalbumin positive (PV+) and parvalbumin negative (PV−) neurons; and
 wherein the discriminator network is configured to distinguish between the PV+ and PV− neurons based on the input data. 
 
     
     
         8 . The method of  claim 5 , comprising optimizing a feature of the synthetic genetic sequence by applying a class label to both the generator network and the discriminator network. 
     
     
         9 . The method of  claim 8 , wherein the class label is configured to force a first probability value for a first input type and a second probability value for a second input type that is different from the first input type. 
     
     
         10 . The method of  claim 8 , wherein the feature represents an enhancer of an activity in a cell type. 
     
     
         11 . The method of  claim 1 , wherein the method further comprises:
 generating a nucleic acid sequence comprising the synthetic genetic sequence.   
     
     
         12 . The method of  claim 11 , wherein the nucleic acid sequence comprises the synthetic genetic sequence operably linked to a nucleotide sequence encoding a marker. 
     
     
         13 . The method of  claim 12 , wherein the marker is a tagged Sun1 fusion polypeptide. 
     
     
         14 . The method of  claim 11 , wherein the nucleic acid sequence further comprises a virus nucleic acid sequence. 
     
     
         15 . The method of  claim 14 , wherein the virus is adeno-associated virus (AAV) or lentivirus. 
     
     
         16 . The method of  claim 1 , comprising:
 receiving results data representing a delivery of a nucleic acid including the synthetic genetic sequence to an organism, the results data representing a successful labeling of the at least one cell type or an unsuccessful labeling of the at least one cell type;   updating the model using the results data; and   generating an updated synthetic genetic sequence based on the updated model.   
     
     
         17 . The method of  claim 1 , further comprising
 receiving training data including the hallmark representing marker positive results marker negative results, or both;   extracting at least one feature corresponding to the marker positive result or to the marker negative result; and   adding the at least one feature to the model.   
     
     
         18 . The method of  claim 17 , wherein the at least one feature is a set of k-mer or gapped k-mer counts, and wherein extracting the at least one feature comprises scanning the sequence to determine the set of k-mer counts that form the sequence. 
     
     
         19 . The method of  claim 1 , wherein the model comprises a support vector machine comprising:
 a feature space; and   a support vector representing a classification border in the feature space,   wherein the method comprises adding, to the support vector of the support vector machine, a given set of k-mer counts within a predefined distance of the classification border in the feature space of the support vector machine.   
     
     
         20 . The method of  claim 1 , wherein the model comprises a neural network. 
     
     
         21 . The method of  claim 20 , wherein the neural network is a convolutional neural network. 
     
     
         22 . The method of  claim 20 , wherein the neural network comprises one or more weight values each associated with a feature of the synthetic genetic sequence. 
     
     
         23 . The method of  claim 20 , wherein the feature comprises a set of k-mer counts of the genetic sequence. 
     
     
         24 . The method of  claim 20 , wherein the feature represents a transcription factor binding motif. 
     
     
         25 . The method of  claim 1 , wherein the synthetic genetic sequence is configured to distinguish between either:
 parvalbumin positive (PV+) and parvalbumin negative (PV−) neurons;   PV+ and excitatory (EXC) neurons: or   PV+ and vasoactive intestinal peptide-expressing (VIP+) neurons.

Join the waitlist — get patent alerts

Track US2022310206A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.