Reducing dimensionality of multimodal embeddings in deep learning models
Abstract
Apparatus and method for reducing the dimensionality of embeddings in a machine learning (ML) system. In some embodiments, original embeddings from a pre-trained deep learning model are extracted for a set of data, the original embeddings having an initial dimensionality. A set of projection vectors are initialized with a specified dimensionality smaller than the initial dimensionality. A neural network is used to optimize the projection vectors by minimizing a loss function based on a similarity matrix. The set of original embeddings are thereafter projected onto the optimized projection vectors to obtain a set of reduced-dimensional embeddings with the specified dimensionality. A neural network of an ML system is thereafter configured using the reduced-dimensional embeddings, such as by a training operation to duplicate operation of the deep learning model in a smaller latent space. The original embeddings may be single modal or multimodal embeddings.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
in a computer processor having program instructions stored in a memory,
configuring the computer processor to carry out the following operations:
initializing projection vectors of a specified dimensionality in the memory;
extracting a set of original embeddings from a pre-trained deep learning model realized by the computer processor for a set of data having at least one modality, the set of original embeddings stored in the memory and having an initial dimensionality larger than the specified dimensionality;
optimizing the projection vectors by minimizing a loss function of a neural network realized by the computer processor in relation to an original similarity matrix and a reduced similarity matrix to generate optimized projection vectors; and
projecting the set of original embeddings onto the optimized projection vectors to obtain a set of reduced-dimensional embeddings having an overall informational content that corresponds to an overall informational content of the set of original embeddings; and
using the reduced-dimensional embeddings to configure a second neural network of a machine learning (ML).
2 . The method of claim 1 , further comprising evaluating performance of the reduced-dimensional embeddings in the ML system using retrieval metrics associated with the set of data.
3 . The method of claim 1 , further comprising applying a mask to each of the original similarity matrix and the reduced similarity matrix during the optimizing of the projection vectors.
4 . The method of claim 1 , wherein the second neural network of the ML system is trained on the reduced-dimensional embeddings.
5 . The method of claim 4 , wherein the trained second neural network duplicates operation of the pre-trained deep-learning model in a reduced latency space.
6 . The method of claim 1 , wherein the projection vectors are optimized using a gradient descent algorithm with early stopping based on evaluation loss.
7 . The method of claim 1 , wherein the accessing step comprises using the reduced-dimensional embeddings in a retrieval task to identify and retrieve data responsive to a search query.
8 . The method of claim 1 , wherein the specified dimensionality is at least 25% smaller than the initial dimensionality.
9 . The method of claim 1 , wherein the specified dimensionality is at least 50% smaller than the initial dimensionality.
10 . The method of claim 1 , wherein the initial dimensionality is 1024 dimensions and the specified dimensionality is 768 dimensions or less.
11 . The method of claim 1 , wherein the computer processor is a server processor or a GPU processor, and the using step comprises transferring the reduced-dimensional embeddings across a network to a memory accessible by a micro-controller of the ML system at a location remote from the computer processor.
12 . The method of claim 1 , wherein the original similarity matrix is constructed using the set of original embeddings, the projected similarity matrix is constructed using the set of reduced-dimensional embeddings, and the optimizing step conforms the projected similarity matrix to the original similarity matrix while preserving relationships encoded in the set of original embeddings.
13 . An apparatus, comprising:
a computer system comprising at least one computer processor having program instructions stored in a memory, the at least one computer processor configured to::
initialize projection vectors of a specified dimensionality in the memory;
extract a set of original embeddings from a pre-trained deep learning model realized by the computer processor for a set of data having at least one modality, the set of original embeddings stored in the memory and having an initial dimensionality larger than the specified dimensionality;
optimize the projection vectors by minimizing a loss function of a neural network realized by the computer processor in relation to an original similarity matrix and a reduced similarity matrix to generate optimized projection vectors;
project the set of original embeddings onto the optimized projection vectors to obtain a set of reduced-dimensional embeddings having an overall informational content that corresponds to an overall informational content of the set of original embeddings; and
configure a neural network of a machine learning (ML) system with the set of reduced-dimensional embeddings.
14 . The apparatus of claim 13 , wherein the at least one processor is further configured to evaluate performance of the reduced-dimensional embeddings using retrieval metrics.
15 . The apparatus of claim 13 , wherein the at least one processor further operates to apply masking to the original similarity matrix and the projected similarity matrix during the optimization of the projection vectors.
16 . The apparatus of claim 13 , wherein the projection vectors are optimized using a gradient descent algorithm with early stopping based on evaluation loss.
17 . The apparatus of claim 13 , wherein the ML system uses the reduced-dimensional embeddings in a retrieval task to identify and retrieve data responsive to a search query.
18 . The apparatus of claim 13 , wherein the computer processor is a server processor or a GPU processor, and the accessing step comprises installing the reduced-dimensional embeddings in a micro-controller of the ML system at a location remote from the computer processor.
19 . The apparatus of claim 13 , wherein the original similarity matrix is constructed using the set of original embeddings, the projected similarity matrix is constructed using the set of reduced-dimensional embeddings, and computer processor conforms the projected similarity matrix to the original similarity matrix while preserving relationships encoded in the set of original embeddings.
20 . The apparatus of claim 13 , wherein the second neural network of the ML system is trained on the reduced-dimensional embeddings to duplicate operation of the pre-trained deep-learning model in a reduced latency space.
21 . The apparatus of claim 13 , wherein the specified dimensionality is at least 25% smaller than the initial dimensionality.
22 . The apparatus of claim 13 , wherein the specified dimensionality is at least 50% smaller than the initial dimensionality.Join the waitlist — get patent alerts
Track US2026050788A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.