Dataset relevance estimation in storage systems
Abstract
The invention is notably directed to computer-implemented methods and systems for managing datasets in a storage system. In such systems, it is assumed that a (typically small) subset of datasets are labeled with respect to their relevance, so as to be associated with respective relevance values. Essentially, the present methods determine, for each unlabeled dataset of the datasets, a respective probability distribution over a set of relevance values. From this probability distribution, a corresponding relevance value can be obtained. This probability distribution is computed based on distances (or similarities), in terms of metadata values, between said each unlabeled dataset and the labeled datasets. Based on their associated relevance values, datasets can then be efficiently managed in a storage system.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method for managing datasets in a storage system, wherein a subset of the datasets is labeled with respect to their relevance, so as to be associated with respective relevance values, the method comprising:
defining supersets of datasets, based on metadata available for the datasets; associating each of the datasets to at least one of the defined supersets by comparing metadata available for each of the defined datasets with metadata used to define the at least one of the defined supersets; defining a heterogeneous bipartite factor graph having two types of nodes, wherein the datasets and the defined supersets are associated with a first type of nodes and a second type of nodes of the defined heterogeneous bipartite factor graph, respectively; associating each unlabeled dataset to at least one of the defined supersets by connecting each unlabeled dataset to at least one of the second type of nodes of the defined heterogeneous bipartite factor graph; determining, for each unlabeled dataset of the datasets, a respective probability distribution over a set of relevance values, to obtain a corresponding relevance value, wherein the respective probability distribution is computed by a message passing algorithm on the defined heterogeneous bipartite factor graph, whereby probability distributions are passed as messages along edges of the defined heterogeneous bipartite factor graph that connect pairs of nodes of different types in the defined heterogeneous bipartite factor graph; and managing the datasets in the storage system based on their associated relevance values, wherein managing the datasets comprises storing the datasets across storage tiers of the storage system based on their relevance values.Join the waitlist — get patent alerts
Track US2019243546A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.