US2018081992A1PendingUtilityA1
Determination of relationships between collections of disparate media types
Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: Jan 28, 2012Filed: Nov 30, 2017Published: Mar 22, 2018
Est. expiryJan 28, 2032(~5.5 yrs left)· nominal 20-yr term from priority
G06F 17/30979G06F 16/90335
45
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Architecture that automatically determines relationships between vector spaces of disparate media types, and outputs ranker signals based on these relationships, all in a single process. The architecture improves search result relevance by simultaneously clustering queries and documents, and enables the training of a model for creating one or more ranker signals using simultaneous clustering of queries and documents in their respective spaces.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system, comprising:
a relationship component that automatically determines relationships between document collections of respective, different file types by
applying a cost function that is based on truth data to measure one or more similarities between concurrently created document clusters and query clusters, the one or more similarities defining the relationships, the collections and the relationships being usable in a query processing to return documents,
computing a total weight of each query in a query collection, and
computing a probability that the query belongs in a given query collection; and
a processor that executes computer-executable instructions associated with at least the relationship component.
2 . The system of claim 1 , wherein the relationships are automatically determined during an offline training phase and the query processing occurs online.
3 . The system of claim 1 , wherein the query does not share a media type with the document collections.
4 . The system of claim 1 , wherein the relationship component:
computes vectors for the document and query collections by application of a collection-wise probabilistic algorithm; and creates a combined model of query-document labeled data by application of a similarity function to the vectors.
5 . The system of claim 4 , wherein the file types include at least two of text, audio, video, and image.
6 . The system of claim 1 , wherein the relationship component computes vectors for the respective collections, and wherein the vectors are generated based on a collection-wise probabilistic algorithm and processed with a similarity function to yield a combined model of query-document labeled data.
7 . The system of claim 1 , further comprising a signaler that generates one or more signals, based on the computed relationships, for an online processing that orders search results.
8 . A system that determines relationships between respective media collections that are each of a different media type, comprising:
a relationship component that, during an offline training phase, determines relationships between the respective media collections, the relationships being usable in subsequent online processing of a query, by
applying a cost function to identify one or more similarities between observed true labels and predicted labels as the relationships;
computing a total weight of each query in a query collection, and
computing a probability that the query belongs in a given query collection; and
a processor that executes computer-executable instructions associated with the relationship component.
9 . The system of claim 8 , wherein the collections and relationships are computed in a single process.
10 . The system of claim 8 , wherein the one or more similarities between expected true labels is based on a comparison of pairs of media collection members.
11 . The system of claim 8 , wherein the collections are clusters that are concurrently created as query clusters and media clusters, and the relationship component computes the relationships between the query clusters and media clusters.
12 . The system of claim 8 , wherein media types include at least two of text, audio, video, and image.
13 . A method, comprising:
obtaining vector collections by
converting a query of a first media type that returns documents of a different media type into collections of multi-dimensional query vectors, and
converting the documents into collections of multi-dimensional document vectors;
computing relationships between the vectors of the respective collections, based on vector probabilities, by applying a cost function that measures similarity between expected true labels and predicted labels, the relationships representing relevance of the query to a one of the documents; and utilizing a processor that executes instructions stored in memory to perform at least one of the acts of processing, converting, computing relationships, or computing a probability.
14 . The method of claim 13 , wherein the method is performed during an offline training phase.
15 . The method of claim 13 , wherein the converting operations are executed together.
16 . The method of claim 13 , further comprising creating a combined model that is usable by a similarity function, based on query and document labeled data.
17 . The method of claim 13 , wherein the cost function is cross-entropy based.
18 . The method of claim 13 , wherein the vectors of the respective collections are generated based on a collection-wise probabilistic algorithm.
19 . The method of claim 18 , further comprising obtaining a combined model of query-document labeled data by applying a similarity function to the computed query vectors and the computed document vectors.
20 . The method of claim 13 , wherein the document types include at least two of text, audio, video, and image.Join the waitlist — get patent alerts
Track US2018081992A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.