US2018081992A1PendingUtilityA1

Determination of relationships between collections of disparate media types

Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: Jan 28, 2012Filed: Nov 30, 2017Published: Mar 22, 2018
Est. expiryJan 28, 2032(~5.5 yrs left)· nominal 20-yr term from priority
G06F 17/30979G06F 16/90335
45
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Architecture that automatically determines relationships between vector spaces of disparate media types, and outputs ranker signals based on these relationships, all in a single process. The architecture improves search result relevance by simultaneously clustering queries and documents, and enables the training of a model for creating one or more ranker signals using simultaneous clustering of queries and documents in their respective spaces.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system, comprising:
 a relationship component that automatically determines relationships between document collections of respective, different file types by
 applying a cost function that is based on truth data to measure one or more similarities between concurrently created document clusters and query clusters, the one or more similarities defining the relationships, the collections and the relationships being usable in a query processing to return documents, 
 computing a total weight of each query in a query collection, and 
 computing a probability that the query belongs in a given query collection; and 
   a processor that executes computer-executable instructions associated with at least the relationship component.   
     
     
         2 . The system of  claim 1 , wherein the relationships are automatically determined during an offline training phase and the query processing occurs online. 
     
     
         3 . The system of  claim 1 , wherein the query does not share a media type with the document collections. 
     
     
         4 . The system of  claim 1 , wherein the relationship component:
 computes vectors for the document and query collections by application of a collection-wise probabilistic algorithm; and   creates a combined model of query-document labeled data by application of a similarity function to the vectors.   
     
     
         5 . The system of  claim 4 , wherein the file types include at least two of text, audio, video, and image. 
     
     
         6 . The system of  claim 1 , wherein the relationship component computes vectors for the respective collections, and wherein the vectors are generated based on a collection-wise probabilistic algorithm and processed with a similarity function to yield a combined model of query-document labeled data. 
     
     
         7 . The system of  claim 1 , further comprising a signaler that generates one or more signals, based on the computed relationships, for an online processing that orders search results. 
     
     
         8 . A system that determines relationships between respective media collections that are each of a different media type, comprising:
 a relationship component that, during an offline training phase, determines relationships between the respective media collections, the relationships being usable in subsequent online processing of a query, by
 applying a cost function to identify one or more similarities between observed true labels and predicted labels as the relationships; 
 computing a total weight of each query in a query collection, and 
 computing a probability that the query belongs in a given query collection; and 
   a processor that executes computer-executable instructions associated with the relationship component.   
     
     
         9 . The system of  claim 8 , wherein the collections and relationships are computed in a single process. 
     
     
         10 . The system of  claim 8 , wherein the one or more similarities between expected true labels is based on a comparison of pairs of media collection members. 
     
     
         11 . The system of  claim 8 , wherein the collections are clusters that are concurrently created as query clusters and media clusters, and the relationship component computes the relationships between the query clusters and media clusters. 
     
     
         12 . The system of  claim 8 , wherein media types include at least two of text, audio, video, and image. 
     
     
         13 . A method, comprising:
 obtaining vector collections by
 converting a query of a first media type that returns documents of a different media type into collections of multi-dimensional query vectors, and 
 converting the documents into collections of multi-dimensional document vectors; 
   computing relationships between the vectors of the respective collections, based on vector probabilities, by applying a cost function that measures similarity between expected true labels and predicted labels, the relationships representing relevance of the query to a one of the documents; and   utilizing a processor that executes instructions stored in memory to perform at least one of the acts of processing, converting, computing relationships, or computing a probability.   
     
     
         14 . The method of  claim 13 , wherein the method is performed during an offline training phase. 
     
     
         15 . The method of  claim 13 , wherein the converting operations are executed together. 
     
     
         16 . The method of  claim 13 , further comprising creating a combined model that is usable by a similarity function, based on query and document labeled data. 
     
     
         17 . The method of  claim 13 , wherein the cost function is cross-entropy based. 
     
     
         18 . The method of  claim 13 , wherein the vectors of the respective collections are generated based on a collection-wise probabilistic algorithm. 
     
     
         19 . The method of  claim 18 , further comprising obtaining a combined model of query-document labeled data by applying a similarity function to the computed query vectors and the computed document vectors. 
     
     
         20 . The method of  claim 13 , wherein the document types include at least two of text, audio, video, and image.

Join the waitlist — get patent alerts

Track US2018081992A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.