US2012143841A1PendingUtilityA1

Methods and apparatuses for searching content

Individually held — no corporate assignee on recordPriority: Jun 12, 2006Filed: Feb 10, 2012Published: Jun 7, 2012
Est. expiryJun 12, 2026(expired)· nominal 20-yr term from priority
G06Q 30/0256G06F 16/951
57
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Embodiments of methods and apparatuses for searching contents, including structured search are described herein. Embodiments of the present invention use tree structures (or more generally, graph structures), layout structures, and/or content category information to capture within search results relevant content that would otherwise be missed, to reduce the incidence of false positives within search results, and to improve the accuracy of rankings within search results. Embodiments of the present invention further use tree structures (or more generally, graph structures), layout structures, and/or content category information to extend search results to include sub-document constituents. Embodiments of the present invention also support the use of distribution properties as criteria for ranking search results. And embodiments of the present invention support search based on structural proximity, search expressions with recursively embedded operators, predicates, and/or quantifiers, and applications to selection of advertisements.

Claims

exact text as granted — not AI-modified
1 . A machine implemented method comprising:
 receiving by a search engine, from a content searching or consuming application, an atomic search term, the search engine and the content searching or consuming application being operated on one or more different or same computing devices;   generating in response, by the search engine, one or more scores indicative of relative relevance of a content or one or more portions of the content to the atomic search term, the generating by the search engine being based at least in part on a structure, a distance function, and a scoring function, the structure structurally describing the content having content nodes and/or text strings, the distance function measuring distances between sub-structures within the structure, and the scoring function being positionally sensitive, yielding different scores for different occurrence positions of the atomic search term in the content; and   conditionally providing or not providing the content or one or more portions of the content, or access information of the content or one or more portions of the content, to the content searching or consuming application, by the search engine, based at least in part on the generated one or more scores,   wherein the generating of one or more scores includes establishing a bound on a number of children content nodes for each content node and/or a bound on a size of each of the text strings.   
     
     
         2 . The method of  claim 1 , wherein the atomic search term comprises a plurality of words. 
     
     
         3 . The method of  claim 1 , wherein the structure comprises one or more strings of words, one or more markup strings, one or more trees corresponding to parsed markup, one or more deduced semantic trees, one or more database records or one or more database objects. 
     
     
         4 . The method of  claim 1 , wherein the content comprises one or more web pages of one or more web applications, one or more XML documents in one or more XML repositories, one or more documents in one or more document corpora, or one or more database objects in one or more databases. 
     
     
         5 . The method of  claim 1 , wherein the structure comprises a tree structure corresponding to parsed markup of the content, annotated with measurement information derived from layout structures associated with the content. 
     
     
         6 . The method of  claim 5 , further comprising deriving the measurement information and annotating the tree structure. 
     
     
         7 . The method of  claim 1  wherein the content comprises a plurality of constituents, and the method further comprises building the structure by recursively forming higher sub-structures from lower sub-structures of the constituents of the content. 
     
     
         8 . The method of  claim 1 , wherein the content comprises a plurality of constituents, and the generating of one or more scores comprises generating said scores for one or more atomic ones of the constituents, one or more aggregate ones of the atomic constituents, one or more aggregate ones of the aggregates, or one or more aggregate ones of the aggregates and atomic constituents. 
     
     
         9 . The method of  claim 8 , wherein the generating of a score for an aggregate comprises calculating an overall score for the aggregate as a match for the atomic search term by calculating c 1 *D+c 2 *Δ+c 3 *ρ, where D is a density of the atomic search term on the aggregate, Δ is a distribution score for the atomic search expression on the aggregate, ρ is the r-value for the atomic search expression on the aggregate, and c 1 , c 2 , and c 3  are non-negative real numbers such that c 1 +c 2 +c 3 ≦1, wherein (Σ 1≦i≦m (c i *P i   e     i   )) * Π m+1≦i≦n  P i   e     i    provides the overall score, P 1 , . . . , P m  being beneficial properties and P m+i , . . . , P n  being detrimental properties. 
     
     
         10 . The method of  claim 9 , wherein the generating further comprises calculating either D, Δ or both, based at least in part on relevance values assigned to children of the aggregate.

Join the waitlist — get patent alerts

Track US2012143841A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.