US2015026170A1PendingUtilityA1

Representative Document Selection for a Set of Duplicate Documents

Assignee: GOOGLE INCPriority: Jul 3, 2003Filed: Oct 9, 2014Published: Jan 22, 2015
Est. expiryJul 3, 2023(expired)· nominal 20-yr term from priority
G06F 17/30867G06F 17/3053G06F 16/951G06F 16/355Y10S707/99935Y10S707/99932Y10S707/99954Y10S707/99931G06F 16/24578G06F 16/9535G06F 16/9538
59
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods are provided for obtaining a plurality of documents. A respective document in the plurality of documents is associated with a score and each document in the plurality of documents is from a different data structure in a plurality of data structures. Each data structure in the plurality of data structures represents a different portion of a document address space. A first document in the plurality of documents is selected in accordance with the score associated with the first document. The first document has a fingerprint that indicates that the first document has substantially identical content to every other document in the plurality of documents. In accordance with the score, the first document is indexed thereby producing an indexed first document. With respect to the plurality of documents, the indexed first document is included in a document index as representative of each document in the plurality of documents.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method, comprising:
 at a computing device having one or more processors and memory:   obtaining a plurality of documents, wherein a respective document in the plurality of documents is associated with a score and wherein each document in the plurality of documents is from a different data structure in a plurality of data structures, each data structure in the plurality of data structures representing a different portion of a document address space;   selecting a first document in the plurality of documents in accordance with the score associated with the first document, wherein
 the first document has a fingerprint that indicates that the first document has substantially identical content to every other document in the plurality of documents; 
   indexing, in accordance with the score, the first document thereby producing an indexed first document; and   with respect to the plurality of documents, including the indexed first document in a document index as representative of each document in the plurality of documents.   
     
     
         2 . The method of  claim 1 , wherein the score includes a document ranking value indicative of document importance. 
     
     
         3 . The method of  claim 1 , wherein indexing the first document includes:
 identifying a canonical document for the plurality of documents, from one of: (i) the first document, or (ii) a second document in the plurality of documents, by:
 comparing the score associated with the first document to a query score associated with the second document; and 
 selecting the first document as the canonical document when the score associated with the first document is higher than the score associated with the second document by more than a predefined threshold. 
   
     
     
         4 . The method of  claim 3 , wherein the score associated with the second document is the highest among a subset of the plurality of documents, and wherein the subset of the plurality of documents excludes the first document. 
     
     
         5 . The method of  claim 1 , further comprising removing at least one document from the plurality of documents when a total number of documents in the plurality of documents exceeds a predefined value. 
     
     
         6 . The method of  claim 5 , wherein
 each document in the plurality of documents is independently assigned a score, and   the score associated with the at least one document is lower than the scores associated with every other document in the plurality of documents.   
     
     
         7 . The method of  claim 1 , wherein the first document has substantially identical content to the second document in the plurality of documents, when:
 (i) the first document and the second document share the same page content;   (ii) the first document and the second document share the same target network identifier;   (iii) a network identifier of the first document is the same as the target network identifier of the second document; or   (iv) the target network identifier of the first document is the same as the network identifier of the second document.   
     
     
         8 . The method of  claim 1 , further comprising:
 obtaining a second document not included in the plurality of documents; and   in accordance with a determination that the second document has substantially identical content to each document in the plurality of documents, adding the second document to the plurality of documents.   
     
     
         9 . The method of  claim 1 , wherein the fingerprint of a document in the plurality of documents is a function of (i) the content of the document, or (ii) a network address of the document. 
     
     
         10 . The method of  claim 1 , wherein indexing the first document includes:
 identifying a canonical document for the plurality of documents, from one of:   (i) the first document or (ii) a second document in the plurality of documents, wherein the second document in the plurality of documents is presently indexed in the document index as representative of each document in the plurality of documents, by:
 comparing the score associated with the first document to the score associated with the second document, and 
 selecting the first document as the canonical document when (i) the score associated with the first document is higher than the score associated with the second document by more than a predefined arithmetic threshold and (ii) the ratio of the score associated with the first document and the score associated with the second document is greater than a predefined multiplicative threshold. 
   
     
     
         11 . A computing system, comprising:
 one or more processors;   memory; and   one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, the one or more programs comprising instructions for:   selecting a first document in the plurality of documents in accordance with the score associated with the first document, wherein
 the first document has a fingerprint that indicates that the first document has substantially identical content to every other document in the plurality of documents; 
   indexing, in accordance with the score, the first document thereby producing an indexed first document; and   with respect to the plurality of documents, including the indexed first document in a document index as representative of each document in the plurality of documents.   
     
     
         12 . The computing system of  claim 11 , wherein the score includes a document ranking value indicative of document importance. 
     
     
         13 . The computing system of  claim 11 , wherein indexing the first document includes:
 identifying a canonical document for the plurality of documents, from one of: (i) the first document, or (ii) a second document in the plurality of documents, by:
 comparing the score associated with the first document to a query score associated with the second document; and 
 selecting the first document as the canonical document when the score associated with the first document is higher than the score associated with the second document by more than a predefined threshold. 
   
     
     
         14 . The computing system of  claim 13 , wherein the score associated with the second document is the highest among a subset of the plurality of documents, and wherein the subset of the plurality of documents excludes the first document. 
     
     
         15 . The computing system of  claim 11 , further comprising removing at least one document from the plurality of documents when a total number of documents in the plurality of documents exceeds a predefined value. 
     
     
         16 . The computing system of  claim 11 , wherein the first document has substantially identical content to the second document in the plurality of documents, when:
 (i) the first document and the second document share the same page content;   (ii) the first document and the second document share the same target network identifier;   (iii) a network identifier of the first document is the same as the target network identifier of the second document; or   (iv) the target network identifier of the first document is the same as the network identifier of the second document.   
     
     
         17 . The computing system of  claim 11 , further comprising:
 obtaining a second document not included in the plurality of documents; and   in accordance with a determination that the second document has substantially identical content to each document in the plurality of documents, adding the second document to the plurality of documents.   
     
     
         18 . The computing system of  claim 11 , wherein the fingerprint of a document in the plurality of documents is a function of (i) the content of the document, or (ii) a network address of the document. 
     
     
         19 . The computing system of  claim 11 , wherein indexing the first document includes:
 identifying a canonical document for the plurality of documents, from one of: (i) the first document or (ii) a second document in the plurality of documents, wherein the second document in the plurality of documents is presently indexed in the document index as representative of each document in the plurality of documents, by:
 comparing the score associated with the first document to the score associated with the second document, and 
 selecting the first document as the canonical document when (i) the score associated with the first document is higher than the score associated with the second document by more than a predefined arithmetic threshold and (ii) the ratio of the score associated with the first document and the score associated with the second document is greater than a predefined multiplicative threshold. 
   
     
     
         20 . A non-transitory computer readable storage medium storing one or more programs to be executed by one or more processing units of a computing device, the one or more programs comprising instructions that, when executed by the one or more processing units, cause the computing device to:
 obtain a plurality of documents, wherein a respective document in the plurality of documents is associated with a score and wherein each document in the plurality of documents is from a different data structure in a plurality of data structures, each data structure in the plurality of data structures representing a different portion of a document address space;   select a first document in the plurality of documents in accordance with the score associated with the first document, wherein
 the first document has a fingerprint that indicates that the first document has substantially identical content to every other document in the plurality of documents; 
   index, in accordance with the score, the first document thereby producing an indexed first document; and   with respect to the plurality of documents, include the indexed first document in a document index as representative of each document in the plurality of documents.

Join the waitlist — get patent alerts

Track US2015026170A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.