US2012290597A1PendingUtilityA1

Detecting Duplicate and Near-Duplicate Files

Individually held — no corporate assignee on recordPriority: Aug 4, 2006Filed: Sep 2, 2011Published: Nov 15, 2012
Est. expiryAug 4, 2026(~0 yrs left)· nominal 20-yr term from priority
G06F 40/194G06F 16/958
47
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Near-duplicate documents may be identified by (a) accepting a set of documents, (b) processing the set of documents to determine a first set of near-duplicate documents using a first document similarity technique, and (c) processing the first set of near duplicate documents to determine a second set of near-duplicate documents using a second document similarity technique. The first document similarity technique might be token order dependent, and the second document similarity technique might be order independent. The first document similarity technique might be token frequency independent, and the second document similarity technique might be frequency dependent. The first document similarity technique might determine whether two documents are near-duplicates using representations based on a subset of the words or tokens of the documents, and the second document similarity technique might determine whether two documents are near-duplicates using representations based on all of the words or tokens of the documents. The first document similarity technique might use set intersection to determine whether or not documents are near-duplicates, and the second document similarity technique might use random projections to determine whether or not documents are near-duplicates.

Claims

exact text as granted — not AI-modified
1 - 29 . (canceled) 
     
     
         30 . A computer-implemented method comprising:
 accepting a set of documents; and   processing the set of documents to determine near-duplicate documents, the processing including, for one or more pairs of documents in the set of documents:
 determining whether the documents in each pair of documents are from the same Website; and 
 determining whether each pair of documents are near-duplicate documents using a first document similarity technique when the pair of documents is from the same Website, and determining whether each pair of documents are near-duplicate documents using a second document similarity technique when the pair of documents is from different Websites, where the first document similarity technique and the second document similarity technique are different document similarity techniques. 
   
     
     
         31 . The method of  claim 30 , wherein the first document similarity technique is token order dependent and token frequency independent. 
     
     
         32 . The method of  claim 30 , wherein the second document similarity technique is token order independent and token frequency dependent. 
     
     
         33 . A system comprising:
 one or more computers configured to perform operations comprising:
 accepting a set of documents; and 
 processing the set of documents to determine near-duplicate documents, the
 determining whether the documents in each pair of documents are from the same Website; and 
 determining whether each pair of documents are near-duplicate documents using a first document similarity technique when the pair of documents are from the same Website, and determining whether each pair of documents are near-duplicate documents using a second document similarity technique when the pair of documents are from different Websites, where the first document similarity technique and the second document similarity technique are different document similarity techniques. 
 
   
     
     
         34 . The system of  claim 33 , wherein the first document similarity technique is token order dependent and token frequency independent. 
     
     
         35 . The system of  claim 33 , wherein the second document similarity technique is token order independent and token frequency dependent. 
     
     
         36 . A machine readable medium having stored thereon machine-executable instructions which, when executed by a machine, cause the machine to perform operations comprising:
 accepting a set of documents; and   processing the set of documents to determine near-duplicate documents, the processing including, for one or more pairs of documents in the set of documents:
 determining whether the documents in each pair of documents are from the same Website; and 
 determining whether each pair of documents are near-duplicate documents using a first document similarity technique when the pair of documents are from the same Website, and determining whether each pair of documents are near-duplicate documents using a second document similarity technique when the pair of documents are from different Websites, where the first document similarity technique and the second document similarity technique are different document similarity techniques. 
   
     
     
         37 . The machine readable medium of  claim 36 , wherein the first document similarity technique is token order dependent and token frequency independent. 
     
     
         38 . The machine readable medium of  claim 36 , wherein the second document similarity technique is token order independent and token frequency dependent.

Join the waitlist — get patent alerts

Track US2012290597A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.