US2020364198A1PendingUtilityA1

Method and system for offline indexing of content and classifying stored data

Assignee: COMMVAULT SYSTEMS INCPriority: Oct 17, 2006Filed: Aug 6, 2020Published: Nov 19, 2020
Est. expiryOct 17, 2026(~0.2 yrs left)· nominal 20-yr term from priority
G06F 16/319G06F 2201/84G06F 16/2372G06F 16/2228G06F 11/1446
70
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method and system for creating an index of content without interfering with the source of the content includes an offline content indexing system that creates an index of content from an offline copy of data. The system may associate additional properties or tags with data that are not part of traditional indexing of content, such as the time the content was last available or user attributes associated with the content. Users can search the created index to locate content that is no longer available or based on the associate attributes.

Claims

exact text as granted — not AI-modified
I/We claim: 
     
         1 . A non-transitory computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method for indexing content offline, the method comprising:
 identifying an offline copy of a production copy of data files created at a production server,
 wherein the offline copy includes one or more data files each having keywords and metadata, and 
 wherein the offline copy is stored in one or more secondary storage devices associated with storage management system that are unavailable to and inaccessible by the production server; 
   accessing the identified offline copy by an intermediate server, wherein the intermediate server is separate and distinct from the production server; and   by the intermediate server:
 restoring the offline copy to the intermediate server to analyze content within the offline copy; 
 identifying content keywords from content within the one or more data files of the offline copy on the intermediate server; and 
 creating or updating a content index based on the keywords identified from the content within the one or more data files of the offline copy on the intermediate server. 
   
     
     
         2 . The non-transitory computer-readable storage medium of  claim 1 , wherein creating or updating the content index is based on an indexing policy. 
     
     
         3 . The non-transitory computer-readable storage medium of  claim 1 , wherein creating or updating the content index comprises at least one of:
 determining whether the identified keywords are encrypted;   determining a state of data protection of the identified keywords;   determining whether the identified keywords have associated access control information;   determining at least a portion of a topology of a network in which the identified keywords is stored; and   determining whether the identified keywords contain one or more specified text strings or words.   
     
     
         4 . The non-transitory computer-readable storage medium of  claim 1 , the method further comprising querying the content index to identify one or more indexed content items that satisfy a received search query. 
     
     
         5 . The non-transitory computer-readable storage medium of  claim 1 , the method further comprising selecting a secondary copy of multiple data files from among multiple secondary copies of the multiple data files based on time required to access each of the multiple secondary copies of the multiple data files. 
     
     
         6 . The non-transitory computer-readable storage medium of  claim 1 , wherein accessing an offline copy of data that is stored in one or more secondary storage devices of the storage management system includes accessing at least one snapshot of the production copy of data files. 
     
     
         7 . The non-transitory computer-readable storage medium of  claim 6 , wherein the at least one snapshot is stored on the one or more secondary storage devices. 
     
     
         8 . The non-transitory computer-readable storage medium of  claim 1 , wherein the identifying an offline copy includes identifying a copy to use from among multiple offline copies based on a time required to access each of the multiple offline copies. 
     
     
         9 . The non-transitory computer-readable storage medium of  claim 1 , wherein the content index classifies the identified keywords based on at least one or more user-defined classifications that include administratively defined groups within an organization or organization departments associated with the storage management system. 
     
     
         10 . The non-transitory computer-readable storage medium of  claim 1 , wherein the intermediate server is distinct from a data server that performed one or more data storage operations to store the offline copy to the one or more secondary storage devices. 
     
     
         11 . The non-transitory computer-readable storage medium of  claim 1 , wherein the updating the content index comprises associating the one or more data files of the offline copy with the data files of the production copy. 
     
     
         12 . A system for indexing content offline, the system comprising:
 one or more computing devices comprising one or more processors and computer memory, the one or more computing devices configured to:
 identify an offline copy of a production copy of data files created at a production server,
 wherein the offline copy includes one or more data files each having keywords and metadata, and 
 wherein the offline copy is stored in one or more secondary storage devices associated with storage management system that are unavailable to and inaccessible by the production server; 
 
 access the identified offline copy by an intermediate server, wherein the intermediate server is separate and distinct from the production server; and 
   the intermediate server configured to:
 restore the offline copy to the intermediate server to analyze content within the offline copy; 
 identify content keywords from content within the one or more data files of the offline copy on the intermediate server; and 
 create or update a content index based on the keywords identified from the content within the one or more data files of the offline copy on the intermediate server. 
   
     
     
         13 . The system of  claim 12 , wherein the intermediate server creates or updates the content index based on an indexing policy. 
     
     
         14 . The system of  claim 12 , wherein the intermediate server is further configured to:
 determine whether the identified keywords are encrypted;   determine a state of data protection of the identified keywords;   determine whether the identified keywords have associated access control information;   determine at least a portion of a topology of a network in which the identified keywords is stored; and   determine whether the identified keywords contain one or more specified text strings or words.   
     
     
         15 . The system of  claim 12 , wherein the one or more computing devices is further configured to:
 query the content index to identify one or more indexed content items that satisfy a received search query.   
     
     
         16 . The system of  claim 12 , wherein the one or more computing devices is further configured to:
 select a secondary copy of multiple data files from among multiple secondary copies of the multiple data files based on time required to access each of the multiple secondary copies of the multiple data files.   
     
     
         17 . The system of  claim 12 , wherein accessing an offline copy of data that is stored in one or more secondary storage devices of the storage management system includes accessing at least one snapshot of the production copy of data files. 
     
     
         18 . The system of  claim 17 , wherein the at least one snapshot is stored on the one or more secondary storage devices. 
     
     
         19 . The system of  claim 12 , wherein the identifying an offline copy includes identifying a copy to use from among multiple offline copies based on a time required to access each of the multiple offline copies. 
     
     
         20 . The system of  claim 12 , wherein the one or more computing devices is further configured to:
 classify the identified keywords based on at least one or more user-defined classifications that include administratively defined groups within an organization or organization departments associated with the storage management system.   
     
     
         21 . The system of  claim 12 , wherein the intermediate server is distinct from a data server that performed one or more data storage operations to store the offline copy to the one or more secondary storage devices. 
     
     
         22 . The system of  claim 12 , wherein the one or more computing devices is further configured to:
 associate the one or more data files of the offline copy with the data files of the production copy.

Join the waitlist — get patent alerts

Track US2020364198A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.