US2020364198A1PendingUtilityA1
Method and system for offline indexing of content and classifying stored data
Est. expiryOct 17, 2026(~0.2 yrs left)· nominal 20-yr term from priority
Inventors:Anand PrahladJeremy A. SchwartzDavid NgoBrian BrockwayMarcus S. MullerParag GokhaleRajiv Kottomtharayil
G06F 16/319G06F 2201/84G06F 16/2372G06F 16/2228G06F 11/1446
70
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A method and system for creating an index of content without interfering with the source of the content includes an offline content indexing system that creates an index of content from an offline copy of data. The system may associate additional properties or tags with data that are not part of traditional indexing of content, such as the time the content was last available or user attributes associated with the content. Users can search the created index to locate content that is no longer available or based on the associate attributes.
Claims
exact text as granted — not AI-modifiedI/We claim:
1 . A non-transitory computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method for indexing content offline, the method comprising:
identifying an offline copy of a production copy of data files created at a production server,
wherein the offline copy includes one or more data files each having keywords and metadata, and
wherein the offline copy is stored in one or more secondary storage devices associated with storage management system that are unavailable to and inaccessible by the production server;
accessing the identified offline copy by an intermediate server, wherein the intermediate server is separate and distinct from the production server; and by the intermediate server:
restoring the offline copy to the intermediate server to analyze content within the offline copy;
identifying content keywords from content within the one or more data files of the offline copy on the intermediate server; and
creating or updating a content index based on the keywords identified from the content within the one or more data files of the offline copy on the intermediate server.
2 . The non-transitory computer-readable storage medium of claim 1 , wherein creating or updating the content index is based on an indexing policy.
3 . The non-transitory computer-readable storage medium of claim 1 , wherein creating or updating the content index comprises at least one of:
determining whether the identified keywords are encrypted; determining a state of data protection of the identified keywords; determining whether the identified keywords have associated access control information; determining at least a portion of a topology of a network in which the identified keywords is stored; and determining whether the identified keywords contain one or more specified text strings or words.
4 . The non-transitory computer-readable storage medium of claim 1 , the method further comprising querying the content index to identify one or more indexed content items that satisfy a received search query.
5 . The non-transitory computer-readable storage medium of claim 1 , the method further comprising selecting a secondary copy of multiple data files from among multiple secondary copies of the multiple data files based on time required to access each of the multiple secondary copies of the multiple data files.
6 . The non-transitory computer-readable storage medium of claim 1 , wherein accessing an offline copy of data that is stored in one or more secondary storage devices of the storage management system includes accessing at least one snapshot of the production copy of data files.
7 . The non-transitory computer-readable storage medium of claim 6 , wherein the at least one snapshot is stored on the one or more secondary storage devices.
8 . The non-transitory computer-readable storage medium of claim 1 , wherein the identifying an offline copy includes identifying a copy to use from among multiple offline copies based on a time required to access each of the multiple offline copies.
9 . The non-transitory computer-readable storage medium of claim 1 , wherein the content index classifies the identified keywords based on at least one or more user-defined classifications that include administratively defined groups within an organization or organization departments associated with the storage management system.
10 . The non-transitory computer-readable storage medium of claim 1 , wherein the intermediate server is distinct from a data server that performed one or more data storage operations to store the offline copy to the one or more secondary storage devices.
11 . The non-transitory computer-readable storage medium of claim 1 , wherein the updating the content index comprises associating the one or more data files of the offline copy with the data files of the production copy.
12 . A system for indexing content offline, the system comprising:
one or more computing devices comprising one or more processors and computer memory, the one or more computing devices configured to:
identify an offline copy of a production copy of data files created at a production server,
wherein the offline copy includes one or more data files each having keywords and metadata, and
wherein the offline copy is stored in one or more secondary storage devices associated with storage management system that are unavailable to and inaccessible by the production server;
access the identified offline copy by an intermediate server, wherein the intermediate server is separate and distinct from the production server; and
the intermediate server configured to:
restore the offline copy to the intermediate server to analyze content within the offline copy;
identify content keywords from content within the one or more data files of the offline copy on the intermediate server; and
create or update a content index based on the keywords identified from the content within the one or more data files of the offline copy on the intermediate server.
13 . The system of claim 12 , wherein the intermediate server creates or updates the content index based on an indexing policy.
14 . The system of claim 12 , wherein the intermediate server is further configured to:
determine whether the identified keywords are encrypted; determine a state of data protection of the identified keywords; determine whether the identified keywords have associated access control information; determine at least a portion of a topology of a network in which the identified keywords is stored; and determine whether the identified keywords contain one or more specified text strings or words.
15 . The system of claim 12 , wherein the one or more computing devices is further configured to:
query the content index to identify one or more indexed content items that satisfy a received search query.
16 . The system of claim 12 , wherein the one or more computing devices is further configured to:
select a secondary copy of multiple data files from among multiple secondary copies of the multiple data files based on time required to access each of the multiple secondary copies of the multiple data files.
17 . The system of claim 12 , wherein accessing an offline copy of data that is stored in one or more secondary storage devices of the storage management system includes accessing at least one snapshot of the production copy of data files.
18 . The system of claim 17 , wherein the at least one snapshot is stored on the one or more secondary storage devices.
19 . The system of claim 12 , wherein the identifying an offline copy includes identifying a copy to use from among multiple offline copies based on a time required to access each of the multiple offline copies.
20 . The system of claim 12 , wherein the one or more computing devices is further configured to:
classify the identified keywords based on at least one or more user-defined classifications that include administratively defined groups within an organization or organization departments associated with the storage management system.
21 . The system of claim 12 , wherein the intermediate server is distinct from a data server that performed one or more data storage operations to store the offline copy to the one or more secondary storage devices.
22 . The system of claim 12 , wherein the one or more computing devices is further configured to:
associate the one or more data files of the offline copy with the data files of the production copy.Join the waitlist — get patent alerts
Track US2020364198A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.