US2018329932A1PendingUtilityA1

Identification of distinguishing compound features extracted from real time data streams

Assignee: CA INCPriority: Mar 25, 2016Filed: May 16, 2018Published: Nov 15, 2018
Est. expiryMar 25, 2036(~9.7 yrs left)· nominal 20-yr term from priority
G06F 17/3053G06F 2221/034G06F 21/56G06F 21/552G06F 17/30303G06F 17/30539G06F 16/2465G06F 16/215G06F 16/24578
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A big data processing system includes a features permutations testing function that separates out from among a set of identified compound features, those compound feature permutations that have better capabilities for distinguishing between anomalies observed in respective multi-dimensional feature spaces having as their axes the features of the identified compound features.

Claims

exact text as granted — not AI-modified
1 .- 21 . (canceled) 
     
     
         22 . A machine-implemented method of presenting hierarchically organized alert information for meaningful comprehension by human administrators, the method comprising:
 automatically ranking compound feature anomaly detections extracted in a de-duplicating manner from normalcy-violating log reports of a logs generating data processing system, the ranking being according to how the de-duplicated compound feature anomaly detections place in an informational entropy space relative to predetermined single feature anomaly detections that serve as benchmarks in the informational entropy space, the compound feature anomaly detections being comprised of respective sets of interrelated single features determined to correlate to respective normalcy-violations within a predetermined time window; and   presenting one or more alert reports in a dashboard such that for each presented alert report, the highest ranked and de-duplicated set of compound features related to that alert report is presented ahead of a next highest ranked and de-duplicated set of compound features related to that alert report.   
     
     
         23 . The method of  claim 22  wherein:
 each presented de-duplicated set of compound features is unfurled or unfurl-able to present an identification of its more relevant features among its respective set of compound features. 
 
     
     
         24 . The method of  claim 23  wherein at least one of the more relevant features is selected from the group consisting of:
 a common set of certain IP addresses from which requests are sourced; 
 a common set of certain HTTP request codes used in requests; 
 a common set of certain geographic addresses from which requests are sourced; 
 a common set of certain user identifications on behalf of which requests are sourced; 
 an above threshold number of requests sourced within the predetermined time window from a common set of certain IP addresses or certain geographic addresses; 
 above threshold response times for a common set of sourced within the predetermined time window; 
 above threshold congestion on a communications link within the predetermined time window; and 
 above threshold congestion in a requests servicing server within the predetermined time window. 
 
     
     
         25 . The method of  claim 23  wherein one or more of the unfurled or unfurl-able identifications of a more relevant feature is expandable to present more detailed information about that identified more relevant feature. 
     
     
         26 . The method of  claim 22  wherein:
 each presented alert report is expandable to show the normalcy-violating log reports on which the presented alert report is based. 
 
     
     
         27 . The method of  claim 22  wherein:
 each presented alert report is expandable to show one or more diagnostic visualizations useful for arriving at a diagnosis of what the probable underlying causes are for the normalcy-violating log reports on which the presented alert report is based. 
 
     
     
         28 . The method of  claim 27  wherein:
 the diagnostic visualizations include a diagnostic insight derived from a knowledge base of probable underlying causes maintained by the data processing system. 
 
     
     
         29 . The method of  claim 28  wherein the knowledge base of probable underlying causes maintained by the data processing system contains a set of fault/failure models that correlate observable failure attributes with corresponding likely fault attributes so that likely underlying causes of observed failures can be inferred. 
     
     
         30 . The method of  claim 29  wherein the fault/failure models include those for different kinds of denial of service (DOS) attacks. 
     
     
         31 . The method of  claim 29  wherein the fault/failure models include those for different kinds hardware performance degradations. 
     
     
         32 . The method of  claim 22  and further comprising:
 determining a type of user to whom the dashboard is being presented; 
 automatically formatting the presented dashboard in accordance with predetermined presentation rules maintained within a knowledge database of the data processing system. 
 
     
     
         33 . The method of  claim 32  and further comprising:
 receiving presentation satisfaction feedback from different types of users with respect to automatically presented formats; and 
 adaptively modifying the presentation rules maintained within the knowledge database in accordance with the received satisfaction feedback. 
 
     
     
         34 . The method of  claim 33  wherein the presentation rules include those for preferred detail visualization formats and the presented dashboard expands to present the preferred detail visualization formats of the corresponding user. 
     
     
         35 . The method of  claim 34  wherein the detail visualization formats include presenting a multi-dimensional graph having a common first feature axis along which are plotted interrelated other features of a set of compound features related to a given alert report presented within the dashboard. 
     
     
         36 . An data processing system having one or more processors and configured to automatically carry out a method comprised of presenting hierarchically organized alert information for meaningful comprehension by human administrators of the data processing system, the method comprising:
 automatically ranking compound feature anomaly detections extracted in a de-duplicating manner from normalcy-violating log reports of a logs generating portion of the data processing system, the ranking being according to how the de-duplicated compound feature anomaly detections place in an informational entropy space relative to predetermined single feature anomaly detections that serve as benchmarks in the informational entropy space, the compound feature anomaly detections being comprised of respective sets of interrelated single features determined to correlate to respective normalcy-violations within a predetermined time window; and   presenting one or more alert reports in a dashboard such that for each presented alert report, the highest ranked and de-duplicated set of compound features related to that alert report is presented ahead of a next highest ranked and de-duplicated set of compound features related to that alert report.   
     
     
         37 . The data processing system of  claim 36  wherein:
 each presented de-duplicated set of compound features is unfurled or unfurl-able to present an identification of its more relevant features among its respective set of compound features. 
 
     
     
         38 . The data processing system of  claim 37  wherein at least one of the more relevant features is selected from the group consisting of:
 a common set of certain IP addresses from which requests are sourced; 
 a common set of certain HTTP request codes used in requests; 
 a common set of certain geographic addresses from which requests are sourced; 
 a common set of certain user identifications on behalf of which requests are sourced; 
 an above threshold number of requests sourced within the predetermined time window from a common set of certain IP addresses or certain geographic addresses; 
 above threshold response times for a common set of sourced within the predetermined time window; 
 above threshold congestion on a communications link within the predetermined time window; and 
 above threshold congestion in a requests servicing server within the predetermined time window. 
 
     
     
         39 . The data processing system of  claim 37  wherein one or more of the unfurled or unfurl-able identifications of a more relevant feature is expandable to present more detailed information about that identified more relevant feature. 
     
     
         40 . The data processing system of  claim 36  wherein:
 each presented alert report is expandable to show the normalcy-violating log reports on which the presented alert report is based. 
 
     
     
         41 . An data processing system having one or more processors and further comprising:
 an anomalies visualization generator configured to generate one or more visualizations for a user of the system, the one or more visualizations providing alert information in a corresponding one or more formats for meaningful comprehension by the user;   an alerts generator configured to generate one or more alerts having alert information that can be presented as one or more visualizations by the anomalies visualization generator; and   a knowledge database operatively coupled to the anomalies visualization generator and to the alerts generator, the knowledge database storing:
 adaptively modifiable expert rules for picking the one or more visualizations providing the alert information in a format that allows for meaningful comprehension by the user; 
 a plurality of end user models configured to model different kinds of users and presentation formats likely to be preferred by the different kinds of users; 
 a plurality of fault/failure models that correlate observable failure attributes with corresponding likely fault attributes so that likely underlying causes of observed failures can be inferred; 
 a plurality of anomaly identification rules; 
 a plurality of single and compound feature identification rules; and 
 a plurality of domain and context specified rules; 
   wherein the knowledge database is used by the anomalies visualization generator and by the alerts generator for selectively picking alert information that is likely to be useful for the user and visualization formats likely to be preferred by the user.

Join the waitlist — get patent alerts

Track US2018329932A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.