Method and system for reducing risk values discrepancies between categories
Abstract
The present teaching generally relates to removing perturbations from predictive scoring. In one embodiment, data representing a plurality of events detected by a content provider may be received, the data indicating a time that a corresponding event occurred and whether the corresponding event was fraudulent. First category data may be generated by grouping each event into one of a number of categories, each category being associated with a range of times. A first measure of risk for each category may be determined, where the first measure of risk indicates a likelihood that a future event occurring at a future time is fraudulent. Second category data may be generated by processing the first category data and a second measure of risk for each category may be determined. Measure data representing the second measure of risk for each category and the range of times associated with that category may be stored.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A method for determining a likelihood that a future event is fraudulent, the method being implemented on at least one computing device having at least one processor, memory, and communications circuitry, and the method comprising:
receiving event data representing a plurality of events detected by a content provider, wherein the event data indicates a feature associated with a corresponding event and whether the corresponding event was fraudulent; generating category data by grouping each of the plurality of events into one of categories, wherein each of the categories is associated with a range associated with the feature; training a predictive model based on the category data; and determining, based on the predictive model, a likelihood that a future event is fraudulent.
2 . The method of claim 1 , wherein the feature corresponds to hours of a day.
3 . The method of claim 2 , wherein each of the categories is associated with a distinct range of times in a day.
4 . The method of claim 1 , wherein the feature corresponds to one of:
a browser cookie age, a distance between a current IP address location of a user device associated with a user interaction and a home IP address associated with a most frequently used IP address of the user device, a number of visits in a certain number of days, a number of clicks for an entity in a certain number of days, an average click-through-rate for an entity in a certain number of days, an average fraud in an entity in a certain number of days, and a ratio of IP addresses over user agents on a website.
5 . The method of claim 1 , further comprising:
applying a smoothing function to the category data to generate smoothed data.
6 . The method of claim 5 , further comprising:
updating the predictive model based on the smoothed data.
7 . The method of claim 1 , further comprising:
allocating a portion of a memory for a data structure including the category data.
8 . A non-transitory computer readable medium comprising instructions for removing perturbations from predictive scoring, wherein the instructions, when read by at least one processor of a computing device, cause the computing device to perform:
receiving event data representing a plurality of events detected by a content provider, wherein the event data indicates a feature associated with a corresponding event and whether the corresponding event was fraudulent; generating category data by grouping each of the plurality of events into one of categories, wherein each of the categories is associated with a range associated with the feature; training a predictive model based on the category data; and determining, based on the predictive model, a likelihood that a future event is fraudulent.
9 . The medium of claim 8 , wherein the feature corresponds to hours of a day.
10 . The medium of claim 9 , wherein each of the categories is associated with a distinct range of times in a day.
11 . The medium of claim 8 , wherein the feature corresponds to one of:
a browser cookie age, a distance between a current IP address location of a user device associated with a user interaction and a home IP address associated with a most frequently used IP address of the user device, a number of visits in a certain number of days, a number of clicks for an entity in a certain number of days, an average click-through-rate for an entity in a certain number of days, an average fraud in an entity in a certain number of days, and a ratio of IP addresses over user agents on a website.
12 . The medium of claim 8 , wherein the instructions, when read by the processor of the computing device, cause the computing device to further perform:
applying a smoothing function to the category data to generate smoothed data.
13 . The medium of claim 12 , wherein the instructions, when read by the processor of the computing device, cause the computing device to further perform:
updating the predictive model based on the smoothed data.
14 . The medium of claim 8 , wherein the instructions, when read by the processor of the computing device, cause the computing device to further perform:
allocating a portion of a memory for a data structure including the category data.
15 . A system for removing perturbations from predictive scoring, the system comprising:
memory storing computer program instructions; and one or more processors that, in response to executing the computer program instructions, effectuate operations comprising:
receiving event data representing a plurality of events detected by a content provider, wherein the event data indicates a feature associated with a corresponding event and whether the corresponding event was fraudulent;
generating category data by grouping each of the plurality of events into one of categories, wherein each of the categories is associated with a range associated with the feature;
training a predictive model based on the category data; and
determining, based on the predictive model, a likelihood that a future event is fraudulent.
16 . The system of claim 15 , wherein the feature corresponds to hours of a day.
17 . The system of claim 16 , wherein each of the categories is associated with a distinct range of times in a day.
18 . The system of claim 15 , wherein the operations further comprise:
applying a smoothing function to the category data to generate smoothed data.
19 . The system of claim 18 , wherein the operations further comprise:
updating the predictive model based on the smoothed data.
20 . The system of claim 15 , wherein the operations further comprise:
allocating a portion of a memory for a data structure including the category data.Join the waitlist — get patent alerts
Track US2022383168A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.