US2025110781A1PendingUtilityA1

Computer-based systems and/or computing devices configured for the scaling of computing resources using a machine learning model trained to monitor and/or predict usage of inference models

Assignee: CAPITAL ONE SERVICES LLCPriority: Nov 30, 2021Filed: Sep 9, 2024Published: Apr 3, 2025
Est. expiryNov 30, 2041(~15.3 yrs left)· nominal 20-yr term from priority
G06F 21/6245G06N 20/00G06F 11/3409G06F 9/5027G06N 5/04G06F 2209/5019G06F 9/505G06F 9/5016
75
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

An example method includes receiving historical usage data associated with computing services provided by distributed servers and an inference model. The inference model is configured to receive a request and make an inference based on the request. The method further includes training a machine learning model to determine a correlation between usage of a first computing service of the and usage of the inference model. The correlation indicates that a first spike in usage of the first computing service precedes a second spike in usage of the inference model. The method further includes receiving, in real-time, current usage data associated with the first computing service. The method further includes determining, based on the current usage data and the correlation, that the current usage data is indicative of the first spike in usage of the first computing service that precedes the second spike in usage of the inference model.

Claims

exact text as granted — not AI-modified
1 - 20 . (canceled) 
     
     
         21 . A method comprising:
 receiving, by one or more processors of one or more computing devices, current usage data associated with at least one first computing service of a plurality of computing services;   determining, by the one or more processors based on the current usage data and a machine learning model configured to determine a correlation between usage of at least one first computing service and usage of an inference model, that the current usage data is indicative of a future increase or decrease in usage of the inference model; and   transmitting, by the one or more processors in response to the determining that the current usage data is indicative of the future increase or decrease in usage of the inference model, at least one command causing an increase or decrease in an amount of computing resources available for an execution of the inference model.   
     
     
         22 . The method of  claim 21 , wherein the machine learning model is trained based on historical usage data to determine the correlation between the usage of the at least one first computing service and the usage of the inference model. 
     
     
         23 . The method of  claim 22 , wherein the historical data is associated with:
 the plurality of computing services provided by a plurality of distributed servers; and   the inference model, wherein the inference model is associated with at least one of the plurality of computing services.   
     
     
         24 . The method of  claim 21 , wherein the inference model is configured to receive a request and make an inference based on the request. 
     
     
         25 . The method of  claim 21 , wherein the correlation indicates that:
 at least one first spike in the usage of the at least one first computing service precedes at least one second spike in the usage of the inference model; or   at least one first decrease in the usage of the at least one first computing service precedes at least one second decrease in the usage of the inference model.   
     
     
         26 . The method of  claim 21 , wherein the at least one command causing the increase or decrease in the amount of computing resources available for the execution of the inference model comprises a command to increase the amount of computing resources, and further wherein the at least one command to increase the amount of computing resources available for the execution of the inference model comprises one or more of:
 instructions to allocate additional graphics processing units (GPUs) or central processing units (CPUs) for use by the inference model;   instructions to allocate additional memory for use by the inference model; or instructions to allocate additional nodes of a cloud computing service for use by the inference model.   
     
     
         27 . The method of  claim 21 , wherein the at least one command causing the increase or decrease in the amount of computing resources available for the execution of the inference model comprises a command to decrease the amount of computing resources, and further wherein the at least one command to decrease the amount of computing resources available for the execution of the inference model comprises one or more of:
 instructions to allocate fewer graphics processing units (GPUs) or central processing units (CPUs) for use by the inference model;   instructions to allocate less memory for use by the inference model; or   instructions to allocate fewer nodes of a cloud computing service for use by the inference model.   
     
     
         28 . The method of  claim 21 , wherein the inference model comprises one or more of a credit checking service, a credit limit estimation service, a line of credit approval service, a transaction fraud protection monitoring service, or an electronic message risk scanning service. 
     
     
         29 . The method of  claim 21 , wherein the at least one first computing service comprises one or more of a website provider service, an advertisement provider service, an in-store traffic monitoring service, a transaction or purchase tracking service, an electronic message sending and receiving service, or a game console service. 
     
     
         30 . The method of  claim 21 , wherein the at least one first computing service is subject to a security protocol comprising a requirement to limit access to historical sensitive usage data of the at least one first computing service. 
     
     
         31 . The method of  claim 30 , wherein the machine learning model is trained using synthetic usage data comprising an anonymized version of the historical sensitive usage data. 
     
     
         32 . The method of  claim 21 , wherein the at least one first computing service is subject to a security protocol comprising a requirement to limit access to sensitive current usage data of the at least one first computing service. 
     
     
         33 . The method of  claim 32 , further comprising generating, by the one or more processors, the current usage data associated with the at least one first computing service by converting the sensitive current usage data to synthetic usage data, wherein the synthetic usage data comprises an anonymized version of the sensitive current usage data. 
     
     
         34 . The method of  claim 21 , wherein the receiving of the current usage data comprises monitoring, by the one or more processors, the usage of the at least one first computing service by monitoring a number of application programming interface (API) calls into and/or out of the at least one first computing service. 
     
     
         35 . A system comprising:
 a memory; and   at least one processor coupled to the memory, the at least one processor configured to:   receive current usage data associated with at least one first computing service of a plurality of computing services;   determine, based on the current usage data and a machine learning model configured to determine a correlation between usage of at least one first computing service and usage of an inference model, that the current usage data is indicative of a future increase in usage of the inference model; and   transmit, in response to the determination that the current usage data is indicative of the future increase in usage of the inference model, at least one command configured to cause an increase in an amount of computing resources available for an execution of the inference model.   
     
     
         36 . The system of  claim 35 , wherein the at least one command configured to cause the increase or decrease in the amount of computing resources available for the execution of the inference model comprises a command to decrease the amount of computing resources, and further wherein the at least one command to decrease the amount of computing resources available for the execution of the inference model comprises one or more of:
 instructions to allocate fewer graphics processing units (GPUs) or central processing units (CPUs) for use by the inference model;   instructions to allocate less memory for use by the inference model; or instructions to allocate fewer nodes of a cloud computing service for use by the inference model.   
     
     
         37 . The system of  claim 35 , wherein the inference model comprises one or more of a credit checking service, a credit limit estimation service, a line of credit approval service, a transaction fraud protection monitoring service, or an electronic message risk scanning service. 
     
     
         38 . The system of  claim 35 , wherein the at least one first computing service comprises one or more of a website provider service, an advertisement provider service, an in-store traffic monitoring service, a transaction or purchase tracking service, an electronic message sending and receiving service, or a game console service. 
     
     
         39 . A non-transitory computer readable medium having instructions stored thereon that, upon execution by a computing device, cause the computing device to perform operations comprising:
 receiving current usage data associated with at least one first computing service of a plurality of computing services;   determining, based on the current usage data and a machine learning model configured to determine a correlation between usage of at least one first computing service and usage of an inference model, that the current usage data is indicative of a future increase in usage of the inference model; and   transmitting, in response to the determination that the current usage data is indicative of the future increase in usage of the inference model, at least one command thereby causing an increase in an amount of computing resources available for an execution of the inference model.   
     
     
         40 . The non-transitory computer readable medium of  claim 39 , wherein the at least one command causing the increase or decrease in the amount of computing resources available for the execution of the inference model comprises a command to decrease the amount of computing resources, and further wherein the at least one command to decrease the amount of computing resources available for the execution of the inference model comprises one or more of:
 instructions to allocate fewer graphics processing units (GPUs) or central processing units (CPUs) for use by the inference model;   instructions to allocate less memory for use by the inference model; or   instructions to allocate fewer nodes of a cloud computing service for use by the inference model.

Join the waitlist — get patent alerts

Track US2025110781A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.