System and method for providing language processing model services on a network
Abstract
Apparatus and method for recommending and configuring LLM models for organizations. For example, LLM model usage requirements of one or more organizations are evaluated, including applications and users associated with each organization. A cost estimation is performed with respect to expected utilization of the plurality of LLM models and a subset of LLM models is recommended for each of the organizations, applications, and users, along with rate limits for each organization and corresponding applications based on a global threshold rate limit specified for the entity. Upon acceptance by an administrator, the global threshold rate limit is partitioned into a corresponding set of per-organization threshold rate limits; each organization threshold rate limit is allocated to a corresponding organization of the one or more organizations, and each respective threshold rate limit is subdivided into portions to be allocated to applications of the respective organization.
Claims
exact text as granted — not AI-modified1 . A method implemented in a set of one or more electronic devices to generate and enforce limits on usage of a plurality of large language models (LLM), the method comprising:
evaluating LLM model usage requirements of one or more organizations of an entity, including applications and users associated with each organization, wherein the organizations include a first organization, wherein the evaluating includes determining, for the applications associated with the first organization or a subset thereof, types of LLM requests required to be serviced, each of the types of LLM requests associated with a different expected number of tokens per minute (TPM) and/or requests per minute (RPM); performing a cost estimation with respect to expected utilization of the plurality of LLM models by the one or more organizations and, the applications; determining a respective recommended subset of LLM models from the plurality of LLM models for each of the one or more organizations based on the usage requirements and the cost estimation, determining respective recommended threshold rate limits, in terms of TPM and/or RPM, for each organization and the associated applications based on a global threshold rate limit specified for the entity; providing the recommended subset of LLM models and rate limits for the first organization to an administrator, including options for accepting the recommended subset of LLM models and the rate limits and/or modifying one or more of the recommended subset of LLM models and/or rate limits, wherein responsive to the administrator accepting the recommendations with modifications or without modifications: allocating the first organization a threshold rate limit partitioned from the global threshold rate limit; subdividing the threshold rate limit into portions to be allocated to respective ones of the applications associated with the first organization; subdividing the portions into corresponding sub-portions to be allocated to users that are associated with the first organization and that use the respective ones of the applications; tracking runtime LLM usage by each organization, application, and user; and individually enforcing respective threshold rate limits allocated to the one or more organizations, including enforcing for the first organization the corresponding portion allocated to each of the respective ones of the applications, and the corresponding sub-portions allocated to the users, wherein the threshold rate limits manage load and enable efficient operation of the plurality of LLM models, thereby preventing overloading and ensuring that the plurality of LLM models can efficiently process incoming requests and provide responses.
2 . The method of claim 1 , further comprising:
presenting the cost estimation with respect to expected utilization of the subset of LLM models recommended for the first organization.
3 . The method of claim 1 , wherein individually enforcing the respective threshold rate limits is performed by a respective LLM governance engines operable within the corresponding organization.
4 . (canceled)
5 . The method of claim 14 , further comprising:
detecting that a threshold RPM or threshold TPM has been reached by one of the organizations, applications, or users; determining if spare RPM or spare TPM resources are available from other organizations, applications, or users; and responsively reallocating at least a portion of the spare RPM or spare TPM resources to the one of the organizations, applications, or users.
6 . The method of claim 1 , wherein providing the recommended subset of LLM models and rate limits for the first organization to an administrator comprises presenting the recommended subset of LLM models and rate limits for the first organization in a graphical user interface (GUI) to be accessed by the administrator, the GUI to provide options for accepting the recommended subset of LLM models and the rate limits and/or modifying one or more of the recommended subset of LLM models and rate limits.
7 . The method of claim 6 , wherein the GUI is to present a listing of applications which are suited for a corresponding recommended LLM, the listing comprising a set of entries, each entry corresponding to a different one of the applications.
8 . The method of claim 7 , wherein each entry is to provide an indication of one or more of: a type of the corresponding application, a currently assigned LLM model, if any, current performance metrics associated with the currently assigned LLM model, potential performance metrics corresponding to the recommended LLM model, current token or request metrics associated with the currently assigned LLM model, and recommended token or request metrics corresponding to the recommended LLM model.
9 . The method of claim 8 , wherein each entry is to further provide a selectable graphical element which, when selected, is to cause the GUI to display additional relevant information related to the application, the current LLM model, and the recommended LLM model, and is to provide a plurality of options for the administrator to adjust specified parameters for operation.
10 . The method of claim 9 , wherein a first option of the plurality of options comprises an input region for adjusting TPM or RPM values associated with the corresponding application and a second option to enable automatic adjustment of the TPM or RPM values by an LLM governance engine operable in the organization.
11 . A non-transitory machine-readable storage medium having program code stored thereon which, when executed by one or more electronic devices, are to cause the one or more electronic devices to generate and enforce limits on usage of a plurality of large language models (LLM) by performance of operations comprising:
evaluating LLM model usage requirements of one or more organizations of an entity, including applications and users associated with each organization, wherein the organizations include a first organization, wherein the evaluating includes determining, for the applications associated with the first organization or a subset thereof, types of LLM requests required to be serviced, each of the types of LLM requests associated with a different expected number of tokens per minute (TPM) and/or requests per minute (RPM); performing a cost estimation with respect to expected utilization of the plurality of LLM models by the one or more organizations and the, associated applications; determining a respective recommended subset of LLM models from the plurality of LLM models for each of the one or more organizations, based on the usage requirements and the cost estimation, determining respective recommending-rate limits, in terms of TPM and/or RPM, for each organization and the associated applications based on a global threshold rate limit specified for the entity; providing the recommended subset of LLM models and rate limits for the first organization to an administrator, including options for accepting the recommended subset of LLM models and the rate limits and/or modifying one or more of the recommended subset of LLM models and/or rate limits, wherein responsive to the administrator accepting the recommendations with modifications or without modifications:
allocating the first organization a threshold rate limit partitioned from the global threshold rate limit;
subdividing the threshold rate limit into portions to be allocated to respective ones of the applications associated with the first organization;
subdividing the portions into corresponding sub-portions to be allocated to users that are associated with the first organization and that use the respective ones of the applications;
tracking runtime LLM usage by each organization, application, and user; and individually enforcing respective threshold rate limits allocated to the one or more organizations, including enforcing for the first organization the corresponding portion allocated to each of the respective ones of the applications, and the corresponding sub-portions allocated to the users of the respective application, wherein the threshold rate limits manage load and enable efficient operation of the plurality of LLM models, thereby preventing overloading and ensuring that the plurality of LLM models can efficiently process incoming requests and provide responses.
12 . The non-transitory machine-readable storage medium of claim 11 ,
further comprising program code to cause the one or more electronic devices to perform the operations of: presenting the cost estimation with respect to expected utilization of the subset of LLM models recommended for the first organization.
13 . The non-transitory machine-readable storage medium of claim 11 ,
wherein individually enforcing the respective threshold rate limits is performed by a respective LLM governance engine operable within the corresponding organization.
14 . (canceled)
15 . The non-transitory machine-readable storage medium of claim 14 ,
further comprising program code to cause the one or more electronic devices to perform the operations of: detecting that a threshold RPM or threshold TPM has been reached by one of the organizations, applications, or users; determining if spare RPM or spare TPM resources are available from other organizations, applications, or users; and responsively reallocating at least a portion of the spare RPM or spare TPM resources to the one of the organizations, applications, or users.
16 . The non-transitory machine-readable storage medium of claim 11 , wherein providing the recommended subset of LLM models and rate limits for the first organization to an administrator comprises presenting the recommended subset of LLM models and rate limits in a graphical user interface (GUI) to be accessed by the administrator, the GUI to provide options for accepting the recommended subset of LLM models and the rate limits and/or modifying one or more of the recommended subset of LLM models and rate limits.
17 . The non-transitory machine-readable storage medium of claim 16 , wherein the GUI is to present a listing of applications which are suited for a corresponding recommended LLM, the listing comprising a set of entries, each entry corresponding to a different one of the applications.
18 . The non-transitory machine-readable storage medium of claim 17 , wherein each entry is to provide an indication of one or more of: a type of the corresponding application, a currently assigned LLM model, if any, current performance metrics associated with the currently assigned LLM model, potential performance metrics corresponding to the recommended LLM model, current token or request metrics associated with the currently assigned LLM model, and recommended token or request metrics corresponding to the recommended LLM model.
19 . The non-transitory machine-readable storage medium of claim 18 , wherein each entry is to further provide a selectable graphical element which, when selected, is to cause the GUI to display additional relevant information related to the application, the current LLM model, and the recommended LLM model, and is to provide a plurality of options for the administrator to adjust specified parameters for operation.
20 . The non-transitory machine-readable storage medium of claim 19 , wherein a first option of the plurality of options comprises an input region for adjusting TPM or RPM values associated with the corresponding application and a second option to enable automatic adjustment of the TPM or RPM values by an LLM governance engine operable in the organization.Join the waitlist — get patent alerts
Track US2026057422A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.