Proactive data prefetch with applied quality of service
Abstract
Examples described herein relate to prefetching content from a remote memory device to a memory tier local to a higher level cache or memory. An application or device can indicate a time availability for data to be available in a higher level cache or memory. A prefetcher used by a network interface can allocate resources in any intermediary network device in a data path from the remote memory device to the memory tier local to the higher level cache. Memory access bandwidth, egress bandwidth, memory space in any intermediary network device can be allocated for prefetch of content. In some examples, proactive prefetch can occur for content expected to be prefetched but not requested to be prefetched.
Claims
exact text as granted — not AI-modified1 - 22 . (canceled)
23 . Accelerator hardware circuitry configurable for use in carrying out artificial intelligence-related operations in association with at least one accelerator-to-accelerator communication interconnect, multiple network interface controllers, multiple central processing units, and at least one processor-to-processor communication interconnect, the multiple network interface controllers and multiple central processing units being comprised in multiple circuit boards that are communicatively coupled together via at least one switch, the accelerator hardware circuitry comprising:
at least one graphics processing unit (GPU) for being assigned, when the accelerator hardware circuitry is in operation, reserved memory space resources and reserved memory bandwidth resources for use in executing at least one workload, the reserved memory bandwidth resources to be allocated based, at least in part, upon quality of service-based scheduling; wherein:
when the accelerator hardware circuitry is in the operation:
the reserved memory space resources are configurable to comprise L2 cache memory;
the accelerator hardware circuitry is to communicate with other accelerator hardware circuitry via the at least one accelerator-to-accelerator communication interconnect;
at least one of the multiple central processing units is to communicate with at least one other of the multiple central processing units via the at least one processor-to-processor communication interconnect; and
the at least one workload is configurable to be container-based or virtual machine-based.
24 . The accelerator hardware circuitry of claim 23 , wherein:
the quality of service-based scheduling is associated with throughput and latency requirements.
25 . The accelerator hardware circuitry of claim 24 , wherein:
the artificial intelligence-related operations are associated, at least in part, with at least one cloud-based service.
26 . The accelerator hardware circuitry of claim 25 , wherein:
the multiple circuit boards comprise the at least one GPU; and the requirements are associated with a customer service level agreement.
27 . The accelerator hardware circuitry of claim 26 , wherein:
the at least one GPU comprises at least one physical GPU.
28 . At least one non-transitory machine-readable storage medium storing instructions for being executed, at least in part, by accelerator hardware circuitry, the accelerator hardware circuitry being configurable for use in carrying out artificial intelligence-related operations in association with at least one accelerator-to-accelerator communication interconnect, multiple network interface controllers, multiple central processing units, and at least one processor-to-processor communication interconnect, the multiple network interface controllers and multiple central processing units being comprised in multiple circuit boards that are communicatively coupled together via at least one switch, the accelerator hardware circuitry comprising at least one graphics processing unit (GPU), the instructions, when executed by the accelerator hardware circuitry, resulting in performance of operations comprising:
configuring the at least one GPU for assignment of reserved memory space resources and reserved memory bandwidth resources for use in executing at least one workload, the reserved memory bandwidth resources to be allocated based, at least in part, upon quality of service-based scheduling; wherein:
when the accelerator hardware circuitry is in the operation:
the reserved memory space resources are configurable to comprise L2 cache memory;
the accelerator hardware circuitry is to communicate with other accelerator hardware circuitry via the at least one accelerator-to-accelerator communication interconnect;
at least one of the multiple central processing units is to communicate with at least one other of the multiple central processing units via the at least one processor-to-processor communication interconnect; and
the at least one workload is configurable to be container-based or virtual machine-based.
29 . The at least one non-transitory machine-readable storage medium of claim 28 , wherein:
the quality of service-based scheduling is associated with throughput and latency requirements.
30 . The at least one non-transitory machine-readable storage medium of claim 29 , wherein:
the artificial intelligence-related operations are associated, at least in part, with at least one cloud-based service.
31 . The at least one non-transitory machine-readable storage medium of claim 30 , wherein:
the multiple circuit boards comprise the at least one GPU; and the requirements are associated with a customer service level agreement.
32 . The at least one non-transitory machine-readable storage medium of claim 31 , wherein:
the at least one GPU comprises at least one physical GPU.
33 . A physical computing platform for use in carrying out, when the physical computing platform is in operation, artificial intelligence-related operations, the physical computing platform comprising:
accelerator hardware circuitry; at least one accelerator-to-accelerator communication interconnect; multiple network interface controllers; multiple central processing units; at least one processor-to-processor communication interconnect; at least one switch; wherein:
the multiple network interface controllers and the multiple central processing units are comprised in multiple circuit boards that are communicatively coupled together via the at least one switch;
the accelerator hardware circuitry comprises at least one graphics processing unit (GPU) for being assigned, when the physical computing platform is in the operation, reserved memory space resources and reserved memory bandwidth resources for use in executing at least one workload;
the reserved memory bandwidth resources are to be allocated based, at least in part, upon quality of service-based scheduling;
when the physical computing platform is in the operation:
the reserved memory space resources are configurable to comprise L2 cache memory;
the accelerator hardware circuitry is to communicate with other accelerator hardware circuitry via the at least one accelerator-to-accelerator communication interconnect;
at least one of the multiple central processing units is to communicate with at least one other of the multiple central processing units via the at least one processor-to-processor communication interconnect; and
the at least one workload is configurable to be container-based or virtual machine-based.
34 . The physical computing platform of claim 33 , wherein:
the quality of service-based scheduling is associated with throughput and latency requirements.
35 . The physical computing platform of claim 34 , wherein:
the artificial intelligence-related operations are associated, at least in part, with at least one cloud-based service.
36 . The physical computing platform of claim 35 , wherein:
the multiple circuit boards comprise the at least one GPU; and the requirements are associated with a customer service level agreement.
37 . The physical computing platform of claim 36 , wherein:
the at least one GPU comprises at least one physical GPU.
38 . At least one non-transitory machine-readable storage medium storing instructions for being executed, at least in part, by a computing platform, the computing platform being associated with accelerator hardware circuitry, the computing platform being for use in carrying out, when the physical computing platform is in operation, artificial intelligence-related operations, the computing platform comprising at least one accelerator-to-accelerator communication interconnect, multiple network interface controllers, multiple central processing units, and at least one processor-to-processor communication interconnect, the multiple network interface controllers and multiple central processing units being comprised in multiple circuit boards that are communicatively coupled together via at least one switch, the accelerator hardware circuitry comprising at least one graphics processing unit (GPU), the instructions, when executed, at least in part, by the computing platform, resulting in performance of operations comprising:
configuring the at least one GPU for assignment of reserved memory space resources and reserved memory bandwidth resources for use in executing at least one workload, the reserved memory bandwidth resources to be allocated based, at least in part, upon quality of service-based scheduling; wherein:
when the computing platform is in the operation:
the reserved memory space resources are configurable to comprise L2 cache memory;
the accelerator hardware circuitry is to communicate with other accelerator hardware circuitry via the at least one accelerator-to-accelerator communication interconnect;
at least one of the multiple central processing units is to communicate with at least one other of the multiple central processing units via the at least one processor-to-processor communication interconnect; and
the at least one workload is configurable to be container-based or virtual machine-based.
39 . The at least one non-transitory machine-readable storage medium of claim 38 , wherein:
the quality of service-based scheduling is associated with throughput and latency requirements.
40 . The at least one non-transitory machine-readable storage medium of claim 39 , wherein:
the artificial intelligence-related operations are associated, at least in part, with at least one cloud-based service.
41 . The at least one non-transitory machine-readable storage medium of claim 40 , wherein:
the multiple circuit boards comprise the at least one GPU; and the requirements are associated with a customer service level agreement.
42 . The at least one non-transitory machine-readable storage medium of claim 41 , wherein:
the at least one GPU comprises at least one physical GPU.Join the waitlist — get patent alerts
Track US2021141731A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.