US2023334316A1PendingUtilityA1
Dynamic distributed training of machine learning models
Est. expiryApr 24, 2037(~10.7 yrs left)· nominal 20-yr term from priority
Inventors:Altug KokerAbhishek R. AppuKamal SinhaJoydeep RayBalaji VembuElmoustapha Ould-Ahmed-VallSara S. BaghsorkhiAnbang YaoKevin NealisXiaoming ChenJohn C. WeastJustin E. GottschlichPrasoonkumar SurtiChandrasekaran SakthivelFarshad AkhbariNadathur Rajagopalan SatishLiwei MaJeremy BottlesonEriko NurvitadhiTravis T. SchluesslerAnkur N. ShahJonathan KennedyVasanth RanganathanSanjeev Jahagirdar
G06N 3/045G06N 3/08G06N 3/098G06N 3/09G06N 3/0895G06N 3/0495G06N 3/0464G06N 3/0442G06N 3/082G06F 11/1482G06F 11/1629G06N 20/00G06N 3/063G06N 3/044G06N 3/048G06T 1/20G06F 9/5027G06F 9/5066
73
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Described herein is a graphics processor comprising a memory device and a graphics processing cluster coupled with the memory device. The graphics processing cluster includes a plurality of graphics multiprocessors interconnected via a data interconnect. A graphics multiprocessor includes circuitry configured to load a modular neural network including a plurality of subnetworks, each of the plurality of subnetworks trained to perform a computer vision operation on a separate subject.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A graphics processor comprising:
a memory device; a graphics processing cluster coupled with the memory device, the graphics processing cluster including a plurality of graphics multiprocessors, the plurality of graphics multiprocessors interconnected via a data interconnect, wherein a graphics multiprocessor of the plurality of graphics multiprocessors includes circuitry configured to load a modular neural network including a plurality of subnetworks, each of the plurality of subnetworks trained to perform a computer vision operation on a separate subject, the graphics multiprocessor configured to:
load weights associated with a baseline set of layers of the modular neural network to the memory device;
determine a first subject associated with a deployment environment;
load weights for a first subnetwork to the memory device, the first subnetwork trained to recognize the first subject; and
perform a first matrix operation associated with the first subnetwork to facilitate a first computer vision operation on a first image captured within the deployment environment, wherein the first image includes the first subject.
2 . The graphics processor of claim 1 , the graphics multiprocessor configured to:
dynamically load weights for a second subnetwork to the memory device, the second subnetwork trained to recognize a second subject associated with the deployment environment; and perform a second matrix operation associated with the second subnetwork to facilitate a second computer vision operation on a second image captured within the deployment environment, wherein the second image includes the second subject.
3 . The graphics processor of claim 2 , wherein the graphics multiprocessor is configured such that output associated with the baseline set of layers of the modular neural network is provided as input the first subnetwork to perform the first computer vision operation and provided as input the second subnetwork to perform the second computer vision operation.
4 . The graphics processor of claim 3 , the graphics multiprocessor configured to:
apply a first priority to the first subnetwork; apply a second priority to the second subnetwork; and adjust a resource allocation within the graphics processing cluster based on the first priority and the second priority.
5 . The graphics processor of claim 4 , the graphics multiprocessor configured to:
load first optical flow data associated with a first sequence of images that includes the first image; and perform a third matrix operation associated with the first subnetwork to facilitate detection of a first set of stationary objects of the first subject, the third matrix operation performed based on the first optical flow data and the first sequence of images.
6 . The graphics processor of claim 5 , the graphics multiprocessor configured to:
load second optical flow data associated with a second sequence of images that includes the second image; and perform a fourth matrix operation associated with the second subnetwork to facilitate detection of a second set of stationary objects of the second subject, the fourth matrix operation performed based on the second optical flow data and the second sequence of images.
7 . The graphics processor of claim 6 , wherein the first optical flow data and the second optical flow data includes dense optical flow data.
8 . The graphics processor of claim 6 , the graphics multiprocessor configured to perform operations associated with the first subnetwork and the second subnetwork, the operations cause the graphics multiprocessor to:
determine velocities of moving objects of the first subject and the second subject; and assign a hazard associated with the moving objects to the first set of stationary objects and the second set of stationary objects based on a velocity and trajectory of the moving objects.
9 . A method comprising:
loading, into a memory device of a graphics processor, a modular neural network including a plurality of subnetworks, each of the plurality of subnetworks trained to perform a computer vision operation on a separate subject; loading, by a graphics processing cluster of the graphics processor, weights associated with a baseline set of layers of the modular neural network to the memory device; determining a first subject associated with a deployment environment; loading weights for a first subnetwork to the memory device, the first subnetwork trained to recognize the first subject; and performing a first matrix operation associated with the first subnetwork to facilitate a first computer vision operation on a first image captured within the deployment environment, wherein the first image includes the first subject.
10 . The method of claim 9 , further comprising:
dynamically loading weights for a second subnetwork to the memory device, the second subnetwork trained to recognize a second subject associated with the deployment environment; and performing a second matrix operation associated with the second subnetwork to facilitate a second computer vision operation on a second image captured within the deployment environment, wherein the second image includes the second subject.
11 . The method of claim 10 , wherein output associated with the baseline set of layers of the modular neural network is provided as input the first subnetwork to perform the first computer vision operation and provided as input the second subnetwork to perform the second computer vision operation.
12 . The method of claim 11 , further comprising:
applying a first priority to the first subnetwork; applying a second priority to the second subnetwork; and adjusting a resource allocation within the graphics processing cluster based on the first priority and the second priority.
13 . The method of claim 12 , further comprising:
loading first optical flow data associated with a first sequence of images that includes the first image; and performing a third matrix operation associated with the first subnetwork to facilitate detection of a first set of stationary objects of the first subject, the third matrix operation performed based on the first optical flow data and the first sequence of images.
14 . The method of claim 13 , further comprising:
loading second optical flow data associated with a second sequence of images that includes the second image; and performing a fourth matrix operation associated with the second subnetwork to facilitate detection of a second set of stationary objects of the second subject, the fourth matrix operation performed based on the second optical flow data and the second sequence of images.
15 . The method of claim 14 , wherein the first optical flow data and the second optical flow data includes dense optical flow data.
16 . The method of claim 14 , further comprising, via the first subnetwork and the second subnetwork:
determining velocities of moving objects of the first subject and the second subject; and assigning a hazard associated with the moving objects to the first set of stationary objects and the second set of stationary objects based on a velocity and trajectory of the moving objects.
17 . A data processing system comprising:
an input/output interface; and a graphics processor coupled with the input/output interface, the graphics processor comprising a memory device and a graphics processing cluster coupled with the memory device, the graphics processing cluster including a plurality of graphics multiprocessors, the plurality of graphics multiprocessors interconnected via a data interconnect, wherein a graphics multiprocessor of the plurality of graphics multiprocessors includes circuitry configured to load a modular neural network including a plurality of subnetworks, each of the plurality of subnetworks trained to perform a computer vision operation on a separate subject, and the graphics multiprocessor is configured to:
load weights associated with a baseline set of layers of the modular neural network to the memory device;
determine a first subject associated with a deployment environment;
load weights for a first subnetwork to the memory device, the first subnetwork trained to recognize the first subject; and
perform a first matrix operation associated with the first subnetwork to facilitate a first computer vision operation on a first image captured within the deployment environment, wherein the first image includes the first subject.
18 . The data processing system of claim 17 , the graphics multiprocessor configured to:
dynamically load weights for a second subnetwork to the memory device, the second subnetwork trained to recognize a second subject associated with the deployment environment; and perform a second matrix operation associated with the second subnetwork to facilitate a second computer vision operation on a second image captured within the deployment environment, wherein the second image includes the second subject.
19 . The data processing system of claim 18 , wherein the graphics multiprocessor is configured such that output associated with the baseline set of layers of the modular neural network is provided as input the first subnetwork to perform the first computer vision operation and provided as input the second subnetwork to perform the second computer vision operation.
20 . The data processing system of claim 19 , the graphics multiprocessor configured to:
apply a first priority to the first subnetwork; apply a second priority to the second subnetwork; and adjust a resource allocation within the graphics processing cluster based on the first priority and the second priority.Join the waitlist — get patent alerts
Track US2023334316A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.