US2026051151A1PendingUtilityA1
Method and device for training an occupancy network
Est. expiryAug 16, 2044(~18 yrs left)· nominal 20-yr term from priority
G06V 20/56G06V 20/64G06V 10/82G06V 2201/07G06V 10/44G06V 10/764G06T 7/50G06T 2207/20081G06T 15/08
58
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A method and a device for training an occupancy network for classification and/or object recognition of an object in a scene.
Claims
exact text as granted — not AI-modified1 - 10 . (canceled)
11 . A method for training an occupancy network for classification and/or object recognition of an object in a scene, the method comprising the following steps:
providing 2D image data of the scene; extracting ground truth features from the 2D image data using a pre-trained feature encoder to provide respective ground truth feature embeddings; extracting image features for each 2D image of the provided 2D image data using a feature extraction network of the occupancy network; transforming the extracted 2D image features into a 3D voxel space using an occupancy transformer of the occupancy network to estimate a 3D occupancy probability and respective queryable 3D features of a particular voxel of a predetermined voxel grid in the voxel space; back-transforming the estimated 3D occupancy probability and the estimated queryable 3D features into a 2D representation using volume rendering; training the occupancy network based on a loss function over the 2D representation of the estimated 3D occupancy probability and the estimated queryable 3D features and based on the ground truth feature embeddings; and providing the trained occupancy network for classification and/or object recognition of an object in a scene.
12 . The method according to claim 11 , wherein the loss function is optimized by backpropagation of a loss.
13 . The method according to claim 11 , wherein the feature extraction network includes a backbone network or feature pyramid network.
14 . The method according to claim 11 , wherein each of the 3D features is defined by a multidimensional vector per voxel, wherein each of the multidimensional vectors describes a semantic content of a particular voxel in a latent space.
15 . The method according to claim 11 , wherein the loss function is further adapted using depth supervising information to form depth information data that are additionally provided.
16 . The method according to claim 11 , wherein the occupancy transformer carries out forward propagation or backward propagation.
17 . An inference method for classification and/or object recognition of an object in a scene using a trained occupancy network, the occupancy network having been trained by a method including the following steps:
providing 2D image data of the scene; extracting ground truth features from the 2D image data using a pre-trained feature encoder to provide respective ground truth feature embeddings; extracting image features for each 2D image of the provided 2D image data using a feature extraction network of the occupancy network; transforming the extracted 2D image features into a 3D voxel space using an occupancy transformer of the occupancy network to estimate a 3D occupancy probability and respective queryable 3D features of a particular voxel of a predetermined voxel grid in the voxel space; back-transforming the estimated 3D occupancy probability and the estimated queryable 3D features into a 2D representation using volume rendering; training the occupancy network based on a loss function over the 2D representation of the estimated 3D occupancy probability and the estimated queryable 3D features and based on the ground truth feature embeddings.
18 . A non-transitory computer-readable data carrier on which is stored program code of a computer program for training an occupancy network for classification and/or object recognition of an object in a scene, the program code, when executed by a computer, causing the computer to perform steps comprising the following:
providing 2D image data of the scene; extracting ground truth features from the 2D image data using a pre-trained feature encoder to provide respective ground truth feature embeddings; extracting image features for each 2D image of the provided 2D image data using a feature extraction network of the occupancy network; transforming the extracted 2D image features into a 3D voxel space using an occupancy transformer of the occupancy network to estimate a 3D occupancy probability and respective queryable 3D features of a particular voxel of a predetermined voxel grid in the voxel space; back-transforming the estimated 3D occupancy probability and the estimated queryable 3D features into a 2D representation using volume rendering; training the occupancy network based on a loss function over the 2D representation of the estimated 3D occupancy probability and the estimated queryable 3D features and based on the ground truth feature embeddings; and providing the trained occupancy network for classification and/or object recognition of an object in a scene.
19 . A device for training an occupancy network for classification and/or object recognition of an object in a scene, the device comprising:
an evaluation and computing unit configured to carry out the following steps:
providing 2D image data of the scene,
extracting ground truth features from the 2D image data using a pre-trained feature encoder to provide respective ground truth feature embeddings,
extracting image features for each 2D image of the provided 2D image data using a feature extraction network of the occupancy network,
transforming the extracted 2D image features into a 3D voxel space using an occupancy transformer of the occupancy network to estimate a 3D occupancy probability and respective queryable 3D features of a particular voxel of a predetermined voxel grid in the voxel space,
back-transforming the estimated 3D occupancy probability and the estimated queryable 3D features into a 2D representation using volume rendering,
training the occupancy network based on a loss function over the 2D representation of the estimated 3D occupancy probability and the estimated queryable 3D features and based on the ground truth feature embeddings, and
providing the trained occupancy network for classification and/or object recognition of an object in a scene.Join the waitlist — get patent alerts
Track US2026051151A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.