US2025356667A1PendingUtilityA1

Method and apparatus with three-dimensional object detection

Assignee: SAMSUNG ELECTRONICS CO LTDPriority: May 16, 2024Filed: Jan 29, 2025Published: Nov 20, 2025
Est. expiryMay 16, 2044(~17.8 yrs left)· nominal 20-yr term from priority
G06T 2200/04G06V 10/764G06V 10/776G06V 10/72G06V 10/751G06V 10/24G06V 10/7715G06T 7/73G06T 7/85G06T 7/50G06V 10/80G06V 20/50G06V 10/74G06V 10/82G06V 20/64G06T 2207/20084G06T 2207/20081G06N 3/0499G06N 3/0464G06N 3/084G06V 10/765G06V 10/40G06T 7/593
57
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method of detecting a three-dimensional (3D) object includes: extracting two-dimensional (2D) image features from images using an image backbone; extracting a 3D feature map, reflecting depth prediction information, from the 2D image features by using a view transformer configured to perform domain generalization; extracting a bird's eye view (BEV) feature from the 3D feature map by using a BEV encoder; and predicting a position of the object and a class of the object from the BEV feature by using a detection head.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of detecting a three-dimensional (3D) object, the method comprising:
 extracting two-dimensional (2D) image features from images using an image backbone;   extracting a 3D feature map, the 3D feature map reflecting depth prediction information, from the 2D image features by using a view transformer configured to perform domain generalization;   extracting a bird's eye view (BEV) feature from the 3D feature map by using a BEV encoder; and   predicting a position of the object and a class of the object from the BEV feature by using a detection head.   
     
     
         2 . The method of  claim 1 , wherein the 3D feature map is extracted by a DepthNet predicting a depth output from the 2D image features and by inputting, into a BEV pool, an outer product of the depth output of the DepthNet and the 2D image features. 
     
     
         3 . The method of  claim 1 , wherein the view transformer is configured to perform a relative depth normalization method that minimizes depth and position prediction errors caused by a difference in intrinsic/extrinsic parameters of a camera that provided one of the images. 
     
     
         4 . The method of  claim 3 , wherein cameras, including the camera, provide the respective images, and wherein the relative depth normalization method comprises calculating a transformation matrix through which geometric transformation is performed between adjacent pairs of the cameras from the intrinsic/extrinsic parameters and the camera. 
     
     
         5 . The method of  claim 4 , wherein the relative depth normalization method obtains a relative depth after projecting an image feature onto an adjacent image feature by using the depth prediction information and the transformation matrix and minimizing a relative depth loss based on a depth loss function. 
     
     
         6 . The method of  claim 1 , wherein the view transformer is configured to perform a photometric matching method using depth prediction to optimize alignment between an image and an adjacent image, based on the photometric matching method. 
     
     
         7 . The method of  claim 1 , wherein the image backbone, the view transformer, the BEV encoder, and/or the detection head comprise respective domain adaptation adapters. 
     
     
         8 . The method of  claim 7 , wherein each domain adaptation adapter is added in parallel to an operation block to enable fine-tuning on parameters. 
     
     
         9 . The method of  claim 7 , wherein each domain adaptation adapter is configured to perform a skip connection in which features input to the view transformer, the BEV encoder, and/or the detection head are received, operated, and summed to update a gradient.  10  The method of  claim 1 , further comprising augmenting the 3D feature map by performing a generalization method of decoupling-based image depth estimation. 
     
     
         11 . An electronic device comprising:
 a memory storing instructions; and   one or more processors,   wherein the instructions, when performed by the one or more processors, cause the one or more processors to
 extract two-dimensional (2D) image features from images using an image backbone, 
 extract a 3D feature map, the 3D feature map reflecting depth prediction information, from the 2D image features by using a view transformer, 
 extract a bird's eye view (BEV) feature from the 3D feature map by using a BEV encoder, and 
 predict a position of the object and a class of the object from the BEV feature by using a detection head. 
   
     
     
         12 . The electronic device of  claim 11 , wherein the 3D feature map is extracted by a DepthNet predicting a depth output from the 2D image features and by inputting, into a BEV pool, an output of the DepthNet and the 2D image features. 
     
     
         13 . The electronic device of  claim 11 , wherein the view transformer is configured to perform a relative depth normalization method that minimizes depth and position prediction errors caused by a difference in intrinsic/extrinsic parameters of a camera that provided one of the images. 
     
     
         14 . The electronic device of  claim 13 , wherein cameras, including the camera, provide the respective images, and wherein the relative depth normalization method comprises calculating a transformation matrix through which geometric transformation is performed between adjacent pairs of the cameras from the intrinsic/extrinsic parameters and the camera. 
     
     
         15 . The electronic device of  claim 14 , wherein the relative depth normalization method obtains a relative depth after projecting an image feature onto an adjacent image feature by using the depth prediction information and the transformation matrix and minimizing a relative depth loss based on a depth loss function. 
     
     
         16 . The electronic device of  claim 11 , wherein the view transformer is configured to perform a photometric matching method using depth prediction to optimize alignment between an image and an adjacent image, based on the photometric matching method. 
     
     
         17 . The electronic device of  claim 11 , wherein the image backbone, the view transformer, the BEV encoder, and/or the detection head have respective domain adaptation adapters. 
     
     
         18 . The electronic device of  claim 17 , wherein the domain adaptation adapters temporarily supplant layers in the image backbone, the view transformer, the BEV encoder, and/or the detection head, respectively. 
     
     
         19 . The electronic device of  claim 17 , wherein each domain adaptation adapter is configured to perform a skip connection in which features input to the corresponding view transformer, the BEV encoder, and/or the detection head are received, operated, and summed to update a gradient. 
     
     
         20 . The electronic device of  claim 11 , wherein the 3D feature map is augmented by performing a generalization method of decoupling-based image depth estimation.

Join the waitlist — get patent alerts

Track US2025356667A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.