US2025356667A1PendingUtilityA1
Method and apparatus with three-dimensional object detection
Assignee: SAMSUNG ELECTRONICS CO LTDPriority: May 16, 2024Filed: Jan 29, 2025Published: Nov 20, 2025
Est. expiryMay 16, 2044(~17.8 yrs left)· nominal 20-yr term from priority
G06T 2200/04G06V 10/764G06V 10/776G06V 10/72G06V 10/751G06V 10/24G06V 10/7715G06T 7/73G06T 7/85G06T 7/50G06V 10/80G06V 20/50G06V 10/74G06V 10/82G06V 20/64G06T 2207/20084G06T 2207/20081G06N 3/0499G06N 3/0464G06N 3/084G06V 10/765G06V 10/40G06T 7/593
57
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A method of detecting a three-dimensional (3D) object includes: extracting two-dimensional (2D) image features from images using an image backbone; extracting a 3D feature map, reflecting depth prediction information, from the 2D image features by using a view transformer configured to perform domain generalization; extracting a bird's eye view (BEV) feature from the 3D feature map by using a BEV encoder; and predicting a position of the object and a class of the object from the BEV feature by using a detection head.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of detecting a three-dimensional (3D) object, the method comprising:
extracting two-dimensional (2D) image features from images using an image backbone; extracting a 3D feature map, the 3D feature map reflecting depth prediction information, from the 2D image features by using a view transformer configured to perform domain generalization; extracting a bird's eye view (BEV) feature from the 3D feature map by using a BEV encoder; and predicting a position of the object and a class of the object from the BEV feature by using a detection head.
2 . The method of claim 1 , wherein the 3D feature map is extracted by a DepthNet predicting a depth output from the 2D image features and by inputting, into a BEV pool, an outer product of the depth output of the DepthNet and the 2D image features.
3 . The method of claim 1 , wherein the view transformer is configured to perform a relative depth normalization method that minimizes depth and position prediction errors caused by a difference in intrinsic/extrinsic parameters of a camera that provided one of the images.
4 . The method of claim 3 , wherein cameras, including the camera, provide the respective images, and wherein the relative depth normalization method comprises calculating a transformation matrix through which geometric transformation is performed between adjacent pairs of the cameras from the intrinsic/extrinsic parameters and the camera.
5 . The method of claim 4 , wherein the relative depth normalization method obtains a relative depth after projecting an image feature onto an adjacent image feature by using the depth prediction information and the transformation matrix and minimizing a relative depth loss based on a depth loss function.
6 . The method of claim 1 , wherein the view transformer is configured to perform a photometric matching method using depth prediction to optimize alignment between an image and an adjacent image, based on the photometric matching method.
7 . The method of claim 1 , wherein the image backbone, the view transformer, the BEV encoder, and/or the detection head comprise respective domain adaptation adapters.
8 . The method of claim 7 , wherein each domain adaptation adapter is added in parallel to an operation block to enable fine-tuning on parameters.
9 . The method of claim 7 , wherein each domain adaptation adapter is configured to perform a skip connection in which features input to the view transformer, the BEV encoder, and/or the detection head are received, operated, and summed to update a gradient. 10 The method of claim 1 , further comprising augmenting the 3D feature map by performing a generalization method of decoupling-based image depth estimation.
11 . An electronic device comprising:
a memory storing instructions; and one or more processors, wherein the instructions, when performed by the one or more processors, cause the one or more processors to
extract two-dimensional (2D) image features from images using an image backbone,
extract a 3D feature map, the 3D feature map reflecting depth prediction information, from the 2D image features by using a view transformer,
extract a bird's eye view (BEV) feature from the 3D feature map by using a BEV encoder, and
predict a position of the object and a class of the object from the BEV feature by using a detection head.
12 . The electronic device of claim 11 , wherein the 3D feature map is extracted by a DepthNet predicting a depth output from the 2D image features and by inputting, into a BEV pool, an output of the DepthNet and the 2D image features.
13 . The electronic device of claim 11 , wherein the view transformer is configured to perform a relative depth normalization method that minimizes depth and position prediction errors caused by a difference in intrinsic/extrinsic parameters of a camera that provided one of the images.
14 . The electronic device of claim 13 , wherein cameras, including the camera, provide the respective images, and wherein the relative depth normalization method comprises calculating a transformation matrix through which geometric transformation is performed between adjacent pairs of the cameras from the intrinsic/extrinsic parameters and the camera.
15 . The electronic device of claim 14 , wherein the relative depth normalization method obtains a relative depth after projecting an image feature onto an adjacent image feature by using the depth prediction information and the transformation matrix and minimizing a relative depth loss based on a depth loss function.
16 . The electronic device of claim 11 , wherein the view transformer is configured to perform a photometric matching method using depth prediction to optimize alignment between an image and an adjacent image, based on the photometric matching method.
17 . The electronic device of claim 11 , wherein the image backbone, the view transformer, the BEV encoder, and/or the detection head have respective domain adaptation adapters.
18 . The electronic device of claim 17 , wherein the domain adaptation adapters temporarily supplant layers in the image backbone, the view transformer, the BEV encoder, and/or the detection head, respectively.
19 . The electronic device of claim 17 , wherein each domain adaptation adapter is configured to perform a skip connection in which features input to the corresponding view transformer, the BEV encoder, and/or the detection head are received, operated, and summed to update a gradient.
20 . The electronic device of claim 11 , wherein the 3D feature map is augmented by performing a generalization method of decoupling-based image depth estimation.Join the waitlist — get patent alerts
Track US2025356667A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.