Photorealistic synthesis of agents in traffic scenes
Abstract
A computer-implemented method for synthesizing an image includes extracting agent neural radiance fields (NeRFs) from driving video logs and storing agent NeRFs in a database. For a driving video log to be edited, a scene NeRF and agent NeRFs are extracted from the driving video log to be edited. One or more agent NeRFs are selected from the database to insert into or replace existing agents in a traffic scene of the driving video log based on photorealism criteria. The traffic scene is edited by inserting a selected agent NeRF into the traffic scene, replacing existing agents in the traffic scene with the selected agent NeRF, or removing one or more existing agents from the traffic scene. An image of the edited traffic scene is synthesized by composing edited agent NeRFs with the scene NeRF and performing volume rendering.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method for synthesizing an image, comprising:
extracting agent neural radiance fields (NeRFs) from driving video logs; storing agent NeRFs in a database; for a driving video log to be edited, extracting a scene NeRF and agent NeRFs from the driving video log to be edited; selecting one or more agent NeRFs from the database to insert into or replace existing agents in a traffic scene of the driving video log based on photorealism criteria; editing the traffic scene by at least one of inserting a selected agent NeRF into the traffic scene, replacing existing agents in the traffic scene with a selected agent NeRF, or removing one or more existing agents from the traffic scene; and synthesizing an image of an edited traffic scene by composing edited agent NeRFs with the scene NeRF and performing volume rendering.
2 . The method of claim 1 , wherein the photorealism criteria include at least one of:
view consistency between inserted agents and the traffic scene, illumination consistency between inserted agents and the traffic scene, or shadow consistency between inserted agents and the traffic scene.
3 . The method of claim 1 , further comprising:
receiving outputs from an external behavior simulator indicating insertion locations, orientations, and driving trajectories for agents to be inserted into the traffic scene.
4 . The method of claim 1 , wherein editing the traffic scene comprises:
randomly identifying one or more existing agents for removal from the traffic scene.
5 . The method of claim 1 , further comprising:
representing each agent as a learnable latent code; and applying a hyper-network to map the latent code to parameters of a feature grid for each agent.
6 . The method of claim 5 , further comprising: using the feature grid to generate feature vectors for three-dimensional points associated with an agent.
7 . The method of claim 6 , wherein synthesizing the image comprises:
decoding the feature vectors using multilayer perceptrons to determine density and color for points along rendering rays; and performing volume rendering using the density and color.
8 . A system for synthesizing an image, comprising:
a hardware processor; and a memory storing a computer program which, when executed by the hardware processor, causes the hardware processor to: extract agent neural radiance fields (NeRFs) from driving video logs; store agent NeRFs in a database; for a driving video log to be edited, extract a scene NeRF and agent NeRFs from the driving video log to be edited; select one or more agent NeRFs from the database to insert into or replace existing agents in a traffic scene of the driving video log based on photorealism criteria; edit the traffic scene by at least one of inserting a selected agent NeRF into the traffic scene, replacing existing agents in the traffic scene with the selected agent NeRF, or removing one or more existing agents from the traffic scene; and synthesize an image of the traffic scene by composing edited agent NeRFs with the scene NeRF and performing volume rendering.
9 . The system of claim 8 , wherein the photorealism criteria include at least one of:
view consistency between inserted agents and the traffic scene, illumination consistency between inserted agents and the traffic scene, or shadow consistency between inserted agents and the traffic scene.
10 . The system of claim 8 , wherein the computer program further causes the hardware processor to:
receive outputs from an external behavior simulator indicating insertion locations, orientations, and driving trajectories for agents to be inserted into the traffic scene.
11 . The system of claim 8 , wherein editing the traffic scene comprises:
randomly identifying one or more existing agents for removal from the traffic scene.
12 . The system of claim 8 , wherein the computer program further causes the hardware processor to:
represent each agent as a learnable latent code; and apply a hyper-network to map the latent code to parameters of a feature grid for each agent.
13 . The system of claim 12 , wherein the computer program further causes the hardware processor to:
use the feature grid to generate feature vectors for three-dimensional points associated with an agent.
14 . The system of claim 13 , wherein synthesizing the image comprises:
decoding the feature vectors using multilayer perceptrons to determine density and color for points along rendering rays; and performing volume rendering using the density and color.
15 . A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform a method for synthesizing an image, the method comprising:
extract agent neural radiance fields (NeRFs) from driving video logs; store agent NeRFs in a database; for a driving video log to be edited, extract a scene NeRF and agent NeRFs from the driving video log to be edited; select one or more agent NeRFs from the database to insert into or replace existing agents in a traffic scene of the driving video log based on photorealism criteria; edit the traffic scene by at least one of inserting a selected agent NeRF into the traffic scene, replacing existing agents in the traffic scene with the selected agent NeRF, or removing one or more existing agents from the traffic scene; and synthesize an image of the traffic scene by composing edited agent NeRFs with the scene NeRF and performing volume rendering.
16 . The non-transitory computer-readable storage medium of claim 15 , wherein the photorealism criteria include at least one of:
view consistency between inserted agents and the traffic scene, illumination consistency between inserted agents and the traffic scene, or shadow consistency between inserted agents and the traffic scene.
17 . The non-transitory computer-readable storage medium of claim 15 , wherein the method further comprises:
receiving outputs from an external behavior simulator indicating insertion locations, orientations, and driving trajectories for agents to be inserted into the traffic scene.
18 . The non-transitory computer-readable storage medium of claim 15 , wherein editing the traffic scene comprises:
randomly identifying one or more existing agents for removal from the traffic scene.
19 . The non-transitory computer-readable storage medium of claim 15 , wherein the method further comprises:
representing each agent as a learnable latent code; and applying a hyper-network to map the latent code to parameters of a feature grid for each agent.
20 . The non-transitory computer-readable storage medium of claim 19 , wherein synthesizing the image comprises:
using the feature grid to generate feature vectors for three-dimensional points associated with each agent; decoding the feature vectors using multilayer perceptrons to determine density and color for points along rendering rays; and performing volume rendering using the density and color.Join the waitlist — get patent alerts
Track US2025148736A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.