Skip to main navigation Skip to search Skip to main content

Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis

  • Lei-Lei Li
  • , Jianwu Fang
  • , Junbin Xiao
  • , Shanmin Pang
  • , Hongkai Yu
  • , Chen Lv
  • , Jianru Xue
  • , Tat-Seng Chua
  • Xi’an Jiaotong University
  • National University of Singapore
  • Nanyang Technological University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

2 Scopus citations

Abstract

Egocentricly comprehending the causes and effects of car accidents is crucial for the safety of self-driving cars, and synthesizing causal-entity reflected accident videos can facilitate the capability test to respond to unaffordable accidents in reality. However, incorporating causal relations as seen in real-world videos into synthetic videos remains challenging. This work argues that precisely identifying the accident participants and capturing their related behaviors are of critical importance. In this regard, we propose a novel diffusion model Causal-VidSyn for synthesizing egocentric traffic accident videos. To enable causal entity grounding in video diffusion, Causal-VidSyn leverages the cause descriptions and driver fixations to identify the accident participants and behaviors, facilitated by accident reason answering and gaze-conditioned selection modules. To support CausalVidSyn, we further construct Drive-Gaze, the largest driver gaze dataset (with 1.54M frames of fixations) in driving accident scenarios. Extensive experiments show that CausalVidSyn surpasses state-of-the-art video diffusion models in terms of frame quality and causal sensitivity in various tasks, including accident video editing, normal-to-accident video diffusion, and text-to-video generation.
Original languageEnglish
Title of host publicationProceedings of the IEEE International Conference on Computer Vision
Place of Publicationusa
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages11208-11218
Number of pages11
ISBN (Electronic)9798331587758
DOIs
StatePublished - Jan 1 2025
Event2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 - Honolulu, United States
Duration: Oct 19 2025Oct 23 2025

Conference

Conference2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
Country/TerritoryUnited States
Period10/19/2510/23/25

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • accident video synthesis
  • causal reasoning
  • driver attention
  • video diffusion models

Cite this