FAVDBench Dataset
In CVPR2023, we proposed the Fine-grained Audible Video Description (FAVD) task, which aims to provide detailed textual descriptions of audible videos, including the appearance and spatial location of each object, the actions of moving objects, and the sounds in the video. We also contributed the first fine-grained audible video description dataset, FAVDBench, to the community. For each video segment, we not only provide a one-sentence video summary but also 4-6 sentences describing the visual details of the video and 1-2 audio-related descriptions, with all annotations available in both Chinese and English.