According the documentation:
“”"
To adapt this dataset for text-conditional sounding video generation, captions for all video clips were automatically generated using LLaVA-Next. These captions are provided along with the video clips.
“”"
But I could not find the captions in the dataset provided.