About the Temporal Alignment Track category

According the documentation:

“”"
To adapt this dataset for text-conditional sounding video generation, captions for all video clips were automatically generated using LLaVA-Next. These captions are provided along with the video clips.
“”"

But I could not find the captions in the dataset provided.