Estonian TV Subtitle Dataset

This dataset contains human-produced subtitles and the accommpanying audio files (extracted from video files), scraped from Estonian public television (ETV). Both the subtitles and audio are in Estonian (i.e subtitles are not translations!)

The dataset contains around 800 hours of audio.

NB! The copyright of the material belongs to ETV.

Downloading

etv-subtitles.tar (43 GB)

Contact

Tanel Alumäe tanel.alumae@taltech.ee

Citing

Artem Fedorchenko, Tanel Alumäe, 2025. Optimizing Estonian TV Subtitles with Semi-supervised Learning and LLMs. In: Proceedings of the 25th Nordic Conference on Computational Linguistics (NoDaLiDa)

@inproceedings{fedorchenko-2025-optimizing,
    title = "Optimizing Estonian {TV} Subtitles with Semi-supervised Learning and {LLMs}",
    author = {Fedorchenko, Artem and Alum{\"a}e, Tanel},
    booktitle = "Proceedings of the 25th Nordic Conference on Computational Linguistics (NoDaLiDa)",
    year = "2025"
}