← All Licenses
YO-CPT-ru Dataset License Notice
Added 2026-08-15
DataCustomHuggingFaceProprietary
Full Text
YO-CPT-ru — License
YO-CPT-ru is a derived dataset: beyond attribution, the authors impose no restrictions of their own —
users are responsible for complying with the upstream licenses below.
Contribution of the authors. The annotations and the compilation of the corpus are released under
[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) — attribution required, no other limitation on
use, including commercial use.
Upstream components. The source audio comes from [YODAS2](https://huggingface.co/datasets/espnet/yodas2),
collected exclusively from YouTube videos published under a Creative Commons license; the recordings remain
the intellectual property of their original creators. The released audio and metadata are produced by the
following components:
| component | released output | license |
|---|---|---|
| YODAS2 (source recordings) | `audio` | CC BY 3.0 |
| Silero VAD | — (segmentation only) | MIT |
| VoxBlink2 ResNet34 | `local_spk_id` | not stated (trained on CC BY-NC-SA 4.0 data) |
| DistillMOS | `mos_score` | MIT |
| Whisper-large-v3-turbo (custom) | `text` (ensemble) | MIT |
| GigaAM v3 RNNT | `text` (ensemble) | MIT |
| Vosk | `text` (ensemble) | Apache-2.0 |
| Spectra-0 | — (spoof filter only) | Apache-2.0 |
| ClearVoice MossFormer2_SE_48K | `audio` (enhanced waveform) | Apache-2.0 |
| wav2vec2-BERT (custom CTC) | `text_alignment` | MIT |
| RUAccent turbo3.1 | stress marks in `text_denorm` / `text_alignment` | MIT |
| voice-gender-classifier (ECAPA-TDNN) | `spk_desc.audio_desc.gender` | MIT |
| OpenAI `gpt-4.1-mini` (API) | `text_denorm` | OpenAI ToS |
| TalkNet-ASD | — (face selection only) | MIT |
| LVFace | `global_spk_id` | code MIT; weights: non-commercial research only |
| Qwen3.5-Flash (API) | `spk_desc.image_desc` | Alibaba Cloud API ToS |
For questions about the dataset, or to request removal of your material (as a rights holder or as a person
appearing in the recordings), contact us at aleksei.gusev@ncspeech.org.