th cleaned speech corpus
This refresh provides a converged full train manifest, a frozen eval manifest, and a token-coverage subset of approximately 5000 hours. All three manifests reuse the existing validated FLAC-in-Parquet audio pool; no audio was copied, repacked, or transcoded for this refresh.
Corpus overview
| Variant | Rows | Hours |
|---|---|---|
| train | 11,756,895 | 13412.724 |
| coverage_5000h | 4,338,927 | 4999.983 |
| eval | 1,772 | 2.001 |
All audio locators are portable and have the form
audio/<subset>/<group>/part-N.parquet#row=N.
The release audit records exact input/output hashes, rejected audio mappings, language-tag validation, and zero train/eval audio-identity overlap.
from datasets import load_dataset
ds = load_dataset("WTForbes/th", "th-5000h", split="train")
- Downloads last month
- 28