Credits
The voices, and who made them.
Every voice in Turing AI OS is generated on your own machine by an open model. Several of those models are published under licences that ask for credit, and this page is where it is given — properly, rather than as a line of small print. Some of it is longer than a licence strictly requires. That is deliberate: the work described here was mostly done by people who were not paid by us, and in two cases not paid at all.
- Kokoro — Apache-2.0, from the Kokoro-82M project. Its training set includes Koniwa (tnc), CC BY 3.0, and the SIWIS French Speech Synthesis Database, CC BY 4.0 — the source of the French voice.
- Piper — MIT, from the Piper project, with espeak-ng (GPL-3.0) as its phonemizer.
- LibriTTS — CC BY 4.0. The English narrative voice.
- Multilingual LibriSpeech (MLS) — CC BY 4.0. The German, French and Dutch voices.
-
ARTUR — the Slovenian voice is trained on studio recordings
from ASR database ARTUR 1.0, published in the
CLARIN.SI repository of Slovenian
language resources under CC BY-SA 4.0.
It is worth saying what that database is. ARTUR 1.0 is 1,067 hours of Slovenian speech, 884 of them transcribed, across four subsets — read speech, public speech, private speech and parliamentary speech — with around 1,000 speakers. It was made by the Faculty of Electrical Engineering and Computer Science at the University of Maribor, the Faculty of Electrical Engineering and the Faculty of Computer and Information Science at the University of Ljubljana, Alpineon d.o.o. and STA, under the Ministry of Culture's project Razvoj slovenščine v digitalnem okolju (C3340-20-278001). It carries twenty-seven named authors, and the first of them — who led the work — is izr. prof. dr. Darinka Verdonik, associate professor at UM FERI and the faculty's coordinator on that project. That is a national-scale, publicly funded, multi-year effort: recruiting and paying speakers, studio time, script design, transcription, annotation, quality control, and the consent and legal framework that makes it publishable at all. A Slovenian voice of this quality exists because that work was done, and paid for, by other people first. -
Peter Pišljar — the Slovenian voice model itself,
sl_SI-artur-medium, published under CC BY 4.0.
A speech database is not a voice. ARTUR was built to teach computers to recognise Slovenian, and turning it into something that can speak it is a separate job. Peter Pišljar did that job: he took the studio recordings of a single speaker out of the corpus, prepared them as a text-to-speech training set of some forty hours, and trained and published the model. He then contributed it to the Piper voices catalogue, which is where we get it from.
It is worth naming the rest of it, because a voice is only as good as the language work underneath it. He also publishedslovene_accentuator,slo_g2p_byt5andslo_g2p_norm_byt5— Slovenian accentuation and grapheme-to-phoneme models, which is the pronunciation groundwork that Slovenian speech synthesis needs and that Piper does not come with. We do not ship those models, but they are the same body of work, and one person doing all of it for a language of two million speakers deserves to be read about rather than inferred from a filename.
He did this on his own and gave it away. The Slovenian voice in Turing AI OS exists because of the institutions who made the recordings and because of one person who made them speak, and we would rather name both than let the second one disappear into a filename. - Public domain and CC0 — the remaining English and Flemish voices ask for no credit; they are named here for completeness.
The full third-party licence text, covering everything else in the system, ships on the machine itself and is in Settings under About. If you believe something here is credited wrongly or insufficiently, write to info@turing.london and we will put it right.