Chatting with TAIDE in Indigenous tongues
NCHC is also developing an AI corpus for Taiwanese Indigenous languages. Says Chang Chau-lyan: “I can proudly declare that as a humanitarian undertaking, this project is utterly unique in the world of AI development.” When this field of application was mentioned at an international conference, people were very enthusiastic, and praised Taiwan for according fair treatment to ethnic minorities and for contributing to the preservation of cultures.
According to NCHC research fellow Shiau Yi-haur, who heads the center’s Indigenous languages project: Over the past two years, the project has collected more than 61,000 items of Truku-language voice data, which add up to a combined length of roughly 125 hours, or five times the volume of the voice data contained in the Indigenous-languages voice corpus of the Indigenous Languages Research and Development Foundation (ILRDF). The project has also collected more than 71,000 items of Tsou-language voice data, adding up to a combined length of some 240 hours, or 13 times the volume of the ILRDF’s voice data.
Says Shiau Yi-haur: “We’re next going to work on the Puyuma language, including the Puyuma, Katratripulr, Makazaya and Kasavakan dialects. And there is Kavalan, the extremely endangered language of the Kavalan people. There are only a few more than 1,700 Kavalan left in all of Taiwan, and most of them can’t speak the Kavalan language.”