Skip to main content
Data in Brief logoLink to Data in Brief
. 2024 Jul 14;55:110753. doi: 10.1016/j.dib.2024.110753

Gigant-KTTS dataset: Towards building an extensive gigant dataset for Kurdish text-to-speech systems

Hawraz A Ahmad a, Tarik A Rashid b,
PMCID: PMC11324836  PMID: 39149720

Abstract

Today, speech synthesis is a part of our daily lives in computers all around the world. Central Kurdish Speech Corpus Construction is a speech corpus that is a primary data source for developing a speech system. There are still two main issues that prevent them from achieving the best possible performance, the lack of efficiency in training and analysis, and the difficulty in modelling. The biggest obstacle against text-to-speech in the Kurdish language is that there is a lack of text and speech recognition tools compounded by the fact that around 30 million people speak the Kurdish language in different countries. To address this issue, this corpus introduced a large vocabulary of Kurdish Text-to-Speech Dataset (KTTS, Gigant), including a pronunciation lexicon and speech corpus for the Central Kurdish dialect. A variety of subjects is comprised to record these sentences. The sentences are recorded in a voice recording studio by a Kurdish man who is a dubber. The goal of the speech corpus is to create a collection of sentences that accurately reflect the real data about the Central Kurdish dialect. A combination of audio and visual sources is used to record the 6,078 sentences of 12 document topics. They were recorded in a controlled environment using microphones that were not noisy. The total record duration is 13.63 h. The recorded sentences are in the “.wav” format.

Keywords: Central Kurdish language, Speech corpus, Speech system, Dataset, Deep learning, Classification segmentation, Detection


Specifications Table

Subject Kurdish TTS
Specific subject area The dataset consists of 6078 Central Kurdish Corpus divided into 12 categories, the details of these categories with how many sentences are, News (888), Sport(631), Health(463), Interview(1240), Science(65), Religion(24), Economic(275), General information(224), Politics(66), Education and literature(1399), Article(420) and Social(383).
Type of data Text and Audio, The central Kurdish writing system is like a phonemic system. Each letter of the language is assigned a phoneme, and some exceptions are made. For instance, the letters “ى” and “j” are both pronounced as “palatal,” while the “i” is a “vocet.” Similarly, the letters “و” and “w” are both pronounced as “bilabial,” respectively. The “و” is also repeated in a repeating form as “وو ‘” and is pronounced as “u/”.
Data collection The data was obtained from the environment using microphones that were not noisy. They were recorded in a controlled environment The total record duration is 13.63 h. The recorded sentences are in the “.wav” format.
Data source location Salahaddin University – Erbil
Data accessibility Repository name: Mendeley Data
Data identification number: 10.17632/zhnvwsd7hs.1
Direct link to the dataset: https://data.mendeley.com/datasets/zhnvwsd7hs/1

1. Value of the Data

Details of data and its values can be described as follows:

  • Central Kurdish Text influences Text-To-Speech generation. The dataset consists of 12 categories.

  • The data set helps researchers for early improving Kurdish TTS thereby reducing the time-consuming for this process.

  • The number of Texts in the dataset can be increased. It allows models to learn additional texts and audio from multiple different situations, to correctly classify new texts.

  • Deep learning and machine learning techniques can simulate the collected data and predict classification outcomes.

2. Background

People of the Kurdish ethnic group use their language, which is known as the Kurdish dialect, to communicate with each other in the Middle East. The two most common forms of this dialect are the Kurmanji and the Sorani [1,2]. For this purpose, linguists will use the Sorani dialect or Central Kurdish Language to collect data related to the Kurdish language. Although the characters used in these languages are similar, they have different Unicode designations. In Sorani Kurdish, 36 alphabets are composed of consonants and vowels. In Sorani Kurdish, the character (ی, و) is used as constants and vowels depending on the position of the word. For instance, in the word (Inline graphic) (gull), the word “Flower” is a vowel, while “و” in “Inline graphic” means “Game,” is a vowel [3,4]. The key objectives behind collecting this dataset can be summarized as follows:

  • 1)

    One major aim of establishing this dataset is to offer a significant resource for Central Kurdish speech data, enabling its employment in deep learning and machine learning models.

  • 2)

    This dataset can add to the expansion and operation of policies and procedures across various domains.

  • 3)

    Scholars, researchers, and students focusing on Kurdish studies can apply this data for a multitude of preprocessing tasks.

3. Data Description

The central Kurdish writing system is like a phonemic system. Each letter of the language is assigned a phoneme, and some exceptions are made. For instance, the letters “ى” and “j” are both pronounced as “palatal,” while the “i” is a “vocet.” Similarly, the letters “و” and “w” are both pronounced as “bilabial,” respectively. The “و” is also repeated in a repeating form as “وو ‘” and is pronounced as “u/”. The Central Kurdish writing system is similar to that used in Arabic and Persian writing, but there are differences. Table 1 shows the difference between the letters in Persian, Arabic, and Kurdish alphabets. These unique letters can be used to differentiate languages. In Persian and Arabic, the Kasre and homograph problems are usually present. On the other hand, in the Kurdish system, the corresponding letters for these problems are present in written texts. This means that these problems do not occur.

Table 1.

Letters of Kurdish in comparison with Persian and Arabic letters.

Language Letters
Kurdish only graphic file with name fx3.gif
Kurdish and Persian graphic file with name fx4.gif
Kurdish, Persian and Arabic graphic file with name fx5.gif

The sentences were collected from various types of news websites. The selection criteria focused on collecting sentences on different topics to ensure the dataset remained unbiased. As the initial pre-processing step, we reviewed the collected sentences to remove any informal or unsuitable content and corrected any misspelt words.

Table 2 reveals the various subjects covered by the chosen sentences. Fig. 1 also shows the meta dataset sentences when opened by Excel using a Unicode character set (UTF-8).

Table 2.

Examples of some sentences with their topics.

graphic file with name fx6.gif

Fig. 1.

Fig. 1

Metadata of Gigant Dataset opened by Excel.

The metadata of the Gigant dataset contains three Excel sheets. The first sheet is composed of seven columns. The first column is the ID of the sentence which is the name of the recorded wav file as well. The second column contains the alphanumeric sentences while the third column contains the same sentences with converting the numbers and dates to texts. The fourth column is the category or the topics of the sentences. The fifth column indicates the recording length in seconds. Furthermore, the sixth and seventh columns are sampling rate and quantization respectively as depicted in Fig. 1. The second sheet contains some details regarding the minimum, maximum, and average recording length.

The dataset corpus includes the “text, audio” pair dataset. It is being recorded in a studio using microphones that are designed to cancel out background noise. The sentences were recorded by a Kurdish man whose occupation is dubber.

These are some of the features that were utilized in the analysis of the audio files.

  • 1)

    The output of the program was recorded at a rate of 22,050 kHz.

  • 2)

    The quantization process was carried out using 16 bits of signed data.

  • 3)

    The output files are compressed into WAV files.

  • 4)

    The stored audio files are in the format known as PCM.

  • 5)

    A mono channel is utilized to record the audio streams.

The audio recording and editing process lasted for 10 days. It involved capturing over 6078 WAV files and over 13.63 h of recorded speech. Table 3 provides an overview of the data.

Table 3.

Overview of the data.

Name of Dataset Length File Format Sampling Rate The Number of Files Longest File Length Shortest File Length Average File Length
Gigant 13.63hrs *.wav 22,050 Hz 6078 16.781 s 0.502 s 8.076 s
*.txt (UTF-8)

The audio files are stored in wave format, while the text sentences are saved in an Excel file. The audio files are organized in a single folder. The audio file's name includes the extension names, while the transcript is the text of the speech referenced to the audio file with an ID which is the name of the audio file. The dataset where prepared in a manner to comply with Gaussian distribution to be more effective in training models avoiding bias in record length. A statistical figure of the dataset has been created to show more clarity on the number of audio records of nearly similar length recordings as depicted in Fig. 2.

Fig. 2.

Fig. 2

Statistics of Gigant Dataset.

4. Experimental Design, Materials and Methods

To collect the sentences from web sources, a crawler software was developed in Python to systematically browse the websites and gather the sentences. Following this, a manual pre-processing step was undertaken to remove unsuitable sentences and label the remaining ones. The processed sentences were then transferred to an Excel file to determine the number of categories and the distribution of sentences within each category. Random samples were subsequently selected from each category. For the recording process, we recruited a professional voice actor with access to a studio to ensure the consistency and quality of the recordings.

Due to the complexity of the Kurdish language, its script is different from the Sorani standard. For instance, some sources use ك instead of ک. Researchers are constantly looking into some subjects of people to improve their efficiency. One of the most important factors they consider when it comes to analyzing people's data is their social media interactions. In this study, we use the social application. However, to predict the right view of people using machines, a good dataset is required. The abundance of channels and websites that allow researchers to collect data has made it easy to analyze people's social media interactions. This study contains 6078 samples of Kurdish sentences and the corresponding audio, which fall under more than 12 classifications listed in Table 1. The samples were gathered from various areas and pages. The total recording duration is 13.63 h. The format of the recorded files is “.wav” format.

The dataset was not generated through a questionnaire or survey; instead, it was compiled by collecting random sentences from various Kurdish news websites. These collected sentences were subsequently categorized based on their contextual content. To mitigate the issue of overfitting, enhance model performance, achieve faster convergence, improve interpretability, and increase robustness, the Gaussian distribution method was employed in the selection of the recorded sentences. The sample collection method utilized was direct, unambiguous, and unbiased.

Limitations

In Kurdish, there are issues with Kasra, homographs, and automatic pronunciation. For instance, some words have the same letters but have different meanings like (Inline graphic Inline graphic).

Unlike other languages, such as Persian, Arabic, and English, these problems are not as common in Kurdish due to their one-to-one mapping between spoken and written terms. One of the most challenging issues that speech synthesizers face is the recognition of the prosodic elements in written text, such as length, stress, and intonation. In continuous speech, the features are influenced by various factors, including the artist's emotions and personality. Unfortunately, there is a lack of knowledge about the various prosodic features of written texts. As a result, many of them are constantly modified as speech is made.

Ethics Statement

The production of speech files was carried out by a dubber man whose name is “Azad Mahmood Hussein”, and related consent was obtained from him. It is worth mentioning that there are no other participants involved in producing the speech files.

CRediT Author Statement

Hawraz A. Ahmad: Data curation, Original draft preparation, Visualization, Conceptualization, Writing, Reviewing and Editing. Tarik A. Rashid: Supervision, Reviewing, and Editing.

Acknowledgments

We would like to express our sincere gratitude to Azad Mahmood Hussein for his invaluable support in recording the corpus.

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

Declaration of Competing Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Data Availability

References

  • 1.Saeed A., Rashid T., Mustafa A., Agha R., Shamsaldin A., Al-Salihi N. An evaluation of Reber stemmer with longest match stemmer technique in Kurdish Sorani text classification. Iran J. Comput. Sci. 2018;1:99–107. doi: 10.1007/s42044-018-0007-4. [DOI] [Google Scholar]
  • 2.Rashid T.A., Mustafa A.M., Saeed A.M. A robust categorization system for Kurdish Sorani text documents. Inf. Technol. J. 2017;16:27–34. [Google Scholar]
  • 3.Rashid T., Mustafa A., Saeed A. Advances in Internetworking, Data & Web Technologies. EIDWT 2017 Lecture Notes On Data Engineering and Communications Technologies. Springer; Cham, Wuhan: 2017. Automatic Kurdish text classification using KDC 4007 dataset. [Google Scholar]
  • 4.Saeed A., Rashid T., Mustafa A., Fattah P., Ismael B. Improving Kurdish web mining through tree data structure and Porter's Stemmer algorithms. UKH J. Sci. Eng. 2018;2:48–54. 10.25079/ukhjse.v2n1y2018.pp48-54 [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement


Articles from Data in Brief are provided here courtesy of Elsevier

RESOURCES