Russian scientists improve neural network’s ability to engage in more natural dialogue

Researchers at St. Petersburg State University (SPbU), a partner of TV BRICS, have integrated an intonation model into a neural network, making the artificial intelligence’s (AI) pronunciation sound more natural and familiar to humans. This development will help improve text-to-speech systems used in voice assistants and humanoid robots. This was reported on the university’s
website.

Most often, language models reproduce the most common variants of word pronunciation or elements of a phrase. AI is largely unable to produce intonation that conveys semantic and emotional nuances. At the same time, an unfortunate combination of intonation patterns in different parts of a sentence can spoil the overall impression, even when the voice generation is of high quality. Russian linguists have attempted to rectify this shortcoming of neural networks. The modifications they proposed have enabled the neural network to overcome the unnaturalness or monotony of intonation.

The AI training was based on an intonation classification system developed by staff at the Department of Phonetics and Methods for Teaching Foreign Languages at St. Petersburg State University – Associate Professor Nina Volskaya and the Head of Department, Professor Pavel Skrelin. This methodology takes into account the type of sentence: a question, a narrative, a command, a request or an address. Furthermore, the system incorporates a wide range of intonation variants.

“Thanks to the correct structuring of the data and the use of labels for intonation variants, the model was able to convey various nuances in intonation, which allows for a more natural and expressive generated sound. ‘This approach makes the synthesis system more flexible and suitable for various applications, including voice control and educational programmes,” explained Ulyana Kochetkova, Associate Professor at St. Petersburg State University.

To assess the success of the neural network’s training, the researchers asked 42 native Russian speakers to answer questions after listening to the synthesised audio recordings. On average, the participants found the audio files to sound natural.

Linguists note that to further improve the quality of the synthesis, the model needs to be fine-tuned, taking into account emotional and stylistic variations.

 

 

Share your love

Leave a Reply