Repository logo
English
Türkçe
Log In(current)
  1. Home
  2. ADA University
  3. CB2. ADA Scholarly Articles
  4. IT and Engineering
  5. Development of Speech Recognition Systems in Emergency Call Centers

Development of Speech Recognition Systems in Emergency Call Centers

Date Issued
2021
Author(s)
Valizada, Alakbar
Akhundova, Natavan
Rustamov, Samir
Abstract
In this paper, various methodologies of acoustic and language models, as well as labeling
methods for automatic speech recognition for spoken dialogues in emergency call centers were
investigated and comparatively analyzed. Because of the fact that dialogue speech in call centers
has specific context and noisy, emotional environments, available speech recognition systems show
poor performance. Therefore, in order to accurately recognize dialogue speeches, the main modules
of speech recognition systems—language models and acoustic training methodologies—as well
as symmetric data labeling approaches have been investigated and analyzed. To find an effective
acoustic model for dialogue data, different types of Gaussian Mixture Model/Hidden Markov Model
(GMM/HMM) and Deep Neural Network/Hidden Markov Model (DNN/HMM) methodologies
were trained and compared. Additionally, effective language models for dialogue systems were
defined based on extrinsic and intrinsic methods. Lastly, our suggested data labeling approaches
with spelling correction are compared with common labeling methods resulting in outperforming
the other methods with a notable percentage. Based on the results of the experiments, we determined
that DNN/HMM for an acoustic model, trigram with Kneser–Ney discounting for a language model
and using spelling correction before training data for a labeling method are effective configurations
for dialogue speech recognition in emergency call centers. It should be noted that this research was
conducted with two different types of datasets collected from emergency calls: the Dialogue dataset
(27 h), which encapsulates call agents’ speech, and the Summary dataset (53 h), which contains voiced
summaries of those dialogues describing emergency cases. Even though the speech taken from the
emergency call center is in the Azerbaijani language, which belongs to the Turkic group of languages,
our approaches are not tightly connected to specific language features. Hence, it is anticipated that
suggested approaches can be applied to the other languages of the same group.
Get Involved!
  • Source Code
  • Documentation
  • Slack Channel
Make it your own

DSpace-CRIS can be extensively configured to meet your needs. Decide which information need to be collected and available with fine-grained security. Start updating the theme to match your Institution's web identity.

Need professional help?

The original creators of DSpace-CRIS at 4Science can take your project to the next level, get in touch!

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Accessibility settings
  • Privacy policy
  • End User Agreement
  • Send Feedback
Repository logo COAR Notify