Repository logo
English
Türkçe
Log In(current)
  1. Home
  2. ADA University
  3. CB5. ADA Theses, Dissertations and Final Projects
  4. School of Information Technologies and Engineering
  5. Spelling Correction for Azerbaijani Language Using Sequence to Sequence Model

Spelling Correction for Azerbaijani Language Using Sequence to Sequence Model

Date Issued
2023-04
Author(s)
Dashdamirli, Asad
Abstract
In natural language processing (NLP), spelling correction is an essential task which seeks to automatically
correct misspelled words in text documents,. This thesis focuses on Azerbaijani language spelling
correction, which presents unique challenges due to its rich morphology and complex orthographic norms.
Beginning with a comprehensive literature review covering the extant approaches and techniques
for spelling correction in various languages, the thesis then proceeds to its methodology. We identify
the limitations of existing methods and propose a novel approach for Azerbaijani orthography
correction based on a sequence-to-sequence (seq2seq) deep neural network.
Our proposed method makes use of seq2seq models, which have demonstrated great success in a
variety of NLP tasks, to discover the mapping between misspelled words and their right counterparts.
In addition, we introduce techniques for generating artificial noise to augment the training data and
enhance the model’s ability to manage various types of misspellings.
We conduct extensive experiments on a large corpus of Azerbaijani text data in order to evaluate
the performance of our approach. We evaluate the results in terms of character error rate, word error
rate and sequence error rate by comparing our method to several other methods. Our experiments
demonstrate that our seq2seq-based approach reaches adequate results, 5.3% character error rate and
25.77% word error rate in text from news which shows potential to enhance the accuracy of
Azerbaijani text spelling correction.
In addition, we analyze the effect of artificial noise generation techniques on the performance of
our model and provide insights into how effective they are in managing various misspelling types. In
addition, we discuss the limitations of our methodology and possible future directions for further
development.
This thesis concludes with a novel approach to Azerbaijani spelling correction using deep neural
networks, specifically the seq2seq model, along with artificial noise generation techniques. The
experimental results demonstrate the viability of our method for enhancing the precision of
Azerbaijani text documents in real-world settings.
Subjects

Natural language proc...

Spelling correction -...

Azerbaijani language ...

IT and Engineering

Get Involved!
  • Source Code
  • Documentation
  • Slack Channel
Make it your own

DSpace-CRIS can be extensively configured to meet your needs. Decide which information need to be collected and available with fine-grained security. Start updating the theme to match your Institution's web identity.

Need professional help?

The original creators of DSpace-CRIS at 4Science can take your project to the next level, get in touch!

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Accessibility settings
  • Privacy policy
  • End User Agreement
  • Send Feedback
Repository logo COAR Notify