Repository logo
English
Türkçe
Log In(current)
  1. Home
  2. ADA University
  3. CB5. ADA Theses, Dissertations and Final Projects
  4. School of Information Technologies and Engineering
  5. Computing Infrastructure and Data Pipeline for Enterprise-scale Data Preparation: Scalability Optimization Study

Computing Infrastructure and Data Pipeline for Enterprise-scale Data Preparation: Scalability Optimization Study

Date Issued
2023-04
Author(s)
Akhund, Sadig
Abstract
In today's data-driven landscape, enterprises face significant challenges in managing and
processing massive amounts of data for meaningful insights and informed decision-making.
Data preparation, a critical process that converts raw data into a usable format, plays a pivotal
role in the data pipeline and significantly impacts downstream data analysis and modeling.
However, traditional data preparation methods may struggle to keep up with the increasing
volumes and complexity of data, leading to scalability issues, inefficiencies, delays, and
suboptimal performance in the data pipeline. This thesis presents a comprehensive scalability
optimization study that analyzes and optimizes the data preparation process in enterprise
grade data pipelines. The study begins by analyzing common components of data pipelines
and identifying limitations and bottlenecks that hinder scalability. It thoroughly examines
existing data preparation methods, tools, and technologies, as well as cutting-edge tools and
methodologies such as Apache Nifi, Apache Atlas, and Apache Spark for addressing
scalability challenges. The research draws insights from literature, industry practices, and
state-of-the-art technologies to propose practical strategies and recommendations for
designing a scalable data strategy in an enterprise setting. The study provides actionable
insights and recommendations to enhance the performance of data pipelines in enterprise
grade data environments. The paper concludes with a summary of key findings, limitations,
and future research directions, emphasizing the need for a well-designed data preparation
pipeline that incorporates scalable data ingestion, efficient data transformation, and
intelligent data storage strategies to ensure reliable and efficient data processing in
enterprises dealing with large volumes of data.
Subjects

Data processing -- Sc...

Big data -- Managemen...

Database management -...

IT and Engineering

Get Involved!
  • Source Code
  • Documentation
  • Slack Channel
Make it your own

DSpace-CRIS can be extensively configured to meet your needs. Decide which information need to be collected and available with fine-grained security. Start updating the theme to match your Institution's web identity.

Need professional help?

The original creators of DSpace-CRIS at 4Science can take your project to the next level, get in touch!

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Accessibility settings
  • Privacy policy
  • End User Agreement
  • Send Feedback
Repository logo COAR Notify