Machine learning systems often fail to reach optimal performance not because of inadequate model architectures, but due to poorly structured data processing pipelines hidden within the codebase. Data-centric refactoring aims to improve model quality through systematic restructuring of code elements responsible for data collection, preprocessing, transformation, validation, and feature engineering. This paper introduces a comprehensive taxonomy of data-centric refactoring strategies, investigates their application across ML-driven software projects, and evaluates their impact on model accuracy, robustness, maintainability, and reproducibility. By bridging software refactoring principles with data-centric AI practices, the proposed framework demonstrates that code-level improvements to data handling routines can yield substantial gains in model performance while reducing technical debt. Experimental results show that systematically refactoring data pipelines leads to more reliable features, reduced noise propagation, and improved generalization. The findings position data-centric refactoring as a key discipline for modern ML engineering, enabling scalable, interpretable, and production-ready models.
Fatou Diop· International Journal of Art...· 0 citations
AI has quickly ceased being a supporting computational device and has become a central decision-maker in various spheres of society, such as health care, finances, criminal justice, and government, education, and employment. Decision-making systems based on AI have an increasing impact on the outcomes that have significant ethical, legal, and social implications on society and individuals. Although these systems have efficacy, scalability, and objectiveness, they also introduce some essential ethical dilemmas associated with bias, fairness, transparency, accountability, privacy, autonomy, and social justice. The paper will be a systematic review of the ethical issues of AI-controlled decision-making in the society. The investigation is an interdisciplinary synthesis of the literature in the domain of computer science, philosophy, law, and social sciences in order to distinguish the major ethical hazards and novel normative structures. It analyzes the origin of algorithmic bias in data, structural model design and institutional conditions and how explainability and trust are endangered through the lack of transparency in complex machine learning models. Specific focus is made on the asymmetries of power that the AI implementation causes in which automated systems impact vulnerable and marginalized people unequally. It is suggested in the paper that a methodology of ethical governance based on principles of responsible AI should be structured, fairness-by-design, transparency, human-in-the-loop oversight, and constant impact assessment. The conceptual model of ethical risk assessment is presented to consider AI systems throughout its lifecycle, including data collection and after-deployment language. The findings underscore the fact that AI systems have ethical failures that are seldom technical but rather socio-technical, which necessitate interventions at the policy, organizational governance, and technical design levels. The paper highlights the necessity of ethical norms that are enforceable, interdisciplinary cooperation, and international harmonization of regulations to make sure that the decisions made by AI could be consistent with the basic human values. The paper ends by identifying the future research directions/decision-making, as well as determining the policy implications of transforming AI systems into trustworthy, accountable, and socially beneficial systems.
Fatou Diop· International Journal of Inn...· 0 citations