Skip to content
Open access

Scalable AutoML Frameworks for End-to-End Optimization of Data Science Workflows

Sep 2026 · Journal of Data Science · 0 citations · 24 references

TL;DR

This study proposes a scalable end-to-end AutoML framework that integrates automated preprocessing, feature engineering, model selection, hyperparameter optimization, and adaptive resource allocation within a unified optimization architecture to improve optimization efficiency while maintaining predictive robustness.

Abstract

The increasing complexity of machine learning workflows has created significant challenges in efficiently developing scalable and high-performing predictive models. Although Automated Machine Learning (AutoML) has emerged as a promising solution for reducing manual intervention in data science pipelines, existing approaches still suffer from limited scalability, high computational overhead, fragmented pipeline optimization, and insufficient resource-aware mechanisms. This study addresses these limitations by proposing a scalable end-to-end AutoML framework that integrates automated preprocessing, feature engineering, model selection, hyperparameter optimization, and adaptive resource allocation within a unified optimization architecture. The proposed methodology employs hierarchical search strategies, Bayesian and evolutionary optimization, and dynamic resource scheduling to improve optimization efficiency while maintaining predictive robustness. Experiments were conducted on medium- and large-scale datasets using repeated cross-validation and benchmark comparisons against conventional machine learning models and existing AutoML systems. The results demonstrate that the proposed framework achieved superior predictive performance, obtaining an accuracy of 0.93 ± 0.01 and an AUC of 0.96 ± 0.01, outperforming baseline approaches by approximately 8-12% across multiple evaluation metrics. In addition, the framework reduced optimization runtime by approximately 68% compared to existing AutoML systems while maintaining strong scalability across increasing dataset sizes. This research aims to advance scalable and efficient AutoML systems by providing a reproducible, resource-aware, and integrated framework for automating end-to-end data science workflows.

Read PDF

Similar papers

Open access Aug 2026

Learning from Prior Experiments: Meta-learning Models of Workflow Performance

This study shows that accurate performance prediction for classical machine learning workflows can be achieved through meta-learning using readily available OpenML meta-data, and indicates that careful selection of regression models is more critical than increased representational complexity.

Roman Neruda, Juan Carlos Figueroa–García, Carlos Franco · 0 citations
Book Open access Aug 2026

Automating End-to-End Hybrid Query Processing: Benchmark, Solution, and Insights

A large-scale benchmark with 60\sim 90× more queries than prior work, built on 3× more databases, an automated pipeline that can execute existing methods without manual intervention, and multi-dimensional, fine-grained evaluation metrics for comprehensive assessment.

Bo Li, Chenzhan Wang, Long-Kang Lin et al. · 0 citations
Open access 2025

An Ensemble-Based End-to-End Framework for Automated Resume Screening Using Machine Learning and Service-Oriented Architecture

This study presents an end-to-end framework for automated resume screening leveraging an ensemble of advanced machine learning techniques within a service-oriented architecture (SOA). The framework integrates a diverse set of predictive algorithms designed to evaluate resumes against job descriptions across various dom...

G. Mallikarjuna, M. Basha, A. S. Kumar et al. · 2 citations
Review Open access Sep 2026

Software engineering for reproducible pipeline development in bioinformatics

Reproducibility in bioinformatics remains challenging despite the availability of workflow management systems and mature computational infrastructures. This work presents a software-engineering perspective for developing reproducible bioinformatics pipelines, with emphasis on pipeline-specific code. We reinterpret the...

D. Pérez-Rodríguez, Alba Nogueira-Rodríguez, Jorge Vieira et al. · 0 citations
Open access Aug 2026

A Novel Iterative Machine Learning-Driven Framework for Reliable and Adaptive Cloud Data Migration

AMF-CloudForge is presented, a unified machine learning-driven framework that integrates migration state analysis, intelligent scheduling, consistency preservation, and real-time adaptive management into a single end-to-end architecture and transforms cloud data migration from a static, tool-centric process into a reli...

S. Sapate · 0 citations
Book Open access Aug 2026

BioFlowBench: A Comprehensive Benchmark for Evaluating Bioinformatics Tool-use Capabilities of LLMs and Agents

A significant gap exists between static QA and dynamic execution tasks, with top LLMs perform well on static QA but falter in real-world execution scenario; specialized agents outperform general models in real-world execution through environmental interaction and iterative refinement; and domain knowledge remains the p...

Yu-Fei Hou, Jiajia Wang, Ke Xiang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.