Skip to content

Rethinking Data Quality for AI-Driven Systems: Evidence from Practitioner Interviews

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

Lifecycle assurance is synthesized through lifecycle assurance: a conceptual framing focused on producing evidence that data can support a specific AI claim when its influence may be embedded in model behavior, model-based judgments, or agent actions.

Abstract

Data quality research has usually treated data as an input that is stored, processed, and validated. In AI-driven software-intensive systems, data also shapes model behavior, evaluation, and lawful use. Empirical evidence remains limited on how practitioners define, assess, and manage quality under these conditions. We interviewed 16 practitioners from nine organizations and analyzed the transcripts using reflexive thematic analysis and developed six themes from participants'accounts. In AI systems, traceability shifted from modular debugging to attributing model behavior, while using models as quality assessors introduced circularity. Agent context and memory became data objects, and synthetic and pseudo-labeled data made authenticity a quality concern. In foundation-model development, lawfulness became a gate for training data, while representativeness was judged through coverage of situations in which the system must behave safely. Prior ML research examines many of these problems separately. Our study provides a practitioner-grounded account of how they are encountered together as an engineering and organizational concern. We also interpret five recurring conditions as helping explain how the themes relate to reduced trust in data and AI outcomes. We synthesize these findings through lifecycle assurance: a conceptual framing focused on producing evidence that data can support a specific AI claim when its influence may be embedded in model behavior, model-based judgments, or agent actions.

View source

Similar papers

Book Open access Aug 2026

Scaffolded AI-Verification: Assessment Patterns for Resource-Constrained Environments.

The widespread availability of generative tools has weakened a long-standing assumption in computing education: that the production of working code can serve as a proxy for student competence. In resource-constrained settings, these tensions are compounded by intermittent power, high data costs, and emergent institutio...

Kehinde D. Aruleba, Kike Ladipo, I. Sanusi et al. · 0 citations
Open access Aug 2026

Beyond Automation: Prompt Design and Trustworthiness in AI-Assisted Inductive Coding

This study introduces a structured prompt framework for AI-assisted inductive qualitative data analysis, demonstrating how carefully designed prompts can guide each stage of the coding process while preserving methodological rigor. A distinctive contribution of the study is that it provides a set of empirically tested,...

Yılmaz Sağlam · 1 citation
#diffusion models Review Open access Oct 2026

From Automated Coding to Qualitative Intelligence: A Human-Governed AI Model for Interpretive Research

The model is based on the principles of interpretivist epistemology, construct-validity theory, and human-in-the-loop (HITL) AI principles and redefines AI as an enhancement tool and not an autonomous interpreter, which sets a conceptual background to future empirical verification of AI systems run by humans.

E. Oladunmoye, M. A. Adewusi, L. O. Oyedele · 0 citations
Review Sep 2026

Helpful but Fallible: Developer Experiences of AI Tools Under a Coordinated Industrial Roll-out

AI-enabled software development tools (AI-devtools) are being industrially adopted under strong expectations of productivity gains, yet developers'experiences of such roll-outs are underexplored. Organizations commit budgets, evaluate staff, and revise practice on a partial picture, since the evidence base is mainly to...

Andreas Bexell, Rushali Gupta, L. Heander et al. · 0 citations
#artificial intelligence Review Sep 2026

Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use

A structured, seeded review of 24 focal empirical publications, finding no validated individual-level instrument in the focal corpus that tests the full combination of agent scope, permissions, recovery, state isolation, independent review, and evidence-based closure.

D. Véri · 0 citations
#artificial intelligence Open access Aug 2026

The Vibe Shift in Software Engineering: Evaluating AI-Led Conversational Programming for Performance, Cognition, and Responsible Adoption

This study evaluates Vibe Coding, an emerging AI-led conversational programming paradigm that enables developers to generate software through natural-language interaction with large language models (LLMs). Using a mixed-methods design, the study assessed performance efficiency, cognitive implications, and responsible a...

Sales Aribe Jr. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.