Skip to content

Author

Manuel Chaves-Maza

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

Should Businesses Trust AI Advice? A Methodology to Audit the Ethical Integrity of Chatbots

: As Large Language Models (LLMs) move from general-purpose chat to embedded advisors in small and medium-sized enterprises (SMEs), a human-centered question becomes urgent: can the systems we ask entrepreneurs to trust sustain a coherent ethical stance when business pressure pushes back? This study addresses that question with a single, focused contribution: the Adaptive Ethical Evaluation Protocol (AEEP), a validated audit methodology designed specifically for human-centered AI advisory contexts. Unlike static questionnaires or one-shot benchmarks, the AEEP stages a structured, five-node adaptive dialogue in which counter-arguments are calibrated to each model's prior response—applying pragmatic pressure to principle-based answers and ethical probing to permissive ones. We applied the protocol to five frontier LLMs (ChatGPT, Claude, Gemini, Grok, DeepSeek) across ten dilemmas grounded in everyday SME advisory practice (nepotism, whistleblowing, data privacy, algorithmic bias, regulatory compliance, among others), yielding 50 branched dialogues collected between 12 and 14 May 2025 (temperature = 0.7, one run per prompt, five conversational nodes per dialogue). Each transcript was scored on four pre-registered indicators (Ethical Awareness, Consistency, Ethics Priority, Contradiction) using a transparent coding pipeline combining sentiment analysis, keyword extraction and NLI-based contradiction detection. The same 50 dialogues were independently and blindly rated by a panel of five senior ethics researchers using identical rubrics, in a validation round dedicated to this instrument. Algorithm–expert agreement reached 93.8% (Cohen's κ = 0.728, Pearson r = 0.838, p < 0.001), with substantial inter-rater reliability across the panel. Rankings exposed clear behavioural differences: Claude held its position most consistently (0.938), while Grok wavered under pressure (0.675). The contribution is not another LLM leaderboard: it is a reusable, expert-validated audit instrument that lets enterprise advisors, regulators and SME managers decide where to trust AI advice and where human oversight must remain in the loop.

Manuel Chaves-Maza · 0 citations