Skip to content
#generative ai Open access

Paper B: Field Baseline Results. Does presentation format affect blind entity retrieval and brand attribution in production generative engines?

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Results for the pre-registered field baseline deposited as version 1.0 of this record (DOI 10.5281/zenodo.22798718, 16 September 2026). This version reports the registered dependent variables, completes Section 9 of the pre-registration from the session logs, and deposits the evidence package. Across 40 confirmatory sessions run from the UK on 1 October 2026, neither engine retrieved any of the four target pages. The design covered four entities across ChatGPT and Gemini, with five fresh sessions for every engine-entity combination, using the registered URL-free prompt through the desktop web interfaces. retrieved was 0/5 in all eight cells, and retrieval_shown was 0/40. Brand attribution could not be estimated because no session retrieved an entity. A supplementary cross-day run on 2 October also returned 0/8. The combined result was 0/48, with a Wilson 95% upper bound of 7.4%. All four pages were confirmed as indexed on Google from the UK on 1 October using the registered site: check. During this test window, that indexation did not lead to blind retrieval on either consumer surface. With retrieval at zero throughout, the registered primary comparison (presentation format on attribution) cannot be estimated. The result supports a narrower statement: during this window, placing a new entity on a single indexed source did not lead ChatGPT or Gemini to retrieve it from a blind prompt. Engine behaviour differed in texture and is reported as observation, not mechanism, in line with Section 12 of the pre-registration. Blind second view: 8 of 40 sessions (20%) coded independently by the co-author, raw agreement 100%. Cohen's kappa is undefined because both coders produced zero variance; the registered halt condition (kappa below 0.60) was not triggered. Deviation log D1 to D7 is included in the results document. Nothing was reconstructed after the fact; where a field was not captured, the record says so. Files: paperB-results-v1.0.pdf and .md (results document); PaperB_field_scoring_v1.0.xlsx (completed scoring workbook, both coders); paperB_raw_confirmatory_2026-10-01_UK.txt (40 confirmatory and 8 void sessions, full response text); paperB_raw_supplementary_2026-10-02_UK.txt (8 sessions); paperB_raw_G2_gate_2026-09-24_UK.txt (24 gate sessions); paperB-prereg-v1.0.pdf and .md (the pre-registration as deposited in version 1.0); MANIFEST.md5. Competing interests: Julio Arévalo Piedra founded and operates KusiGEO, a commercial GEO-audit platform. This study did not test KusiGEO or any other commercial product, and no KusiGEO tooling was used in its design, deployment or measurement. Julio did not run any confirmatory sessions; his measurement role was limited to the blind second-view coding. Artur Ferreira operates The GEO Lab, an independent GEO research site, and declares no competing interest in the outcome. Drafting disclosure: the confirmatory sessions, gate checks, coding, verification and data handling were carried out by the authors. The text of the results document was drafted with AI assistance using Anthropic Claude, working from the session logs and the deposited pre-registration. The authors reviewed, corrected and finalised the document and take full responsibility for its contents.

View source

Similar papers

#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Open access Mar 2024

LLM-based agents for automating the enhancement of user story quality: An early report

The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.

Zheying Zhang, M. Rayhan, Tomas Herda et al. · 48 citations · ⚡4
#computer vision Review Mar 2024

System for systematic literature review using multiple AI agents: Concept and an empirical evaluation

This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.

Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al. · 44 citations · ⚡2
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations

Related blog posts

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.