Skip to content
Preprint

FunnelAudit: Responsibility Auditing in Multi-Route Recommender Systems

Sep 2026 · 0 citations · 43 references
Computer Science

TL;DR

Findings demonstrate the importance of explicit serving semantics and checkable witnesses for recommender accountability, and introduce FunnelAudit, an executable framework for incident-level responsibility auditing.

Abstract

Multi-route recommender systems combine retrieval, allocation, fusion, and ranking, making individual inclusions and exclusions difficult to audit. Route overlap can hide effects from one-at-a-time ablations, while freezing downstream stages produces counterfactuals inconsistent with serving behavior. We introduce FunnelAudit, an executable framework for incident-level responsibility auditing. An accountability contract specifies the disputed Top-K event, controls and owners, permitted reference actions, and replay semantics. FunnelAudit evaluates every permitted control configuration and applies graded actual responsibility to find the smallest outcome-preserving contingency that makes each control pivotal. Its certificate records the contingency and paired serving executions needed to verify the judgment. We instantiate the framework in two-stage, nine-route funnels using fixed union, weighted quota allocation, or weighted reciprocal-rank fusion, followed by SASRec ranking. Across 258,809 user-target incidents from three real interaction datasets, 4.24-16.24% admit a responsible control. Among responsible incident-control pairs, 92.55-99.64% require a nonempty contingency, so single-control ablation recovers only 0.36-7.45%. Policies differing in factual outcomes on only 0.31-2.39% of incidents yield 21.44-54.05% Jaccard distance between responsible-route sets on matched exclusions. Independent replay reproduces all 9,121,792 checked target-world outcomes; exhaustive search and a generic mixed-integer linear program agree with every sampled judgment. These findings demonstrate the importance of explicit serving semantics and checkable witnesses for recommender accountability.

View source

Similar papers

Preprint Aug 2026

PILOT Technical Report

PILOT (Proactive Insight Learner for Online Tree-Experiments), an LLM-agent framework that organizes three roles within a constrained control loop where deterministic services enforce all safety, statistical, and permission boundaries, is presented.

Jiuning Lin, Ruiquan Lan, Xiaodong Zhu et al. · 0 citations
Preprint Aug 2026

DREAM Technical Report

This work presents DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them, supporting agentic meta-control as a viable paradigm for industrial recomm...

Bin Zhang, Bo-Wen Zheng, Chao Yi et al. · 0 citations
Open access 2026

Hodge-Guided Active Preference Elicitation for Efficient Pairwise Ranking

Pairwise comparison is a standard method for eliciting user preferences in recommender systems and group decision-making. The number of required comparisons grows quadratically with the number of alternatives, creating a substantial burden for users and platforms. This paper asks whether Hodge decomposition, which spli...

Eskander Bejaoui, M. O. Aoueileyine, R. Bouallègue · 0 citations
#natural language process... Preprint Sep 2026

FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

FinFIRST is the first financial benchmark to jointly evaluate answers and supporting evidence through atomic rubrics, retaining final-answer correctness as the primary objective while making the supporting research process measurable, verifiable, and diagnosable.

Wen-Qing Wang, Hai-Tao Xiang, Xin-Yi Zhao et al. · 0 citations
Book Open access Sep 2026

TRACE: Targeted Ranking-Aware Counterfactual Explanation for Sequential Recommendation

Ranking-constraint counterfactual explanation for sequential recommendation requires query-limited search to decide where to edit and what to substitute—the bottlenecked for query efficiency lies more in how the search space is structured than in the mutation rate alone. We propose TRACE (Targeted Ranking-Aware Counter...

Ungsik Kim, Sang-Min Choi, Gun-Woo Kim et al. · 1 citation
Open access Sep 2026

TrustCompute: Result Validation, Reputation-Aware Scheduling and Quality-Based Payments for Outsourced LoRA Fine-Tuning

Outsourced model fine-tuning requires reliable task allocation, result validation, and incentives for quality. We present TrustCompute, a compute-sharing platform integrating two-stage adapter verification, reputation-aware scheduling, quality-based payments, and auditable ledger records for LoRA fine-tuning. Our evalu...

Lin-Fa Lee, Yi-Yu Chang, Kuo-Hui Yeh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.