LLM-Click Agreement: Harmonizing Implicit Feedback and Semantic Judgments for Enterprise Search
Message search and retrieval on enterprise collaboration platforms is challenging in ''eye-off'' environments, where explicit human relevance labels cannot be collected and engagement signals such as clicks and dwell times are noisy and behaviorally biased. Large Language Models (LLMs) offer an alternative source of semantic supervision, but models trained solely on LLM-derived labels often regress on engagement-based metrics. We present LLM-Click Agreement Labeling, an industrial-scale supervision strategy that retains only those query--message pairs where click-based labels and LLM-generated labels agree. This selective filtering reduces supervision noise and maintains a balance between user interaction patterns and semantic relevance. In a human-annotated pilot, agreement-based labels improved accuracy by +18% relative to click-only supervision. A worldwide A/B deployment further showed that the approach preserves traditional search quality while delivering statistically significant gains in conversational grounding (e.g., +1.4% CiteDCG, +5.7% Good Citation Count). These results highlight that improving supervision quality, rather than modifying model architecture, is the most effective lever for advancing retrieval performance in large-scale enterprise systems.