Skip to content

KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback

Jul 2026 · arXiv.org · Vol abs/2607.20556 · 0 citations · 44 references
Computer Science

Abstract

In large-scale text analysis tasks, pre-trained language models are often used to embed text corpora for downstream analysis. However, such models may struggle to capture domain-specific semantics and adapting them typically requires large amounts of labeled data and technical expertise to implement training pipelines. Recent approaches have demonstrated how visual interactions in document projections can capture human feedback as training signals for model tuning. However, these methods operate on document-level feedback, which requires users to open and assess individual documents in order to provide effective feedback. In this paper, we propose KeySI, an interaction framework that enables feature-level feedback through keyword-based concept specification. Users specify feedback by organizing extracted keywords into groups representing concepts, which KeySI translates into document-level supervision for subsequent tuning. By operating on keywords as the primary interaction medium, KeySI reduces the need for manual document inspection and labeling and lowers the barrier to adapting embedding models. We present a prototype implementation that, given a corpus, curates representative keywords, visualizes keywords and document embeddings via dimensionality reduction, allows interactive specification of keyword groups, and supports iterative refinement through system feedback. We evaluate KeySI through a user study, usage scenarios, and quantitative experiments demonstrating its effectiveness in capturing user intent and improving embedding alignment.

View source

Similar papers

EviMap: Evidence-Grounded Hierarchical Topic Maps for Exploring Unlabeled Corpora

Research teams and organizations often explore unfamiliar free-text collections, from survey comments and reviews to reports and domain documents, before labels, queries or coding schemes exist. At this stage, the first thematic map shapes what users notice, prioritize and carry into downstream analysis, so it should b...

Zhiyin Tan, Changxu Duan · 0 citations
#artificial intelligence Preprint Sep 2026

Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?

Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-...

Jin-Hao Zhang, Ze-Yu Liu, Zi-Cheng Yan et al. · 0 citations
Review Oct 2026

Data and Knowledge Dual-Driven Text Embedding Approaches: A Survey

Text embedding has emerged as a pivotal technique in natural language processing, facilitating the effective understanding and processing of textual information by machines. With the continuous advancement of data-driven methods like large language models (LLMs), text embeddings have become richer and of higher quality...

Xiao-Yin Chen, Jia-Qing Zhan, Jia-Yi Lin et al. · 0 citations
#artificial intelligence Preprint Aug 2026

EidosDoc: Implicit Structure Encoding for Cost-Effective Semi-Structured Document QA

Semi-structured documents are ubiquitous in scientific reports, financial statements, and technical manuals. Question answering over such documents requires simultaneous understanding of text, tables, charts, and complex hierarchical layouts. Existing methods either rely on repeatedly calling large language models for...

Teng Lin, Yu-Yu Luo, Nan Tang · 0 citations
#artificial intelligence Preprint Sep 2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al. · 1 citation
Preprint Aug 2026

SAM3-LoRA: Parameter-Efficient Adaptation of a Concept-Promptable Foundation Model for Multi-Class Structural Defect Segmentation

Promptable segmentation foundation models such as SAM3 accept an open-vocabulary text concept and return every instance matching it, but adapting them to a specialized domain by full fine-tuning is computationally prohibitive for the organizations that would benefit most. This study applies Low-Rank Adaptation (LoRA) t...

P. Malaisree, S. Youwai, S. Janrungautai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.