Skip to content
Open access

DuaLoc: Dual-Encoder Bug Localization with Bug-Report-Conditioned Attention and Contrastive Learning

Sep 2026 · Electronics · Vol 15, pp. 3978 · 0 citations · 28 references

TL;DR

DuaLoc combines two pre-trained language models: UniXcoder for the semantic understanding of source code and GraphCodeBERT for awareness of data-flow structure and fine-tuned with a contrastive objective that shapes the embedding space around the localization task.

Abstract

Bug localization is the task of automatically identifying the source files responsible for a reported defect. It is a critical step in software maintenance that accelerates defect resolution. Information retrieval (IR) methods are simple and effective at exploiting historical signals such as bug-fixing recency and frequency, but they struggle to bridge the lexical gap between natural-language bug reports and programming-language identifiers. Recent work increasingly leverages pre-trained language models (PLMs) for code to close this gap. However, current PLM-based approaches still rely on a single code encoder that ignores program structure and aggregates function-level signals into file-level representations via uniform pooling. We propose a dual-encoder bug localization (DuaLoc) framework that jointly addresses these limitations. DuaLoc combines two pre-trained language models: UniXcoder for the semantic understanding of source code and GraphCodeBERT for awareness of data-flow structure. Both encoders are fine-tuned with a contrastive objective that shapes the embedding space around the localization task. A bug-report-conditioned attention mechanism then aggregates function embeddings into query-dependent file representations. The resulting neural similarity scores are then fused with classical IR features in a learning-to-rank model. DuaLoc outperforms representative classical and PLM-based baselines across most evaluation settings on a widely used benchmark of six open-source Java projects.

Read PDF

Similar papers

CodeBERT-SENet: Adaptive Syntax-Semantic Fusion via Gated Attention for Python Bug Detection and Localization

This research proposes the first end-to-end multi-task architecture that jointly detects Python syntax errors and localizes their exact line position by adaptively fusing CodeBERT’s semantic embeddings with handcrafted syntactic features via a lightweight gating mechanism, without relying on Abstract Syntax Trees.

Rozali Ilham, Bahtiar Imran, Hasan Basri et al. · 0 citations
#machine learning Preprint Sep 2026

From Codebase to Culprit (C2C): Reducing the Search Space for Bugs with Semantic Retrieval and Hierarchical Reinforcement Learning

We introduce C2C (From Codebase to Culprit), a framework for precise bug localization that progressively reduces the debugging search space across multiple levels of granularity: files, functions, and lines of code. To mirror developer's natural top-down debugging workflows, C2C integrates semantic retrieval and Hierar...

Ankur Garg, Corey Yang-Smith, Rishav Rishav et al. · 0 citations
Preprint Sep 2026

Detecting Argument-Swap Bugs Using Context-Enhanced Code Representations

Names of source code elements convey rich semantic information and have been widely used in software engineering tasks such as bug detection, code completion, type prediction, and code classification. Prior studies exploit lexical similarity between method arguments and formal parameter names to detect bugs caused by i...

Subrata Das, Ali Aman, Muhammad Asaduzzaman et al. · 0 citations
#software testing Open access Sep 2026

SeqRankFL: Sequence-Aware Ranking of LLM-Based Code Representations for Statement-Level Fault Localization

Software reliability is an important part of reliability assurance for complex engineering systems, and timely fault diagnosis supports safe and continuous operation. After a test failure, statement-level fault localization ranks source-code lines for early inspection. Representation-based approaches can operate withou...

D. An, Shi-Hai Wang, Bin Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning

LWVIC4Code is proposed, a non-contrastive representation learning approach specifically designed for Type-IV clone detection that achieves competitive or superior performance without negative samples, benefits from layer-wise supervision, and generalizes effectively from Python to other languages, particularly Java and...

Luciano Marchezan, Kévin Delcourt, Eugene Syriani et al. · 0 citations
Open access 2026

Fine-Tuning UniXcoder for Code Smell Detection in Java Projects

This study investigates the application of UniXcoder, a pre-trained transformer model for source code, to classify Java source code methods across multiple projects as smell or clean, with a particular focus on the Switch Statements smell, and confirms that pre-trained transformer models, particularly UniXcoder, are ca...

Hanson Prihantoro Putro, Umi Laili Yuhana, E. M. Yuniarno et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.