Skip to content

Category

software testing

591 papers

#small language model Open access Sep 2026

Digital Language Laboratories in Indian Secondary Education: Effectiveness, Pedagogical Integration, Implementation Challenges and Policy Alignment

Digital language laboratories have re-emerged in Indian school-education discussions as networked environments for listening, speaking, pronunciation, multimedia practice, formative feedback and digitally mediated interaction. Their educational value, however, cannot be inferred from the presence of computers, headsets or language software. This critical narrative review evaluates the effectiveness of digital language laboratories for secondary education, examines the pedagogical conditions under which they are most defensible, analyses implementation constraints in India, and situates laboratory adoption within the national policy and curriculum landscape. Literature from education, applied linguistics and educational technology was synthesised alongside India-specific empirical studies and official policy documents. The strongest international evidence indicates that technology-supported language learning can improve selected language outcomes, but effects are heterogeneous and depend on task design, learner engagement, feedback, teacher mediation and opportunities to transfer laboratory practice into authentic language use. Evidence directly testing language laboratories in Indian secondary schools is comparatively sparse, geographically concentrated and often methodologically weak, with small or single-site samples, short interventions and limited use of standardised outcome measures. Consequently, positive local findings should be treated as proof of feasibility rather than proof of scalable effectiveness. Indian policy strongly supports digital infrastructure, multilingual education, competency-oriented pedagogy and technology-enabled teaching, yet does not establish a dedicated national standard for language laboratories. The central implication is that a digital language laboratory should be treated as a pedagogical service model rather than a hardware project. Sustainable value requires curriculum-linked tasks, teacher professional learning, reliable maintenance, equitable access, multilingual and accessible content, valid assessment, and routine evaluation of use and outcomes. Future research should prioritise multi-state comparative designs, authentic oral-language measures, long-term transfer, implementation fidelity and cost-effectiveness.

N. Sylvester Raja, G. Kalaiyarasan · 0 citations
#artificial intelligence Open access Sep 2026

The Generation-Governance Impedance Mismatch: Protocol-Governed Systems in the AI Era

The increasing adoption of generative artificial intelligence (AI) in software development introduces a structural asymmetry: implementation generation now occurs at machine speed, while behavioral governance remains constrained by institutional deliberation. These processes are not merely mismatched in velocity; they are orthogonal in function. Generation produces executable artifacts. Governance establishes permissible behavior. This paper formalizes the generation-governance impedance mismatch as a structural property of AI-accelerated systems. We argue that conventional governance mechanisms—code review, testing, and audit processes—operate at the implementation layer and therefore cannot scale to match machine-speed generation without structural reform.

Bhash Ganti, Bhash Ganti · 0 citations
#artificial intelligence Open access Sep 2026

The Generation-Governance Impedance Mismatch: Protocol-Governed Systems in the AI Era

The increasing adoption of generative artificial intelligence (AI) in software development introduces a structural asymmetry: implementation generation now occurs at machine speed, while behavioral governance remains constrained by institutional deliberation. These processes are not merely mismatched in velocity; they are orthogonal in function. Generation produces executable artifacts. Governance establishes permissible behavior. This paper formalizes the generation-governance impedance mismatch as a structural property of AI-accelerated systems. We argue that conventional governance mechanisms—code review, testing, and audit processes—operate at the implementation layer and therefore cannot scale to match machine-speed generation without structural reform.

Bhash Ganti, Bhash Ganti · 0 citations

Multi-Agent Strategies for Bridging Programming Language Gaps in Code Generation

Real-world software systems are inherently multilingual, but current large language models are not equally consistent across programming languages. This mismatch limits code generation usefulness, especially for underrepresented languages. Existing approaches improve code generation through fine-tuning, multi-agent reasoning, or translation, but remain language-specific, assume existing source code, or rely on per-language test suites. We introduce XL-CoGen, a three-stage multilingual code-generation pipeline that starts from a natural-language specification and a shared test list to generate correct implementations across multiple target languages. XL-CoGen first validates direct generation by constructing and correcting the test harness; when it fails, it transfers through empirically selected intermediate languages and translates validated solutions; it then repairs the best candidate through diagnosis and minimal patching. This design requires neither target-language training nor language-specific test suites. Across two benchmarks and multiple LLMs, XL-CoGen consistently improves over direct generation, with the largest gains on low-performing languages. In our Rust fine-tuning case study, XL-CoGen outperforms the best fine-tuned baseline by 22 percentage points and improves challenging languages by up to 33 points on multilingual benchmarks. Ablation results show that transfer and repair are complementary: repair suffices on easier tasks, whereas transfer becomes more important as difficulty increases, especially for weak target languages.

Micheline Bénédicte Moumoula, Serge Lionel Nikiema, Albérick Euraste Djiré et al. · 0 citations
#software testing Open access Aug 2026

Pengaruh Integrated Marketing Communication terhadap Minat Nasabah pada Produk Pembiayaan Gadai Emas di Bank Riau Kepri Syariah KCP Siak Perawang

This study investigates the influence of Integrated Marketing Communications (IMC) on customer interest in Rahn (gold pawn) financing at Bank Kepri Syariah KCP Siak Perawang, Riau, utilizing Partial Least Squares (PLS) analysis. The research is motivated by the observation that marketing communication efforts for the Rahn product remain underdeveloped, despite its significant potential as a Sharia-compliant financial solution. Employing an associative quantitative research design, the study utilizes a purposive sampling method to select participants from the bank's customer base.Data were collected via a questionnaire developed from the indicators of Integrated Marketing Communication (IMC) and customer interest. The data were then analyzed using simple linear regression within the SPSS software environment. The analytical procedure included validity and reliability testing, classical assumption diagnostics, the determination of the coefficient of determination (R2), and partial hypothesis testing (t-test). These findings are expected to provide empirical evidence regarding the impact of IMC on customer interest, thereby offering Bank Riau Kepri Syariah KCP Siak Perawang actionable insights to develop more cohesive and effective marketing strategies to increase adoption of the Rahn financing product. Abstrak Penelitian ini bertujuan untuk menganalisis pengaruh Integrated Marketing Communication (IMC) terhadap minat nasabah dalam menggunakan pembiayaan Rahn di Bank Kepri Syariah KCP Siak Perawang, Riau, dengan menggunakan metode Partial Least Squares (PLS). Latar belakang studi ini didasari oleh kurang optimalnya sosialisasi produk Rahn melalui strategi pemasaran, padahal produk ini memiliki potensi besar sebagai instrumen pembiayaan syariah. Dengan pendekatan kuantitatif asosiatif, penelitian ini melibatkan sampel yang dipilih melalui teknik purposive sampling dari para nasabah di kantor cabang tersebut.Data penelitian dikumpulkan menggunakan instrumen kuesioner yang disusun berdasarkan indikator variabel Integrated Marketing Communication (IMC) dan minat nasabah. Analisis data dilakukan melalui metode regresi linear sederhana dengan tahapan yang meliputi pengujian validitas, reliabilitas, dan uji asumsi klasik menggunakan perangkat lunak SPSS. Uji hipotesis dilakukan dengan meninjau koefisien determinasi (R2) serta pengujian secara parsial (uji t). Hasil penelitian ini diharapkan dapat memberikan gambaran empiris terkait pengaruh IMC terhadap minat nasabah, serta menjadi acuan bagi Bank Riau Kepri Syariah KCP Siak Perawang dalam menyusun strategi komunikasi pemasaran yang lebih efektif untuk meningkatkan ketertarikan masyarakat terhadap pembiayaan Rahn. Kata kunci: Integrated Marketing Communication, Minat Nasabah, Pembiayaan Rahn.

Devina Devina · 0 citations
#software testing Open access Aug 2026

Optimasi Pemanfaatan Ruang Parkir Pinggir Jalan Berbasis Internet of Things (IoT) untuk Mendukung Smart City

IoT utilization holds significant potential to support the digital transformation of efficient, transparent, and sustainable urban transportation management and empirical insights showing that complex urban parking occupancy relies on real-time multi-variable data integration beyond simple traffic flow are provided.

Muhammad Ghrandiaz Tengku Idris, Adi Widiantono · 0 citations
#software testing Open access Aug 2026

Volatile Organic Compounds as In Vitro Biomarkers of Leishmania major

The in vitro VOC profile of Leishmania major represents a metabolic fingerprint of parasite activity and highlights candidate biomarkers with diagnostic potential and support further investigation of VOC-based tools for automated culture detection and noninvasive diagnosis in resource-limited settings.

Alex Green, A. Zabala, Phillip Scott et al. · 0 citations

From Tool Use to Technological Agency: LoopCAT as a Local-First, Open-Source Tool for Translation Technology Education

Translation students need to learn both how to use translation technologies and how to judge the choices those technologies make available. This article presents LoopCAT, an Apache-2.0-licensed, local-first computer-assisted translation environment co-created with OpenAI Codex using GPT-5.5 and GPT-5.6, and proposes a framework connecting workflow competence, evaluative judgement, and technological agency. The account draws on repository history, implementation inspection, and the verification records of an identified development build. LoopCAT combines local project storage, translation memories, terminology, quality assurance, document exchange, and optional connections to local or hosted AI services. Its English, Catalan, and Turkish interface catalogs also make the application itself available as teaching material: students can translate English UI strings into another language, review the existing automatically generated target drafts, import their revisions, and test the interface. We organize these opportunities around four forms of participation: operating a workflow, evaluating outputs, inspecting and configuring mechanisms, and making or defending a bounded intervention. A six-session sequence, a UI-localization assignment, a placeholder example, and an assessment rubric specify how teachers could use the framework. The paper separates implemented capabilities from proposed educational benefits; it reports no new student-learning outcomes. It distinguishes the latest package checks from earlier regression evidence and sets out a protocol for classroom evaluation. LoopCAT provides an inspectable setting for teaching how translation decisions interact with data, interfaces, and software rules. Whether these activities improve judgement, transfer, or participation remains an empirical question.

Gökhan Doğru, Adrià Martín Mor · 0 citations
#artificial intelligence Preprint Sep 2026

Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation

Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-benchmark performance, but are largely result-only: they fit historical pass/fail response matrices or static task semantics, discarding how agents solve problems. We propose PTA-IRT, a Privileged Trajectory-Aware Item Response Theory framework that fuses process and outcome signals. Historical execution trajectories supply process-level evidence beyond pass/fail, such as explored context, attempted edits, and solving paths, which PTA-IRT uses as privileged information for calibration subset selection and ability estimation. Under low calibration budgets, PTA-IRT consistently outperforms prior IRT baselines on score and ranking recovery across four SWE benchmarks. Code and data are publicly available at https://github.com/DeepSoftwareAnalytics/PTA-IRT.

Kefeng Duan, Dewu Zheng, Yan-Lin Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers

Agent systems are commonly described by the model and harness that currently produce their behavior. That boundary is useful for one execution but underspecifies a long-lived agent that may change models, orchestration harnesses, interaction sessions, and host servers while retaining one identity, memory, and executable code lineage. We present a runtime-independent architecture for persistent agents. A continuity-bearing substrate $P_t=(I_t,M_t,B_t)$ contains an architectural identity representation, private durable memory, and a versioned software body. A replaceable deployment binding comprises an execution substrate $E_t=(R_t,H_t,D_t)$, which supplies a reasoner, harness, and host, and a set of interaction surfaces $S_t$, such as chat, API, or user interface bindings. A deployed execution is $A_t=P_t\triangleright(E_t,S_t)$; changing either replaceable layer is migration, not agent creation, when an authorized protocol preserves attributable lineage and transfers continuation authority within a governed deployment boundary. We define six continuity invariants and a quiesce--checkpoint--validate--bind--rehydrate--resume protocol. Enoch realizes the design as a reusable body plus private installed identity, memory, workflow state, and continuation authority, with infrastructure dependencies behind versioned provider contracts. A clean-room run of the frozen public commit passes 833 core tests and 92 provider and library tests executed separately from the core suite; deployments have exercised reasoner-version, interaction-surface, and host-machine substitutions while retaining continuity-bearing state. This evidence supports mechanical substitutability and authorized system continuity, not behavioral invariance or exhaustive pairwise evaluation. The downstream measurement question is whether an authorized continuation still recalls, composes, and enacts its identity.

Zhe Zhao, Roy Zhao Independent Researcher, P. G. A. S. O. C. ScienceEngineering et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

This paper studies autonomous software development, in which LLM-based coding agents transform high-level requirements into complete, functional, and usable software systems without human intervention. We introduce Harness-of-Harness (HoH), a framework that enables coding agents to continually improve software during autonomous development. HoH operates on existing coding-agent harnesses, and organizes their executions into iterative planning-coding-testing loops. To sustain improvement across loops, HoH balances repair with capability growth, scopes development into small and verifiable increments, separates implementation-time testing from independent evaluation, and constrains verifiable outputs rather than prescribing agent workflows. It progressively exposes deliverables, role-specific tools, and skills, encourages reuse rather than recreation, and maintains versioned project histories. On GameCraft-Bench, FrontierSWE, and ProgramBench, three harness-model pairs (Codex with GPT-5.5, OpenCode with DeepSeek-V4-Pro, and Pi with MiniMax-M3), HoH consistently outperforms the corresponding standalone harnesses, achieving an average relative gain of 52.25 percent and a maximum gain of 82.86 percent after three iterations. In a multi-day deployment with more than 70 iterations, HoH autonomously develops a first-person-shooter game, featuring a coherent storyline, fully implemented core mechanics, human-playable experience, polished visuals and integrated audio. Github: https://github.com/Flesymeb/HarnessOfHarness Project Page: https://flesymeb.github.io/HarnessOfHarness/

Hao Yan, Min-Le Su, Hangfan Zhang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

AcrossWAM1.0:A Modular Latent World-Action Stack for Compact Robot Policies

Latent world-action models avoid rendering future pixels by predicting an action-relevant visual subgoal in feature space. LaWAM established this formulation, but its original presentation left the world model, multimodal backbone, and deployment checkpoint tightly coupled. We introduce AcrossWAM1.0, a modularization and scaling study of this latent world-action stack. Rather than presenting latent subgoals as a new algorithm, we make the module boundary explicit: a policy adapter produces latent-action and action-generation contexts; a retained latent world decoder grounds the predicted transition in the current scene;and a flow-matching expert generates continuous action chunks. We further separate training-only teachers from the inference graph and provide a verifiable deployment export. On 2,000 paired LIBERO episodes, replacing a Qwen3-VL-2B backbone with Qwen3.5-0.8B yields 97.45% success versus 98.00% for the 2B model (a-0.55percentage-point difference; exact McNemarp=0.266). This does not prove equivalence, but it meets a prespecified two-point retention criterion. The compact, inference-reachable checkpoint contains 1,472.6M unique parameters, 42.4% fewer than the original 2B policy, while all retained tensors are bitwise identical to the source checkpoint. Cross-family execution is additionally checked with a MiniCPM-V adapter smoke test; closed-loop cross-family transfer remains an open evaluation. AcrossWAM1.0 therefore contributes an auditable software and evaluation boundary for compact latent world-action policies, distinct from LaWAM's original latent-subgoal contribution.

Yafen Zhang, Nan Wu · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.