Skip to content

Category

large language models

513 papers

#large language models Open access Sep 2026

Autonomous Trust: Self-Gating Evaluation as a Prerequisite for Agent-to-Agent Communication at Scale

As autonomous agents backed by large language models (LLMs) move from single-user assistants toward peer systems that transact directly with one another, the volume of agent-to-agent (A2A) communication is projected to reach a scale comparable to today’s host-to-host network traffic. At that scale, no human reviewer can vet each outbound action, yet LLM output quality is non-stationary: a single model can produce an excellent decision at one turn and an unsafe or incoherent one at the next, with no guarantee tied to prior behavior. This paper argues that trust in such systems cannot be modeled on human trust, which is accumulated through track record, nor can it be delegated to the LLM itself, since the model is fundamentally an input-output function with no internal mechanism for self-policing. We propose Autonomous Trust, a framework in which trustworthiness is engineered as a property of the agent as a whole (LLM plus an external evaluation layer) rather than of the model in isolation. The framework rests on four design principles: (1) pre-transmission self-gating, in which the sending agent, not only the receiver, is responsible for intercepting its own unsafe outputs before they leave the system; (2) a verifiability taxonomy that routes each action to an appropriate gating strategy, from deterministic checks to LLM-panel adjudication; (3) continuous self-reported quality telemetry, analogous to application health metrics, that exposes an agent’s own output degradation to external monitoring in real time; and (4) explicit treatment of the gameable verifier" failure mode, in which self-evolving agents can degrade their own automated checks (for example, by authoring tests engineered to always pass). We formulate the design space, analyze failure modes including judge-panel correlated blind spots, and outline an empirical evaluation plan. This work contributes a concrete engineering framework, not a purely theoretical trust model, toward the near-term infrastructural challenge of building safe, unsupervised, internet-scale agent ecosystems.

Nehal Sangoi · 0 citations
#large language models Open access Sep 2026

Does Becoming Exceptional at One Domain Reduce Transfer Elsewhere

Foundation models derive their value from broad general capability across domains, yet deployment usually rewards specialization - creating a fundamental question for general intelligence: when performance is pushed upward in one region of capability space, is competence elsewhere conserved, redistributed, or destroyed? Synthesizing evidence from continual learning, transfer learning, multi-task optimization, parameter-efficient adaptation, model merging, vision-language adaptation, alignment, and 2023–2026 large-language-model studies, we find that the evidence rejects a universal specialization tax: domain-adaptive pretraining can produce positive transfer, whereas sequential fine-tuning, narrow supervised adaptation, and conflicting objectives can cause catastrophic forgetting, feature distortion, degraded zero-shot transfer, weakened instruction following, or loss of safety behavior depending on task relatedness, update locality, data mixture, optimization geometry, and model capacity. We propose the Generality-Specialization Frontier (GSF), a deployment-oriented framework that treats specialization gain and transfer retention as a Pareto problem, introducing distance-stratified transfer evaluation, invariant retention tests, worst-case regression reporting, and a normalized transfer-elasticity measure. We further propose G-S Bench, an evaluation protocol comparing full fine-tuning, replay, parameter-efficient updates, modular routing, weight interpolation, and non-parametric alternatives under matched target gains, concluding that becoming exceptional at one domain reduces transfer elsewhere only when specialization overwrites shared representations faster than the system preserves broadly useful structure - meaning the tradeoff is an architectural and optimization choice rather than an inevitable law of intelligence.

Sahir Maharaj · 0 citations
#large language models Open access Aug 2026

Using large language models to investigate patients' and caregivers' perceptions on SUDEP: A case study with an online epilepsy population.

BACKGROUND Sudden Unexpected Death in Epilepsy (SUDEP) is a leading cause of epilepsy-related mortality, yet remains under-communicated in clinical practice. Social Media Listening (SML) is a novel method using natural language processing and machine learning to retrieve real-world data. This study uses SML and explores patient and caregiver reports surrounding SUDEP on topics such as information provision, emotional impact, and preventive behavior. METHODS A retrospective observational study was conducted using Artificial Intelligence (AI) and Natural Language Processing (NLP)-powered patient-centricity solutions across online communities from 09/2020-11/2024. Posts from 23,584 authors were analyzed, with 1,381 posts by 789 individuals explicitly mentioning SUDEP. Patient and caregiver narratives were annotated, semantically tagged, analyzed quantitatively and qualitatively using Pharos AnalyticsTM and PatientGPT. RESULTS Most patients and caregivers reported learning about SUDEP through independent online research, often years after diagnosis, triggering emotions such as fear, shock, frustration. Lack of counseling by healthcare professionals was linked to feelings of betrayal and mistrust. Knowledge about SUDEP promoted adherence to medication, lifestyle adjustments, and use of preventive tools (e.g., seizure alarms, anti-suffocation pillows). Patients emphasized the need for early, transparent SUDEP discussions, ideally at diagnosis of epilepsy; caregivers focused on monitoring and care burden. Counseling was not associated with long-term psychological harm and welcomed for its empowering potential. CONCLUSION Despite clinicians' reluctance to discuss SUDEP due to fear of increasing patient anxiety, this study shows that knowledge about SUDEP may lead to proactive risk management without long-term emotional harm. Early, honest communication is desired by patients and caregivers and vital for implementing preventive strategies.

Nicole Brazda, T. Andreu, Philipp Cimiano et al. · 0 citations
#large language models Open access Aug 2026

A New Architecture of Agent

According to the drawbacks of large language model, I design a new agent which can overcome the drawbacks. The agent contains four deep neural networks, which is abstract network, concrete network, decision network, execution network. Each of the four networks just do one thing, so it focus what it can do just like brain works. The input is perception information by sensor. The output1 are classes and attributes(common sense), the output2 is memory or consciousness, the output3 is logic and theory, the output4 is action and practice. The No.1 and No.2 constitute an auto-encoder. The dimensions of classes are very high if the grain size is very small when one-hot encoding used, so binary encoding can be used. With the help of muti-level classes and sentence structure, output1 is produced as one sentence and common sense. With the help of two order dimensions, the output3 is a sentence, so the reasoning speed is more higher than LLM. The probability of sentence is joint probability of words, so only high probability of words can produce, so the hallucination problem alleviates. With the help of select gate tanh, when n is small, the computation time is one half of self-attention. When n is large, the time decreases much more. With the help of multi-value functions, the independent consciousness is produced. It is initiative, not rely on prompt. Causal reasoning and Continuous reasoning are realized by concatenate output3 with the input of No.3 as the new input of No.3.

Jinxin Wei, Zhe Hou · 0 citations
#large language models Open access Sep 2026

ACADEMIC PERMISSIBILITY IN LLM-ASSISTED EFL WRITING: HOW CONTEXTUAL JUSTIFICATIONS SHAPE STUDENTS' MORAL JUDGMENTS IN A WITHIN-SUBJECTS VIGNETTE SURVEY

AbstractThis paper investigates EFL students’ perceptions of ethical acceptability judgments of using Large Language Models (LLMs) into academic writing. In contrast to the simplistic view of acceptable/unacceptable use of LLMs, the present study models how specific contextual justifications shape students’ moral evaluations of LLM-assisted writing. A cross-sectional within-subject design was used with 220 third-year EFL students at three public universities in Laghouat, Algeria, to rate the ethical acceptability of LLMs use to complete four writing assignments. Under a neutral baseline condition, participants assessed the use of LLM for completing the four assignments, as well as four single condition contexts (disclosure, accuracy verification, syllabus permission, and learning intent). Their ratings were analyzed using Δ-effect scores (conditional minus baseline) to quantify condition effects above baseline and to describe task-level differences in ethical acceptability. Contextual conditions increased ethical acceptability to different extents, with learning intent producing the largest positive shift (ΔM ≈ 1.6, d ≈ 0.7) and disclosure exerting only a small, non-robust effect (ΔM ≈ 0.2, d ≈ 0.1). The baseline ratings also followed a clear gradient, which saw AI-assisted proofreading as the most acceptable while paraphrasing was consistently least acceptable. Theoretically, the study adds value by supporting a conditional-ethics perspective because it demonstrates how students’ moral considerations of LLMs use are dependent on learning-oriented and epistemically responsible frameworks rather than procedural cues like disclosure or syllabus permissions with no behavioral consequences. Practically, the paper advices that AI policies and pedagogy should be designed to encourage learning intent and verification activities as opposed to disclosure as a standalone requirement.Keywords: Academic integrity; large language models (LLMs); EFL writing; conditional ethics

Mohamed SEDDIKI, Souhila Korichi · 0 citations
#large language models Open access Sep 2026

ACADEMIC PERMISSIBILITY IN LLM-ASSISTED EFL WRITING: HOW CONTEXTUAL JUSTIFICATIONS SHAPE STUDENTS' MORAL JUDGMENTS IN A WITHIN-SUBJECTS VIGNETTE SURVEY

AbstractThis paper investigates EFL students’ perceptions of ethical acceptability judgments of using Large Language Models (LLMs) into academic writing. In contrast to the simplistic view of acceptable/unacceptable use of LLMs, the present study models how specific contextual justifications shape students’ moral evaluations of LLM-assisted writing. A cross-sectional within-subject design was used with 220 third-year EFL students at three public universities in Laghouat, Algeria, to rate the ethical acceptability of LLMs use to complete four writing assignments. Under a neutral baseline condition, participants assessed the use of LLM for completing the four assignments, as well as four single condition contexts (disclosure, accuracy verification, syllabus permission, and learning intent). Their ratings were analyzed using Δ-effect scores (conditional minus baseline) to quantify condition effects above baseline and to describe task-level differences in ethical acceptability. Contextual conditions increased ethical acceptability to different extents, with learning intent producing the largest positive shift (ΔM ≈ 1.6, d ≈ 0.7) and disclosure exerting only a small, non-robust effect (ΔM ≈ 0.2, d ≈ 0.1). The baseline ratings also followed a clear gradient, which saw AI-assisted proofreading as the most acceptable while paraphrasing was consistently least acceptable. Theoretically, the study adds value by supporting a conditional-ethics perspective because it demonstrates how students’ moral considerations of LLMs use are dependent on learning-oriented and epistemically responsible frameworks rather than procedural cues like disclosure or syllabus permissions with no behavioral consequences. Practically, the paper advices that AI policies and pedagogy should be designed to encourage learning intent and verification activities as opposed to disclosure as a standalone requirement.Keywords: Academic integrity; large language models (LLMs); EFL writing; conditional ethics

Mohamed SEDDIKI, Souhila Korichi · 0 citations
#computer vision Preprint Feb 2024

CodePori: Large-Scale System for Autonomous Software Development Using Multi-Agent Technology

Context: LLM-based multi-agent systems enable automation and decision support in software development, yet existing studies rely on benchmark datasets offering only binary pass-or-fail results, limiting insight into real-world applicability. Objective: This study empirically investigates the potential and limitations of LLM-based agents in autonomous software development tasks. Method: A two-phase approach was employed: developing a multi-agent system, CodePori, for automated code generation, and conducting participant-based evaluation to assess practical performance. Results: Participant feedback reveals key strengths, challenges, and areas for improvement in LLM-based multi-agent systems, highlighting aspects missed by standard code-generation benchmarks. Conclusions: While LLM-based multi-agent systems show potential for large-scale software development, successful integration requires addressing challenges such as memory limitations, hallucinations, and code smells, alongside a practitioner-centric perspective.

Z. Rasheed, Muhammad Waseem, Kai-Kristian Kemell et al. · 31 citations
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

Context: Manual qualitative data analysis is time-intensive and can compromise validity and replicability, affecting analysis design, implementation, and reporting. Large Language Models (LLMs) enable human-bot collaboration in Software Engineering (SE), but their potential for qualitative data analysis in SE remains largely unexplored. Objective: The objective of this study is to design and develop an LLM-based multi-agent system that synergizes human decision support with AI to automate various qualitative data analysis approaches. Methods: We used LLM-based multi-agents systems to assist the qualitative data analysis process, deploying 27 agents, each responsible for a specific task, such as text summarization, initial code generation, and extracting themes and patterns. Results: The main findings are: (1) the LLM-based multi-agent system accelerates the qualitative data analysis process, (2) the system effectively automates tasks such as text summarization, initial code generation, and theme extraction, and (3) the publicly accessible code facilitates validation and further evaluation. Conclusion: The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners. Future improvements focus on enhancing multilingual performance and integrating continuous expert feedback. The source code of proposed system and system details can be found here: https://github.com/GPT-Laboratory/Qualitative-Analysis-with-an-LLM-Based-Agentts

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 40 citations
#computer vision Review Feb 2026

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review

Large Language Models (LLMs) have enabled multi-agent systems to perform autonomous code generation for complex tasks. Despite the recent growth in research and industrial applications in this area, there is little work on synthesizing evidence from both academic and industrial sources to capture the current state of research on LLM-based multi-agent systems for code generation. To this end, we conducted a Multi-Vocal Literature Review (MLR), combining insights from both academia and industry, including peer-reviewed studies and grey literature. The aim of this study is to systematically synthesize and analyze existing knowledge on LLM-based multi-agent systems for code generation. Specifically, the review examines the motivations for their use, employed benchmarks and models, key challenges, proposed solutions, and potential directions for future research. We selected and reviewed 114 studies, and the key findings are: 1) the identified reasons for adopting multi-agent systems for code generation were classified into nine categories; 2) the models and evaluation benchmarks utilized across the studies were systematically analyzed to provide a structured overview of commonly adopted LLM configurations and assessment practices; 3) the reported challenges and corresponding solutions were synthesized into six main categories and 26 subcategories; and 4) future research directions were identified and organized into six main categories and 18 subcategories. The results of this MLR will assist researchers and practitioners in pursuing further studies and supporting the real-world adoption of multi-agent systems in industrial settings.

Z. Rasheed, Muhammad Waseem, Kai-Kristian Kemell et al. · 2 citations
#computer vision Apr 2026

Agentic Frameworks for Reasoning Tasks: An Empirical Study

Recent advances in agentic frameworks have enabled AI agents to perform complex reasoning and decision-making. However, evidence comparing their reasoning performance, efficiency, and practical suitability remains limited. To address this gap, we empirically evaluate 22 widely used agentic frameworks across three reasoning benchmarks: BBH, GSM8K, and ARC. The frameworks were selected from 1,200 GitHub repositories collected between January 2023 and July 2025 and organized into a taxonomy based on architectural design. We evaluated them under a unified setting, measuring reasoning accuracy, execution time, computational cost, and cross-benchmark consistency. Our results show that 19 of the 22 frameworks completed all three benchmarks. Among these, 12 showed stable performance, with mean accuracy of 74.6-75.9%, execution time of 4-6 seconds per task, and cost of 0.14-0.18 cents per task. Poorer results were mainly caused by orchestration problems rather than reasoning limits. For example, Camel failed to complete BBH after 11 days because of uncontrolled context growth, while Upsonic consumed USD 1,434 in one day because repeated extraction failures triggered costly retries. AutoGen and Mastra also exhausted API quotas through iterative interactions that increased prompt length without improving results. We also found a sharp drop in mathematical reasoning. Mean accuracy on GSM8K was 44.35%, compared with 89.80% on BBH and 89.56% on ARC. Overall, this study provides the first large-scale empirical comparison of agentic frameworks for reasoning-intensive software engineering tasks and shows that framework selection should prioritize orchestration quality, especially memory control, failure handling, and cost management.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 1 citation
#computer vision Open access Nov 2023

Autonomous Agents in Software Development: A Vision Paper

Large Language Models (LLM) and Generative Pre-trained Transformers (GPT), are reshaping the field of Software Engineering (SE). They enable innovative methods for executing many software engineering tasks, including automated code generation, debugging, maintenance, etc. However, only a limited number of existing works have thoroughly explored the potential of GPT agents in SE. This vision paper inquires about the role of GPT-based agents in SE. Our vision is to leverage the capabilities of multiple GPT agents to contribute to SE tasks and to propose an initial road map for future work. We argue that multiple GPT agents can perform creative and demanding tasks far beyond coding and debugging. GPT agents can also do project planning, requirements engineering, and software design. These can be done through high-level descriptions given by the human developer. We have shown in our initial experimental analysis for simple software (e.g., Snake Game, Tic-Tac-Toe, Notepad) that multiple GPT agents can produce high-quality code and document it carefully. We argue that it shows a promise of unforeseen efficiency and will dramatically reduce lead-times. To this end, we intend to expand our efforts to understand how we can scale these autonomous capabilities further.

Z. Rasheed, Muhammad Waseem, Kai-Kristian Kemell et al. · 35 citations · ⚡2
#large language models Open access Sep 2026

Engines of Escalation: How the Hybrid Public Sphere Amplifies Political Violence Targeting Women in Low-Resource Contexts

Political Violence Targeting Women (PVTW) has reached record highs in fragile, low-resource settings. This paper presents a structural analysis integrating both theory and empirical evidence to examine how a rapidly evolving hybrid public sphere—where the intertwined logics of traditional news and social media create a bidirectional feedback loop—amplifies hatred and accelerates offline physical harm against women in Nigeria. Adapting the framework of discursive opportunities, we unpack the mechanisms escalating misogyny into violence. We analyze a corpus of over 1.6 billion X (formerly Twitter) posts, traditional news articles, and conflict event data from ACLED. We utilize NaijaXLM-T, a custom Large Language Model tailored to Nigerian text, to accurately measure the visibility (volume) and resonance (intensity) of gender-specific hate, and to extract keywords and topics related to gender-targeting news. Furthermore, we map social interaction networks to isolate 'true contagion' from structural homophily. The hub-and-spoke topology observedamong hateful users serves as our proxy for legitimacy; we posit that hatred disseminated by influential network hubs acts as a digital form of social authorization, actively lowering the friction for offline mobilization. Empirically, we first establish a robust linkage where digital hostility directly catalyzes offline PVTW. Moving beyond this baseline effect, we unpack the structural pathways driving this acceleration by demonstrating how visibility, resonance, and network legitimacy fuel this spillover. We then answer critical questions regarding the differing velocities within this system, detailing the distinct temporal scales at which social media and traditional news amplify offline harm. Acknowledging observational and algorithmic limitations, we interpret these findings as a robust temporal linkage rather than strict causality. Ultimately, this research provides the foundational framework for monitoring and early-warning infrastructures, equipping policymakers and NGOs to prevent enabling stakeholders to pre-position protective resources to prevent digital hate from escalating into lethal PVTW in fragile settings.

Shaocheng Huang, Daniel Barkoczi · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.