This release presents Contract-Grounded Cognitive Composition (CGCC), an engineering framework for improving the reliability of LLM multi-agent systems through deterministic workflow topology, shared interface contracts, and local validation. The study reports three exploratory experiments conducted with a local qwen3.5:9b model. The experiments examine sequential reasoning with deterministic validation, control-flow ablations across evidence retrieval, planning, and code generation, and frontend-backend integration under free communication, natural-language API documentation, and shared JSON Schema conditions. The results suggest that LLM cognition can be composed, but reliable composition depends on three distinct conditions: correct workflow progression, compatible interfaces, and validated local execution. Fixed topology reduces premature termination and routing errors, while shared API contracts reduce cases in which independently generated modules are locally plausible but fail during integration. This archive includes the English working paper, complete experimental code, raw model prompts and responses, routing and validation traces, aggregate results, supporting experiment notes, a data dictionary, and reproducibility materials. The experiments use one local 9B model and small synthetic task sets. The results should therefore be interpreted as an exploratory mechanism study rather than a general performance benchmark.
Zhongren Wang· Zenodo (CERN European Organi...· 0 citations
This study aimed to develop and evaluate an Android-based learning application that integrates Augmented Reality (AR) and artificial intelligence to support Indonesian Sign Language System (SIBI) learning for deaf students. The study employed the Research and Development (R&D) method using the Four-D (4D) model consisting of Define, Design, Develop, and Disseminate stages. The gesture recognition module combined MediaPipe for hand landmark detection with MobileNetV2 for gesture classification, achieving an approximate recognition accuracy of 89 %. The model was trained using 80% of the training data and 20% of the testing data. Expert validation involved two content experts and two media experts, while practicality testing involved 12 respondents (nine deaf students and three teachers). Effectiveness was evaluated using a one-group pretest-posttest design involving nine deaf students across four learning sessions. Data were collected through observations, interviews, documentation, questionnaires, and learning achievement assessments. The developed application obtained an overall expert validation score of 4.67 (Very Valid). The software quality evaluation based on the selected ISO/IEC 25010 characteristics showed 100% functional suitability, 80% portability, and 100% compatibility. The practicality testing produced an average score of 4.63 (Very Practical). The effectiveness evaluation indicated a significant improvement in learning outcomes (p = 0.001), with a moderate N-gain (0.56) and 77.78% classical learning mastery. These findings indicate that the developed application is valid, practical, and capable of supporting SIBI learning in preliminary classroom implementation. Nevertheless, the findings should be interpreted within the scope of a limited trial involving a small sample size and the absence of a control group.
Taufik Bahtiar, Haripuddin Haripuddin, Hasrul Bakri et al.· Jurnal Media Elektrik· 0 citations
The spread of Generative Artificial Intelligence and large language models (LLMs) has opened new possibilities for non-specialists to conceptualize and build functioning software on their own. This study exploratorily analyzes how a single user without formal programming training employed generative LLMs to develop, revise, and refine EasyHR Office and a self-understanding web application titled "나는 어떤 일을 어떻게 할 때 가장 잘 작동하는가?". The two programs originated from a combination of the desire to create something directly through generative AI and a practical concern with HR work in small workplaces as well as questions about how people function. In addition, recurring patterns observed across the two cases are tentatively organized under the label of Deliberate Deficit-Consolidation (DDC). Keywords: Generative Artificial Intelligence, Large Language Model, LLM, Non-specialist Development, Self-directed Learning, HR Practice, Software Development, Web Application, Case Study, DDC
Myung‐Jun Lee· Zenodo (CERN European Organi...· 0 citations
Current mainstream artificial intelligence models including convolutional neural networks and large language models rely on statistical fitting over fragmented input symbols. Vision models fit pixel distributions, while language models predict next‑token probabilities. This paradigm is essentially pattern‑matching and probabilistic speculation rather than genuine semantic understanding, and it inherently produces hallucinations, semantic drift, long‑tail failures and uncontrolled emergence. Breaking away from statistical‑learning frameworks, this paper proposes an AI‑native cognitive dynamical architecture built upon read‑only fixed semantic anchors. Instead of adopting human surface‑level grammar as internal computation rules, we construct a stack of layered transformation functions together with a graded semantic‑matching validation mechanism. Without token‑level probability sampling or massive pre‑training, five‑step cognitive dynamics (anchor assembling, verb stacking, complement modification, word‑order rearrangement and equivalent word replacement) reproduce core human‑like language comprehension and generation. A small‑scale sandbox experiment with ten basic semantic anchors demonstrates that the architecture eliminates root‑cause hallucinations and semantic drift, featuring full traceability, low computational cost and strong generalization. It represents a new human‑like cognitive paradigm alternative to statistical AI. Note: The mathematical dynamical function for semantic coupling, which inherits the prior “glue‑temporal‑grid” hypothesis, is not developed in this paper and will be presented in a follow‑up independent publication. Keywords AI‑native cognition; semantic anchor; transformation‑function stack; AI‑native grammar; statistics‑free modelling; explainable AI; hallucination mitigation
You Zhang· Zenodo (CERN European Organi...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Initial Research Release This is the first public release of Security in LLM-Generated Code. This release contains the research materials, experimental code, datasets, analysis scripts, findings, and supporting documentation associated with the study. Contents Research paper and supporting documentation Experimental datasets LLM-generated code security analysis Analysis and auditing scripts Research findings Reproducibility materials Version v1.0.0 This release represents the initial public version of the project and is intended to provide a stable, citable snapshot of the research.
Talha Imran· Zenodo (CERN European Organi...· 0 citations
This bilingual Chinese–English research monograph develops a controlled structural correspondence between a psychological–phenomenological account of third-order consciousness and the mathematics of AF-by-discrete groupoids. On the psychological side, the work asks how experience becomes organized as “my experience,” how separated episodes become recognized as “this is happening again,” and how the observer that names, compares, evaluates, and narrates experience can itself become part of what is observed. On the mathematical side, it independently develops the required language from finite graphs, Cantor path spaces, groupoids and germs, AF groupoids, graphs with boundary, tile inflation, traverses, returns, and incompressibility. The two domains are connected only through explicitly delimited structural correspondences and a stated translation discipline. A central working hypothesis is that subject-organizing complexity need not be uniformly distributed across experience: a large background may be organized through a small number of repeatedly re-identifiable, future-relevant interfaces. These interfaces are not stipulated in advance; they must be forced into view by repeated observation, return structure, failures of local continuation, and stable changes in what becomes possible afterward. The manuscript develops finite-scale formal tools, counterexamples, failure conditions, and candidate routes back to psychological observation and experimental design. The work does not identify consciousness with a groupoid and does not claim to provide a validated neural mechanism, clinical model, or ontological theory. Its aim is to make cross-disciplinary claims more discriminating, more explicit about their evidential burden, and easier to reject when the proposed correspondence fails. This manuscript has not undergone formal peer review. Zenodo is used here as a permanent archival and citation venue for the complete bilingual research text. Access to the deposited files is currently restricted; access may be granted by the author upon request.
Zheng Kuang· Zenodo (CERN European Organi...· 0 citations
A glass-box creativity engine, the Concept Collider, rests on a single primitive: a concept is not merely a point but a structure held under tension, and pressing it along one of seven fracture types breaks it in a characteristic way. That taxonomy is load-bearing—every profile, every homology score, every negative-space detector is computed from it—yet it had never been audited. We give it the first empirical audit and find it imperfect: redundant, non-orthogonal, incomplete, and in one case a consequence miscast as a mechanism. From the audit we reconstruct rather than discard: we demote that consequence, promote a recurring self-defeating tension, and propose an empirically grounded tree of fractures. But the deeper finding is that no taxonomy is canonical: standard decomposition criteria each return a different basis for the same data. Computational creativity is a relativistic discipline—novelty is relative to a base, value to an observer, and the tension taxonomy to a method—so the reportable object is not a basis but the invariants: a small core of fractures (circularity, contradiction, self-defeating means) that survives every change of decomposition method and, across five model families spanning both Western and Chinese training ecosystems (Anthropic, Meta, OpenAI, Alibaba, Moonshot), every change of vendor—though the agreement weakens once the tensions are judged in Chinese, so the invariance holds within a language more than across it. That partial cross-ecosystem agreement weakens—without eliminating—the worry that such consensus merely records shared training text; we therefore report the invariants as robust descriptors under our measurement, not as proof that concepts possess a mind-independent structure. The contribution is a transferable criterion—an invariant is what stays put when both method and model family vary. --- Changes in this version (v2): This second version incorporates the mid-2026 literature on inter-model agreement: Ding (2026) audits agreement as a confidence signal and finds it a positive but weak predictor of correctness, while Liu (2026) supplies the mechanism — error decorrelation across independently trained models — and names its ceiling as a shared-error floor. Both are used to state the size of the observer-independence problem rather than only its direction. The approach is also situated in the psychometric lineage of van der Wal et al. (JAIR 79, 2024), who bring construct validity and reliability to bear on bias measures in NLP. Reproducibility artefacts are deposited with this version: the anonymised 380×7 fracture-profile matrix with its provenance metadata, the per-judge classifications from five model families spanning Western and Chinese training ecosystems, and the analysis and figure scripts. Concept names, tension texts and prompts are withheld; the matrix carries opaque identifiers, which changes no published value — verified by recomputation — while keeping the knowledge base of the audited system out of the release. One correction: the adjusted Rand index of the emergence test has been recomputed from the source corpus and revised from 0.02 to 0.065, and the silhouette is reported as flat across every number of clusters rather than monotonically rising. The conclusion is unchanged — the taxonomy does not emerge from the tension descriptions — and a language control (ARI = −0.001) is now reported alongside it.
Sebastian Wahl· Zenodo (CERN European Organi...· 0 citations
According to the drawbacks of large language model, I design a new agent which can overcome the drawbacks. The agent contains four deep neural networks, which is abstract network, concrete network, decision network, execution network. Each of the four networks just do one thing, so it focus what it can do just like brain works. The input is perception information by sensor. The output1 are classes and attributes(common sense), the output2 is memory or consciousness, the output3 is logic and theory, the output4 is action and practice. The No.1 and No.2 constitute an auto-encoder. The dimensions of classes are very high if the grain size is very small when one-hot encoding used, so binary encoding can be used. With the help of muti-level classes and sentence structure, output1 is produced as one sentence and common sense. With the help of two order dimensions, the output3 is a sentence, so the reasoning speed is more higher than LLM. The probability of sentence is joint probability of words, so only high probability of words can produce, so the hallucination problem alleviates. With the help of select gate tanh, when n is small, the computation time is one half of self-attention. When n is large, the time decreases much more. With the help of multi-value functions, the independent consciousness is produced. It is initiative, not rely on prompt. Causal reasoning and Continuous reasoning are realized by concatenate output3 with the input of No.3 as the new input of No.3.
Jinxin Wei, Zhe Hou· Zenodo (CERN European Organi...· 0 citations
AbstractThis paper investigates EFL students’ perceptions of ethical acceptability judgments of using Large Language Models (LLMs) into academic writing. In contrast to the simplistic view of acceptable/unacceptable use of LLMs, the present study models how specific contextual justifications shape students’ moral evaluations of LLM-assisted writing. A cross-sectional within-subject design was used with 220 third-year EFL students at three public universities in Laghouat, Algeria, to rate the ethical acceptability of LLMs use to complete four writing assignments. Under a neutral baseline condition, participants assessed the use of LLM for completing the four assignments, as well as four single condition contexts (disclosure, accuracy verification, syllabus permission, and learning intent). Their ratings were analyzed using Δ-effect scores (conditional minus baseline) to quantify condition effects above baseline and to describe task-level differences in ethical acceptability. Contextual conditions increased ethical acceptability to different extents, with learning intent producing the largest positive shift (ΔM ≈ 1.6, d ≈ 0.7) and disclosure exerting only a small, non-robust effect (ΔM ≈ 0.2, d ≈ 0.1). The baseline ratings also followed a clear gradient, which saw AI-assisted proofreading as the most acceptable while paraphrasing was consistently least acceptable. Theoretically, the study adds value by supporting a conditional-ethics perspective because it demonstrates how students’ moral considerations of LLMs use are dependent on learning-oriented and epistemically responsible frameworks rather than procedural cues like disclosure or syllabus permissions with no behavioral consequences. Practically, the paper advices that AI policies and pedagogy should be designed to encourage learning intent and verification activities as opposed to disclosure as a standalone requirement.Keywords: Academic integrity; large language models (LLMs); EFL writing; conditional ethics
Mohamed SEDDIKI, Souhila Korichi· Zenodo (CERN European Organi...· 0 citations
AbstractThis paper investigates EFL students’ perceptions of ethical acceptability judgments of using Large Language Models (LLMs) into academic writing. In contrast to the simplistic view of acceptable/unacceptable use of LLMs, the present study models how specific contextual justifications shape students’ moral evaluations of LLM-assisted writing. A cross-sectional within-subject design was used with 220 third-year EFL students at three public universities in Laghouat, Algeria, to rate the ethical acceptability of LLMs use to complete four writing assignments. Under a neutral baseline condition, participants assessed the use of LLM for completing the four assignments, as well as four single condition contexts (disclosure, accuracy verification, syllabus permission, and learning intent). Their ratings were analyzed using Δ-effect scores (conditional minus baseline) to quantify condition effects above baseline and to describe task-level differences in ethical acceptability. Contextual conditions increased ethical acceptability to different extents, with learning intent producing the largest positive shift (ΔM ≈ 1.6, d ≈ 0.7) and disclosure exerting only a small, non-robust effect (ΔM ≈ 0.2, d ≈ 0.1). The baseline ratings also followed a clear gradient, which saw AI-assisted proofreading as the most acceptable while paraphrasing was consistently least acceptable. Theoretically, the study adds value by supporting a conditional-ethics perspective because it demonstrates how students’ moral considerations of LLMs use are dependent on learning-oriented and epistemically responsible frameworks rather than procedural cues like disclosure or syllabus permissions with no behavioral consequences. Practically, the paper advices that AI policies and pedagogy should be designed to encourage learning intent and verification activities as opposed to disclosure as a standalone requirement.Keywords: Academic integrity; large language models (LLMs); EFL writing; conditional ethics
Mohamed SEDDIKI, Souhila Korichi· Zenodo (CERN European Organi...· 0 citations
Context: Manual qualitative data analysis is time-intensive and can compromise validity and replicability, affecting analysis design, implementation, and reporting. Large Language Models (LLMs) enable human-bot collaboration in Software Engineering (SE), but their potential for qualitative data analysis in SE remains largely unexplored. Objective: The objective of this study is to design and develop an LLM-based multi-agent system that synergizes human decision support with AI to automate various qualitative data analysis approaches. Methods: We used LLM-based multi-agents systems to assist the qualitative data analysis process, deploying 27 agents, each responsible for a specific task, such as text summarization, initial code generation, and extracting themes and patterns. Results: The main findings are: (1) the LLM-based multi-agent system accelerates the qualitative data analysis process, (2) the system effectively automates tasks such as text summarization, initial code generation, and theme extraction, and (3) the publicly accessible code facilitates validation and further evaluation. Conclusion: The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners. Future improvements focus on enhancing multilingual performance and integrating continuous expert feedback. The source code of proposed system and system details can be found here: https://github.com/GPT-Laboratory/Qualitative-Analysis-with-an-LLM-Based-Agentts
Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al.· arXiv.org· 40 citations
Large Language Models (LLMs) have enabled multi-agent systems to perform autonomous code generation for complex tasks. Despite the recent growth in research and industrial applications in this area, there is little work on synthesizing evidence from both academic and industrial sources to capture the current state of research on LLM-based multi-agent systems for code generation. To this end, we conducted a Multi-Vocal Literature Review (MLR), combining insights from both academia and industry, including peer-reviewed studies and grey literature. The aim of this study is to systematically synthesize and analyze existing knowledge on LLM-based multi-agent systems for code generation. Specifically, the review examines the motivations for their use, employed benchmarks and models, key challenges, proposed solutions, and potential directions for future research. We selected and reviewed 114 studies, and the key findings are: 1) the identified reasons for adopting multi-agent systems for code generation were classified into nine categories; 2) the models and evaluation benchmarks utilized across the studies were systematically analyzed to provide a structured overview of commonly adopted LLM configurations and assessment practices; 3) the reported challenges and corresponding solutions were synthesized into six main categories and 26 subcategories; and 4) future research directions were identified and organized into six main categories and 18 subcategories. The results of this MLR will assist researchers and practitioners in pursuing further studies and supporting the real-world adoption of multi-agent systems in industrial settings.
Z. Rasheed, Muhammad Waseem, Kai-Kristian Kemell et al.· arXiv.org· 2 citations
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.