Skip to content

Category

large language models

553 papers

#artificial intelligence Review Open access Nov 2026

A comparative review of modern large language model paradigms: GPT-4, BERT, Gemini, and DeepSeek

Comparison of GPT-4, BERT (bidirectional encoder representations from transformers), Gemini, and DeepSeek large language models (LLM), focusing on architectures, training methodologies, and real-world applications reveals GPT-4 excels in natural language generation and complex reasoning, supporting up to 128K tokens with moderate latency and higher costs making it effective for conversational artificial intelligence (AI).

Kavish Sanghvi, Aparna S. Sharma, Surbhi Hooda · 0 citations

An Explicit Interaction-Prompted Diffusion Framework for High-Fidelity 3D Molecular Generation.

Current structure-based drug design generative models often struggle to faithfully recapitulate genuine ligand-protein binding interactions. Instead, under the coupling of implicit learning architectures and biased training data, they tend to learn spurious statistical correlations. To address this, we propose EIP-Diff (Explicit Interaction-Prompted Diffusion), an architecture featuring a novel explicit interaction-prompt embedding mechanism that is better suited for real-world target-specific drug design. This architecture replaces biased implicit learning with explicit, residue-level biological guidance, thereby promoting more fine-grained geometric fidelity and more precise interaction-aware conditioning. To fully realize the capabilities of EIP-Diff and provide a reliable basis for performance evaluation, we further constructed CrystalData set, which provides higher-fidelity and less-biased structural supervision than existing data sets. This explicit architecture markedly improves distribution consistency: even when trained on the crossdocked data set, EIP-Diff achieves the highest alignment with authentic pharmacological distributions among evaluated models. Training on CrystalData set further enhances this alignment and improves 3D geometric accuracy, while retaining strong controllability, high chemical space coverage, and near-perfect uniqueness. In addition, target-based validation on KAT6A and YTHDC1 confirmed that EIP-Diff accurately recapitulates native-like binding modes. Furthermore, in a real-world drug design task against IDO1, we successfully designed a novel lead compound with nanomolar potency (IC50 = 0.31 nM). These results demonstrate that the EIP-Diff architecture can explicitly leverage experimentally derived structural data and biologically meaningful interaction information for target-specific molecular generation, thereby enabling its effective application to real-world structure-based drug design.

Huabin Du, Mingyang Wang, M. Luo et al. · 0 citations
#natural language process... Book Open access Aug 2026

Caduceus: MoE Foundation Models for Unifying Biological and Natural Language

This paper introduces Caduceus, a family of MoE-enhanced foundation models built with a hierarchical pre-training paradigm to jointly integrate biological and natural language, and incorporates a multi-task instruction tuning phase, enabling robust protein parsing and natural language question answering.

Mingze Yin, Yiheng Zhu, Jialu Wu et al. · 0 citations
#artificial intelligence Preprint Jul 2026

Constitutional Midtraining: Content Presence Drives Alignment Gains

Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce durable alignment when cleanly isolated from post-training. We build a 394M-token constitutional corpus from Anthropic's Constitution and apply constitutional midtraining at 120B scale, where principled, values-based content is inserted into midtraining. A 2x2 design (curriculum ordering x deliberative reasoning) was used to produce four constitutionally midtrained conditions, plus a control, which were evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment. All models were evaluated across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperformed the control on alignment generalization and durability, notably on blackmail: SFT instilled a blackmail propensity in all models, but constitutional midtraining blunted it, with the advantage surviving benign fine-tuning (-17.5pp). This durability did not extend to settings that required active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also mattered more than its structure, and constitutional midtraining incurred no capability cost, on average, at any stage (MMLU, ARC-Easy, piqa, GSM8K). A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.

Desiree Cho, Cameron Tice, Bernie Hogan et al. · 0 citations

E<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="feng-ieq1-3676689.gif"/></alternatives></inline-formula>LLM: Structure-Guided Efficient Inference for LLMs in Distributed Edge

Large language models (LLMs) are increasingly deployed in edge computing environments to reduce latency and preserve privacy. However, their inference process presents fundamental challenges for resource-constrained IoT devices. LLM inference involves computationally asymmetric stages: parallelizable prompt processing and sequential token decoding. This asymmetry creates deployment bottlenecks where IoT devices lack capacity for prompt processing while edge nodes suffer from inefficient sequential decoding. This paper presents <italic>E<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="feng-ieq3-3676689.gif"/></alternatives></inline-formula>LLM</italic>, an efficient distributed inference framework for large language models in heterogeneous edge-IoT environments. <italic>E<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="feng-ieq4-3676689.gif"/></alternatives></inline-formula>LLM</italic> leverages high-capacity edge devices for structural planning and introduces auxiliary lightweight models to generate segment-specific key-value (KV) caches. These minimal inference artifacts enable collaborative parallel decoding across IoT devices without requiring full model instantiation. The framework employs static-dynamic KV cache separation to minimize communication overhead while maintaining semantic coherence through structure-guided coordination. Extensive evaluation on realistic edge testbeds demonstrates significant performance improvements. Under diverse deployment settings, <italic>E<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="feng-ieq5-3676689.gif"/></alternatives></inline-formula>LLM</italic> achieves 74% –87.7% end-to-end latency reduction compared with several state-of-the-art baselines, while maintaining comparable generation quality; meanwhile, it also delivers a 34.6% –72.2% reduction in communication overhead, improves 9-12 × in energy efficiency. The framework exhibits strong scalability under bandwidth-limited conditions, enabling efficient LLM deployment across heterogeneous edge-IoT environments.

Xingyu Feng, Huanqi Yang, Zhuangzhuang Chen et al. · 0 citations
#large language models Open access Sep 2026

EXPLAINING TOURISM AND HOSPITALITY STUDENTS' ADOPTION OF LARGE LANGUAGE MODELS IN HIGHER EDUCATION: AN INTEGRATED TAM–UTAUT FRAMEWORK USING PLS-SEM AND NECESSARY CONDITION ANALYSIS

The rapid integration of large language models (LLMs) in higher education has transformed students' learning practices, particularly in applied disciplines such as tourism and hospitality education. Yet, limited empirical research explains the factors driving their adoption. Drawing on the Technology Acceptance Model (TAM) and the Unified Theory of Acceptance and Use of Technology (UTAUT), this study examines tourism and hospitality students' behavioural intention to use LLM-based learning tools by incorporating content reliability, learner motivation, and social influence as extended antecedents. Data were collected from 365 university students enrolled in tourism, hospitality and management courses in addition to the students enrolled in other allied programs having tourism as an elective course in India and analysed using partial least squares structural equation modelling (PLS-SEM) and Necessary Condition Analysis (NCA). The findings indicate that perceived usefulness remains central to adoption, while learner motivation and social influence play critical enabling roles. NCA further reveals that perceived usefulness, learner motivation, and social influence constitute necessary conditions for achieving high adoption intention. By integrating net-effect and necessity-based approaches, the study advances technology acceptance theory in AI-enabled education in tourism and hospitality. It offers practical insights for the responsible integration of LLMs in professional learning contexts.

Priya Singh, Mandeep Bharti, S. Chugh et al. · 0 citations
#large language models Open access Sep 2026

DECODING SOMATIC COMMUNICATION: A TRIADIC HUMAN-AI CO-CREATION FRAMEWORK FOR POST-CANCER PATIENT AUTONOMY

This paper outlines a systematic framework designed to integrate generative AI modalities into the field of restorative tattoo art, specifically targeting psychological and somatic rehabilitation post-oncological disease.

I. Ihnatenko, Liudmyla Vyshkvarok, Diana Raschupkina · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.