Category
large language models
553 papers
Do people apply different norms to humans and large language models acting on their behalf? Evidence from norm elicitations in two canonical economic games
Decoding hate: The rise of transformer and large language model architectures in automated hate speech detection
A comparative review of modern large language model paradigms: GPT-4, BERT, Gemini, and DeepSeek
Comparison of GPT-4, BERT (bidirectional encoder representations from transformers), Gemini, and DeepSeek large language models (LLM), focusing on architectures, training methodologies, and real-world applications reveals GPT-4 excels in natural language generation and complex reasoning, supporting up to 128K tokens with moderate latency and higher costs making it effective for conversational artificial intelligence (AI).
Reach audiences
Advertise in front of researchers, engineers, and readers.
An Explicit Interaction-Prompted Diffusion Framework for High-Fidelity 3D Molecular Generation.
Current structure-based drug design generative models often struggle to faithfully recapitulate genuine ligand-protein binding interactions. Instead, under the coupling of implicit learning architectures and biased training data, they tend to learn spurious statistical correlations. To address this, we propose EIP-Diff (Explicit Interaction-Prompted Diffusion), an architecture featuring a novel explicit interaction-prompt embedding mechanism that is better suited for real-world target-specific drug design. This architecture replaces biased implicit learning with explicit, residue-level biological guidance, thereby promoting more fine-grained geometric fidelity and more precise interaction-aware conditioning. To fully realize the capabilities of EIP-Diff and provide a reliable basis for performance evaluation, we further constructed CrystalData set, which provides higher-fidelity and less-biased structural supervision than existing data sets. This explicit architecture markedly improves distribution consistency: even when trained on the crossdocked data set, EIP-Diff achieves the highest alignment with authentic pharmacological distributions among evaluated models. Training on CrystalData set further enhances this alignment and improves 3D geometric accuracy, while retaining strong controllability, high chemical space coverage, and near-perfect uniqueness. In addition, target-based validation on KAT6A and YTHDC1 confirmed that EIP-Diff accurately recapitulates native-like binding modes. Furthermore, in a real-world drug design task against IDO1, we successfully designed a novel lead compound with nanomolar potency (IC50 = 0.31 nM). These results demonstrate that the EIP-Diff architecture can explicitly leverage experimentally derived structural data and biologically meaningful interaction information for target-specific molecular generation, thereby enabling its effective application to real-world structure-based drug design.
ClauseMiner: Prompt-engineered large language model for accurate, scalable legal clause extraction
Caduceus: MoE Foundation Models for Unifying Biological and Natural Language
This paper introduces Caduceus, a family of MoE-enhanced foundation models built with a hierarchical pre-training paradigm to jointly integrate biological and natural language, and incorporates a multi-task instruction tuning phase, enabling robust protein parsing and natural language question answering.
Constitutional Midtraining: Content Presence Drives Alignment Gains
Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce durable alignment when cleanly isolated from post-training. We build a 394M-token constitutional corpus from Anthropic's Constitution and apply constitutional midtraining at 120B scale, where principled, values-based content is inserted into midtraining. A 2x2 design (curriculum ordering x deliberative reasoning) was used to produce four constitutionally midtrained conditions, plus a control, which were evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment. All models were evaluated across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperformed the control on alignment generalization and durability, notably on blackmail: SFT instilled a blackmail propensity in all models, but constitutional midtraining blunted it, with the advantage surviving benign fine-tuning (-17.5pp). This durability did not extend to settings that required active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also mattered more than its structure, and constitutional midtraining incurred no capability cost, on average, at any stage (MMLU, ARC-Easy, piqa, GSM8K). A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.
E<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="feng-ieq1-3676689.gif"/></alternatives></inline-formula>LLM: Structure-Guided Efficient Inference for LLMs in Distributed Edge
Large language models (LLMs) are increasingly deployed in edge computing environments to reduce latency and preserve privacy. However, their inference process presents fundamental challenges for resource-constrained IoT devices. LLM inference involves computationally asymmetric stages: parallelizable prompt processing and sequential token decoding. This asymmetry creates deployment bottlenecks where IoT devices lack capacity for prompt processing while edge nodes suffer from inefficient sequential decoding. This paper presents <italic>E<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="feng-ieq3-3676689.gif"/></alternatives></inline-formula>LLM</italic>, an efficient distributed inference framework for large language models in heterogeneous edge-IoT environments. <italic>E<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="feng-ieq4-3676689.gif"/></alternatives></inline-formula>LLM</italic> leverages high-capacity edge devices for structural planning and introduces auxiliary lightweight models to generate segment-specific key-value (KV) caches. These minimal inference artifacts enable collaborative parallel decoding across IoT devices without requiring full model instantiation. The framework employs static-dynamic KV cache separation to minimize communication overhead while maintaining semantic coherence through structure-guided coordination. Extensive evaluation on realistic edge testbeds demonstrates significant performance improvements. Under diverse deployment settings, <italic>E<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="feng-ieq5-3676689.gif"/></alternatives></inline-formula>LLM</italic> achieves 74% –87.7% end-to-end latency reduction compared with several state-of-the-art baselines, while maintaining comparable generation quality; meanwhile, it also delivers a 34.6% –72.2% reduction in communication overhead, improves 9-12 × in energy efficiency. The framework exhibits strong scalability under bandwidth-limited conditions, enabling efficient LLM deployment across heterogeneous edge-IoT environments.
EXPLAINING TOURISM AND HOSPITALITY STUDENTS' ADOPTION OF LARGE LANGUAGE MODELS IN HIGHER EDUCATION: AN INTEGRATED TAM–UTAUT FRAMEWORK USING PLS-SEM AND NECESSARY CONDITION ANALYSIS
The rapid integration of large language models (LLMs) in higher education has transformed students' learning practices, particularly in applied disciplines such as tourism and hospitality education. Yet, limited empirical research explains the factors driving their adoption. Drawing on the Technology Acceptance Model (TAM) and the Unified Theory of Acceptance and Use of Technology (UTAUT), this study examines tourism and hospitality students' behavioural intention to use LLM-based learning tools by incorporating content reliability, learner motivation, and social influence as extended antecedents. Data were collected from 365 university students enrolled in tourism, hospitality and management courses in addition to the students enrolled in other allied programs having tourism as an elective course in India and analysed using partial least squares structural equation modelling (PLS-SEM) and Necessary Condition Analysis (NCA). The findings indicate that perceived usefulness remains central to adoption, while learner motivation and social influence play critical enabling roles. NCA further reveals that perceived usefulness, learner motivation, and social influence constitute necessary conditions for achieving high adoption intention. By integrating net-effect and necessity-based approaches, the study advances technology acceptance theory in AI-enabled education in tourism and hospitality. It offers practical insights for the responsible integration of LLMs in professional learning contexts.
DECODING SOMATIC COMMUNICATION: A TRIADIC HUMAN-AI CO-CREATION FRAMEWORK FOR POST-CANCER PATIENT AUTONOMY
This paper outlines a systematic framework designed to integrate generative AI modalities into the field of restorative tattoo art, specifically targeting psychological and somatic rehabilitation post-oncological disease.
Physics-informed large language model with contrastive temporal embedding for transient thermal-hydraulic prediction in loss-of-coolant accidents
From tech blogs
See all →Introducing WeatherNext 3, our most advanced and accurate global weather AI model
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
Training a coding model to paint watercolours with TRL and OpenEnv
Real-Time Intelligence with IBM Time Series Models on Confluent
TimesFM-3: A zero-shot foundation model for multivariate forecasting
Data Management
GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.