Skip to content

Category

federated learning

398 papers

#machine learning Open access Aug 2026

AI-driven PROTAC design overcomes oncogenic resilience by eliminating the CLIP1-LTK fusion protein.

The discovery of CAP-Gly domain-containing linker protein 1(CLIP1)-Leukocyte tyrosine kinase (LTK) as an oncogenic fusion reveals a unique dependency not only on LTK kinase activity but also on CLIP1-mediated multimerization, a noncatalytic function that drives oncogenic signaling. While this fusion is currently targeted with anaplastic lymphoma kinase inhibitors, their exclusive focus on kinase inhibition leaves the scaffolding function intact, necessitating a complete protein clearance strategy. Here, we report the AI-guided development of a first-in-class proteolysis-targeting chimera (PROTAC) designed to selectively degrade the CLIP1-LTK fusion protein. By integrating deep learning models for ternary complex prediction with structure-based molecular optimization, we designed DCL05, an orally bioavailable degrader of CLIP1-LTK fusion protein, achieving picomolar degradation potency (DC50 = 40 pM) and robust antitumor activity. DCL05 consistently outperformed existing kinase inhibitors across a broad spectrum of LTK resistance-associated mutations, both in vitro and in vivo. Collectively, our study explores resistance-associated contexts of LTK and establishes a structure-guided PROTAC development pipeline, providing a promising therapeutic strategy for overcoming acquired resistance in kinase-driven cancers.

Shicheng Chen, Haiting Duan, S. Zhong et al. · 0 citations
#machine learning Open access Apr 2026

Accurate and task-agnostic modeling of enzymatic reactions through multimodal relational learning

Enzymatic reactions play an emerging role in a broad spectrum of scientific and industrial applications. The inherent complexity of enzymes, such as their substrate specificity, conformational flexibility, and the vast diversity of reactions involved, poses substantial challenges for the advanced computational prediction of enzymatic reactions with desirable accuracy. Moreover, existing approaches are mostly tailored for a specific sub-task, such as substrate prediction or binding site annotation, which limits their applicability. In this study, we introduce ERAM, a task-agnostic multimodal learning framework capable of addressing a broad range of downstream applications with both accuracy and efficiency. ERAM aligns pre-trained molecular representations from Protein Language Model with the knowledge of enzyme catalysis by modeling enzymatic reactions as multi-relational data. In enzyme retrieval tasks, ERAM achieves an improvement of 28.31% in mean average precision compared with the state-of-the-art (SOTA) method, CREEP. In substrate prediction tasks, ERAM outperforms the SOTA method ESP, achieving average improvements of 35.53% and 22.97% in Matthews correlation coefficient across two datasets. Additionally, ERAM exhibits commendable interpretability by assigning higher attention weights to binding sites, resulting in lower false-positive rates (42.36%) and higher overlap scores (70.59%) in the unsupervised binding site prediction task compared to RXNAA Mapper. By learning embeddings of substrates, enzymes, and products within a unified knowledge graph latent space, ERAM demonstrates its potential as a versatile and effective tool for enzyme catalysis research.

Yuansheng Huang, Lanqing Li, Wenjia Qian et al. · 2 citations
#machine learning Open access Jul 2026

BBBP-Atlas: Unified Interpretable Modeling of Blood–Brain Barrier Permeability across Small Molecules and Peptides

Accurate prediction of blood-brain barrier permeability (BBBP) is essential for central nervous system drug discovery, yet existing models are often limited by their reliance on predefined physicochemical descriptors, small-molecule-centered training sets, or conformation-dependent representations, which restricts their transferability across chemically diverse modalities especially peptides. In addition, publicly available BBBP datasets remain fragmented, inconsistently standardized, and weakly controlled for molecular redundancy, increasing the risk of data leakage and overestimated model performance. In this study, we propose BBBP-Atlas, a structure-aware BBB permeability prediction model designed for unified modeling of small molecules and peptides with the first cross-modal dataset OmniBBBP. Designed to bypass descriptor and conformation dependencies, our model represents standardized molecular structures as atom-level graphs to capture local atom-bond environments and long-range topological dependencies associated with BBB transport. This design enables direct learning of structure-permeability relationships from molecular topology. For model training and evaluation, we curated a cross-modal, redundancy-filtered database OmniBBBP that seamlessly unifies small molecules and complex peptides, containing 10,218 unique compounds with 9,316 small molecules and 902 peptides. BBBP-Atlas achieved an accuracy of 0.8914 and an MCC of 0.7678 on the independent test set. On a balanced external benchmark of 200 compounds, our model reached an AUC of 0.9108, an accuracy of 0.8500, and an MCC of 0.7000, outperforming LightBBB by an absolute MCC gain of 6%. Case studies further showed that BBBP-Atlas captured clinically meaningful BBB permeability patterns, correctly identifying lorlatinib as BBB-permeable and vancomycin as BBB-impermeable with high confidence. The OmniBBBP-backed BBBP-Atlas offers a versatile and cross-modal approach for single-compound prediction, batch screening, and dataset exploration for CNS drug discovery. BBBP-Atlas is available at https://cadd.drugflow.com/bbbp/.

Xin Shen, Qun Su, Hao Luo et al. · 0 citations
#machine learning Open access Jun 2026

Targeting the intrinsically disordered AR-NTD through a machine learning-based enhanced sampling workflow

Targeting the intrinsically disordered N-terminal domain of the androgen receptor (AR-NTD) represents a promising strategy to overcome resistance in prostate cancer. However, its inherent lack of a stable tertiary structure and highly dynamic conformational ensemble pose formidable challenges for rational drug design. This study introduces an integrated computational workflow that combines enhanced sampling techniques and machine learning collective variables to identify druggable conformations of the AR-NTD and elucidate the binding mechanism of its modulator, EPI-002. We characterize nine metastable states of the Tau-5 region and reveal that ligand recognition is driven by π–π stacking and structured water-mediated hydrogen bonds. Leveraging these insights, we perform structure-based virtual screening based on the identified druggable conformations and identify K53, a rationally designed AR-NTD antagonist, which exhibits potent anti-proliferative activity in enzalutamide-resistant prostate cancer cells. K53 directly binds the AR-NTD, suppresses AR transcriptional activity, and demonstrates high selectivity for cancer cells. This work provides a rational design paradigm for targeting intrinsically disordered proteins and offers a therapeutic candidate for resistant prostate cancer. In this work, the authors develop a machine learning–based enhanced sampling workflow to target the intrinsically disordered AR-NTD, identifying druggable conformations and enabling transferable modeling of ligand binding for rational drug discovery.

Kai Zhu, Huating Wang, Jintu Zhang et al. · 0 citations
#federated learning Open access Aug 2026

Information-Theoretic Framework for Trustworthy Federated Learning

Federated Learning (FL) offers a promising paradigm for decentralized machine learning, enabling collaborative model training without direct data sharing. However, traditional privacy-preserving techniques within FL often rely on ad-hoc assumptions and lack a rigorous theoretical basis. This work introduces an information-theoretic framework to address this limitation. We define a "Privacy Loss Function" predicated on mutual information between local models and global updates, providing a quantifiable measure of information leakage. The framework leverages established techniques such as differential privacy and homomorphic encryption to minimize this loss, ultimately leading to more robust and trustworthy FL systems. Our approach moves beyond intuitive notions of privacy, offering a mathematically sound foundation for designing and analyzing FL protocols, facilitating the development of truly secure and efficient distributed learning solutions. The core contribution is the formalization of privacy risk in FL using information-theoretic principles, enabling a more precise understanding and control over data leakage.

Jincheng Zhang · 0 citations
#federated learning Open access Aug 2026

Privacy-Enhancing Federated Learning Models for Cybersecurity in IoT Networks

The rapid expansion of the Internet of Things (IoT) has intensified cybersecurity risks by exposing distributed connected devices to increasingly complex and pervasive threats. Conventional centralized security mechanisms often struggle to accommodate the heterogeneous and decentralized structure of IoT networks. This study investigates Federated Learning (FL) as a decentralized approach to intrusion detection that enables local model training on IoT edge devices while transmitting only encrypted model updates to a central server, thereby preserving data privacy and reducing communication overhead. A novel FL-based Intrusion Detection System (IDS) architecture was developed using Convolutional Neural Networks (CNNs) for anomaly detection and the Federated Averaging (FedAvg) algorithm for aggregating local model updates. The framework was evaluated on standard IoT datasets under non-independent and identically distributed (non-IID) data conditions to simulate heterogeneous real-world environments. Experimental results demonstrate that the proposed system achieved a detection accuracy of 94.6%, an F1-score of 93.8%, and a recall of 92.7%, outperforming centralized and standalone local learning methods. The framework also reduced communication overhead by 35% and achieved convergence 28% faster than conventional approaches. These findings demonstrate that FL can provide a scalable, privacy-preserving, and computationally efficient foundation for strengthening IoT cybersecurity. This study contributes a decentralized machine-learning architecture for real-time, adaptive, and privacy-conscious intrusion detection in large-scale IoT environments.

Mohammed Ajuji, Yusuf Musa Malgwi, Asabe Sandra Ahmadu et al. · 0 citations
#federated learning Open access Aug 2026

Decentralized Federated Learning with Byzantine Fault Tolerance using Blockchain-Based Verification

Federated learning (FL) offers a promising approach to training machine learning models on decentralized datasets without directly sharing the data itself. However, existing FL systems are susceptible to Byzantine attacks, where malicious participants can inject poisoned data or manipulate model updates, ultimately compromising the global model's integrity. This paper proposes a novel decentralized federated learning system incorporating Byzantine fault tolerance (BFT) achieved through a blockchain-based verification layer. The system leverages blockchain technology to cryptographically verify model updates from each participant before they are aggregated into the global model. This ensures data integrity and enables the detection and mitigation of Byzantine attacks. The core claim is that existing FL systems are vulnerable to these attacks. The proposed mechanism implements a decentralized FL system utilizing a blockchain-based verification layer. This represents a new approach to robust FL training. The system utilizes the following key components: participant nodes, a blockchain network, and a global model aggregator. The blockchain network is composed of multiple nodes that maintain a ledger of all model updates. The global model aggregator utilizes the blockchain network to verify model updates before aggregating them into the global model. The proposed system enhances the security and robustness of FL training by providing a tamper-proof audit trail and enabling the detection of malicious participants.

Jincheng Zhang · 0 citations
#federated learning Open access Aug 2026

A machine learning framework for privacy preserving personalized multimodal emotion recognition

This paper presents a novel machine learning framework, Privacy-Preserving Personalized Multimodal Emotion Recognition (P3MER), that simultaneously addresses three fundamental challenges in affective computing: achieving state-of-the-art recognition accuracy, ensuring robust privacy protection, and enabling effective personalization to individual users. The framework integrates hierarchical multimodal fusion with federated learning and differential privacy to enable collaborative model training without centralized data collection, thereby preserving the confidentiality of sensitive biometric data such as facial expressions, speech recordings, and physiological signals. A key innovation is the incorporation of federated meta-learning that allows rapid personalization of global models to individual expression patterns with minimal local data, while maintaining formal privacy guarantees. Extensive experimental evaluation across three benchmark datasets (CMU-MOSEI, DEAP, and MAHNOB-HCI) demonstrates that P3MER achieves an average improvement of 4.1% in recognition accuracy over state-of-the-art centralized models, while providing formal \((\epsilon , \delta )\) -differential privacy guarantees. At a privacy budget of \(\epsilon = 3.0\) , the framework maintains 95.8% of the non-private federated performance, significantly outperforming conventional differentially private federated learning approaches. The meta-learning personalization mechanism yields an average personalization gain of 14.7% with only five adaptation steps, effectively addressing the inherent heterogeneity in emotional expression across individuals. Furthermore, the framework demonstrates exceptional robustness to real-world challenges, including extreme data heterogeneity (39% reduction in performance variance compared to existing personalized federated approaches), modality incompleteness (maintaining 86.3% of full-modality performance when physiological signals are unavailable), and few-shot learning scenarios (achieving 78% of maximum personalization gain with only 20-50 local samples). These results collectively validate that P3MER successfully reconciles the competing objectives of accuracy, privacy, and personalization in multimodal emotion recognition, offering a practical pathway toward deployable, ethical affective computing systems that respect user privacy while maintaining adaptive intelligence. The proposed framework establishes new standards for privacy-preserving affective computing and provides both theoretical foundations and practical implementations for developing emotion-aware technologies that earn user trust through their technical capability and ethical design. By demonstrating that privacy protection and personalization need not come at the expense of recognition accuracy, this work advances the field toward human-centered AI systems that are simultaneously intelligent, adaptive, and respectful of fundamental privacy rights.

Sanjay Agal · 0 citations
#federated learning Review Open access Aug 2026

The Decentralized Classroom: A Narrative Review of Federated Learning from Google's Keyboards to the Privacy's Frontier

Federated learning---the artificial intelligence whose subject is the decentralized classroom and whose lesson is the model's travel---moved from Dwork's 2006 differential privacy and Shokri and Shmatikov's 2015 gradients through Konečný's 2016 compression, McMahan's 2017 FedAvg, and Bonawitz's 2017 aggregation to Zhao's 2018 non-IID, Kairouz's 2021 survey, and Zhu's 2019 leakage. This article presents a narrative review of that arc's canonical line: Dwork's 2006 ICALP, Shokri and Shmatikov's 2015 CCS, Konečný and colleagues's 2016 strategies, McMahan, Moore, Ramage, Hampson, and Arcas's 2017 FedAvg, Bonawitz and colleagues's 2017 secure aggregation, Zhao and colleagues's 2018 non-IID, Hard and colleagues's 2018 keyboard, Zhu, Liu, and Han's 2019 gradients, Yang and colleagues's 2019 concept, Li and colleagues's 2020 convergence, Li and colleagues's 2020 challenges, and Kairouz and colleagues's 2021 advances. The review is organized around three themes: the privacy's premise and the communication's bottleneck, in which the Dwork's noise and the Shokri-Shmatikov's gradients founded the distributed's training; the algorithm's and the deployment's era, in which the FedAvg's averaging, the secure's aggregation, and the keyboard's deployment gave the federation its engine; and the heterogeneity's and the frontier's era, in which the non-IID's data, the gradient's leakage, the convergence's proofs, and the open's problems carried the field into the privacy's science. It is concluded that federated learning is the machine learning's decentralization---and that its arc is the classroom's reading from the centralized's server to the privacy's frontier.

Zen Revista, 10 IA · 0 citations
#federated learning Open access Aug 2026

Information-Theoretic Framework for Trustworthy Federated Learning

Federated Learning (FL) offers a promising paradigm for decentralized machine learning, enabling collaborative model training without direct data sharing. However, traditional privacy-preserving techniques within FL often rely on ad-hoc assumptions and lack a rigorous theoretical basis. This work introduces an information-theoretic framework to address this limitation. We define a "Privacy Loss Function" predicated on mutual information between local models and global updates, providing a quantifiable measure of information leakage. The framework leverages established techniques such as differential privacy and homomorphic encryption to minimize this loss, ultimately leading to more robust and trustworthy FL systems. Our approach moves beyond intuitive notions of privacy, offering a mathematically sound foundation for designing and analyzing FL protocols, facilitating the development of truly secure and efficient distributed learning solutions. The core contribution is the formalization of privacy risk in FL using information-theoretic principles, enabling a more precise understanding and control over data leakage.

Jincheng Zhang · 0 citations
#federated learning Open access Aug 2026

Decentralized Federated Learning with Differential Privacy for Scientific Data

This paper presents a novel approach to collaborative scientific data analysis leveraging Decentralized Federated Learning with Differential Privacy (DFLDP). The core challenge in many scientific domains is the reluctance to share raw data due to stringent privacy regulations and intellectual property protections. Traditional Federated Learning (FL) solutions, while offering a degree of data privacy, still rely on centralized aggregation, a point of vulnerability. Our proposed DFLDP framework addresses this limitation by adopting a decentralized architecture where individual researchers maintain complete control over their datasets. Crucially, we integrate differential privacy mechanisms directly into the aggregation process, adding a quantifiable layer of protection against data leakage. This ensures that the learned model benefits from the collective knowledge of multiple researchers without revealing individual data contributions. The system utilizes a gossip-based communication protocol for model updates, minimizing communication overhead. We formally define the mathematical framework, outlining the key components and their interactions. The system's performance is evaluated in a simulated environment, demonstrating the effectiveness of the DFLDP approach in achieving accurate models while upholding stringent privacy guarantees. The core claim of this work is that sharing raw scientific data for federated learning is often prohibited due to privacy concerns and intellectual property restrictions. The core mechanism implemented is the realization of a federated learning system that utilizes differential privacy to protect data during aggregation, while also employing a decentralized architecture where individual researchers retain control over their data. This new approach combines federated learning with differential privacy and decentralization, enabling collaborative scientific discovery without compromising data privacy or intellectual property rights.

Jincheng Zhang · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.