Skip to content
Open access

HELM-BERT: Topology-Aware Representations for Chemically Modified Peptides

Jul 2026 · Journal of Chemical Information and Modeling · Vol 66, pp. 7900 - 7916 · 0 citations · 31 references
Medicine

TL;DR

These findings suggest that pretraining directly on notation that makes structural constraints explicit offers a transferable strategy for biomolecular modalities that fall between small-molecule chemistry and protein sequence.

Abstract

Chemically modified and macrocyclic peptides are increasingly important therapeutics, yet current molecular representation models do not natively represent chemical modification and covalent topology in a unified way. Atom-level strings obscure macrocyclic connectivity, whereas protein sequence models cannot encode noncanonical residues and explicit cross-links. Here we pretrain an encoder-only transformer directly on Hierarchical Editing Language for Macromolecules (HELM) notation, which specifies monomer identity and connectivity. In this work, we show that the resulting representations achieve best mean performance in cyclic peptide membrane permeability prediction (random split R 2 = 0.668; retaining best mean performance under a Murcko scaffold split), exceeding external pretrained SMILES-based encoders. An architecture-matched SMILES control narrowed the HELM–SMILES gap under full finetuning, whereas HELM-BERT retained clearer advantages in frozen-representation settings. HELM-BERT also preserves HELM-specified macrocyclic topology in a linearly accessible form and supports competitive peptide–protein interaction prediction across complementary Propedia and ChEMBL benchmarks. More broadly, these findings suggest that pretraining directly on notation that makes structural constraints explicit offers a transferable strategy for biomolecular modalities that fall between small-molecule chemistry and protein sequence.

Read PDF

Similar papers

Preprint Jul 2026

Vilya-1: An all-atom foundation model for macrocycle structure prediction and design

Vilya-1 is introduced, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability.

Vilya Research Pascal Sturmfels, M. Salem, Naozumi Hiranuma et al. · 1 citation
Preprint Jul 2026

Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2

Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.

Vilya Research Pascal Sturmfels, Naozumi Hiranuma, M. Salem et al. · 0 citations
Open access Aug 2026

OmniScore: Universal Scoring of Diverse Biomolecular Complexes via Equivariant Geometry-Aware Discrete Representation Learning

Scoring biomolecular complexes is central to structure assessment and drug discovery, yet the complexes themselves vary widely in pose, size, and molecular composition. A scoring function tuned for one interaction type rarely carries over to another, and most existing methods compound the problem by leaning heavily on task-specific labels. We introduce OmniScore, a universal structure-based framework that learns a shared geometry-aware representation of complexes once and then adapts it to downstream scoring through lightweight task-specific heads. OmniScore couples a graph view and a sequence view of each structure, encodes its three-dimensional geometry, and compresses representations into a compact latent space that a reconstruction module and prediction heads can reuse. We pretrain this backbone on diverse datasets including complexes, monomers, and small molecules with complementary objectives: coordinate recovery, correcting corrupted input tokens, predicting molecular identity, and grounding the representation in structure-level physical quantities. Across the evaluated benchmarks, OmniScore gave the best antibody-antigen and nanobody-antigen quality assessment on all reported metrics compared to state-of-the-art baselines. Its frozen residue embeddings matched the state-of-the-art protein-tokenization method with an average functional-site accuracy of 71.8% on a standard residue-level benchmark. On protein-ligand scoring and ranking benchmarks, it performed on par with methods built specifically for that single task. These results suggest that geometry-aware pretraining can provide a reusable scoring backbone for tasks that depend on interfacial and residue-level structure, within the evaluated settings.

Tien-Cuong Bui, Junsu Ko, Juyong Lee · 0 citations
Aug 2026

HighMorph: De Novo Cyclic Peptide Sequence Design via Protein–Protein Interaction Recapitulation

Cyclic peptides have emerged as a compelling class of bioactive scaffolds, but de novo design of target-binding cyclic peptides from protein structures remains challenging. Here, we present HighMorph, an interaction-guided framework that combines protein–protein interaction information with artificial intelligence for rational cyclic peptide design. HighMorph integrates Monte Carlo tree search with a Transformer-based policy-value network to efficiently explore cyclic peptide sequence space, while incorporating explicit atomic-level hydrogen bond constraints extracted from reference protein–protein complexes to guide sequence optimization. The framework is systematically validated on two clinically relevant targets, programmed death-ligand 1 (PD-L1) and kallikrein-related peptidase 4 (KLK4). Notably, 33.3% and 40% of the generated candidates are active against PD-L1 and KLK4, respectively, with active cyclic peptides exhibiting micromolar binding affinities (approximately 10–6 M). These results validate our approach for cyclic peptide design. Additionally, interaction analysis provides insights for developing therapeutics targeting challenging protein interfaces.

M. Lan, Chengyun Zhang, Wentong Wang et al. · 0 citations
Aug 2026

Beyond Simple Mimicry: Next-Generation Geometric Architectures and Future Paradigms in Small-Molecule and Macrocyclic Peptidomimetics.

Peptidomimetics have matured from motif‑based inhibitors into a structural engineering discipline that systematically translates peptide recognition surfaces into drug‑like scaffolds. Driven by the urgent clinical demand to overcome the inherent pharmacological liabilities of biomolecules, the field is undergoing a decisive Peptide-to-Small Molecule paradigm shift-functionally converting peptide-derived recognition motifs into orally bioavailable synthetic therapeutics. This Perspective highlights how foundational geometric design principles-linear repetition, convergent fusion, and cyclization-define next‑generation architectures capable of targeting complex protein-protein interactions (PPIs). Repeating‑unit oligomers exemplify linear projection strategies, heterocycle‑centered scaffolds embody the convergent fusion of recognition motifs, and macrocyclic frameworks pre-organize bioactive conformations while enabling access to non‑canonical topologies. Beyond simple mimicry, these architectures increasingly embrace dynamic responsiveness, aggregation remodeling, and universal multi‑structure platforms. We argue that the convergence of geometric logic with automated synthesis and AI‑driven design will transform peptidomimetics into a primary modality for decoding and therapeutically engaging the human interactome, including historically "undruggable" PPIs.

Jesang Lee, Sumin Son, Jeong Yeon Yoo et al. · 0 citations