Skip to content
Open access

Expanding all-α-helical protein space through rational computational design

Jul 2026 · bioRxiv · 0 citations · 1 references
Biology

Abstract

De novo protein design is advancing rapidly1,2. This is being driven by AI to generate protein backbones, sequences, and structural models3–7. As a result, de novo designed proteins are becoming larger and more complex8–10, and increasingly explore new protein structures11,12. By contrast, natural proteins have evolved structural and functional complexity by modular combination of recurring protein domains13. Approximately 25% of these natural domains are mostly α-helical structures14. Here we show how these can be expanded using rational computational design. Following the domain classification scheme CATH15, we build complex all-α de novo proteins hierarchically using sequence-to-structure relationships for helix-helix interactions, systematic rules to connect helices, computational tools to design loops, and in silico evaluation. The pipeline starts with a target architecture of free-standing helices. These are connected into a topology by considering local arrangements of helical bundles using understood sequence-to-structure relationships for helix packing. Single-chain sequences are completed using template- and AI-based methods. Finally, AlphaFold models are assessed to give small numbers of designs for experimental validation. We test 31 designs for 14 different architectures and 25 topologies. 75% of these express as stable, monomeric, water-soluble proteins; and >30% yield X-ray crystal structures matching the designs to atomic accuracy and with new-to-nature structures. Finally, several of the scaffolds are functionalised through one-shot designs to deliver ion, small-molecule and protein binders.

Read PDF

Similar papers

Open access Aug 2026

Intra-Protein Interfaces Control Folding Dynamics and Mechanical Stability in a De Novo Designed Repeat Protein

Together, these results show that designed repeat-protein folding is governed by seed formation, interface propagation, and terminal boundary conditions, and establish intramolecular crosslinking as a strategy for rationally reshaping folding landscapes in designed proteins.

Melanie Weiß, Anna Lisa Heit, L. Milles et al. · 0 citations
Aug 2026

HighMorph: De Novo Cyclic Peptide Sequence Design via Protein–Protein Interaction Recapitulation

Cyclic peptides have emerged as a compelling class of bioactive scaffolds, but de novo design of target-binding cyclic peptides from protein structures remains challenging. Here, we present HighMorph, an interaction-guided framework that combines protein–protein interaction information with artificial intelligence for rational cyclic peptide design. HighMorph integrates Monte Carlo tree search with a Transformer-based policy-value network to efficiently explore cyclic peptide sequence space, while incorporating explicit atomic-level hydrogen bond constraints extracted from reference protein–protein complexes to guide sequence optimization. The framework is systematically validated on two clinically relevant targets, programmed death-ligand 1 (PD-L1) and kallikrein-related peptidase 4 (KLK4). Notably, 33.3% and 40% of the generated candidates are active against PD-L1 and KLK4, respectively, with active cyclic peptides exhibiting micromolar binding affinities (approximately 10–6 M). These results validate our approach for cyclic peptide design. Additionally, interaction analysis provides insights for developing therapeutics targeting challenging protein interfaces.

M. Lan, Chengyun Zhang, Wentong Wang et al. · 0 citations
Aug 2026

Nucleotide-binding motifs nucleated folding of the first enzymes

It is concluded that the early emergence of the Rossmann fold reflects the chemical and physical constraints of protein folding, explaining both its profound antiquity and sustained longevity.

Koh Seya, Tatsuya Corlett, Hamza Giaffar et al. · 0 citations
Jul 2026

Rep3D : an algorithm to identify structurally similar motifs

Repeats in protein structures act as essential structural building blocks, commonly forming multiple complex structures and functional units. Identifying sequence repeats in the primary structure alone is not sufficient to find the function of the proteins. Therefore, a new method, repeats in the three-dimensional structure of proteins ( Rep3D ), has been developed using a dynamic programming approach. This method enables rapid and accurate identification of structural repeats by calculating the distance between Cα atoms in the peptide backbone. A standalone computing version of the tool has been developed and implemented in Python. The Rep3D source code, along with documentation, examples of use and instructions, is available in the Rep3D repository (https://github.com/srimaha0801/Rep3d-Stand_Alone_v1/).

Gurleen Kaur, Madhumathi Sanjeevi, Srimaha Gandhi et al. · 0 citations