Skip to content
#protein folding Dataset Open access

Quadrupling the protein family space with global metagenomics

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Here you can find the data from the study "Quadrupling the protein family space with global metagenomics". For Protein Families: iso_clusters25_names.tsv.bz2 Description: TSV file of isolate clusters with more than 25 members. Columns: (1)Cluster name (2)Representative isolate protein header (3) Isolate protein member header (4)Isolate protein member sequence metag_clusters25_names.tsv.bz2 Description: TSV file of metagenomic clusters with more than 25 members. Columns: (1) Cluster name (2) Representative metagenomic protein header (3) Metagenomic protein member header (4) Metagenomic protein member sequence For Protein Folds structures.tar.gz Description: Contains three subfolders with protein structure models HQ:High-Quality (pTM ≥ 0.7) models (26,767 pdb files) MQ:Medium-Quality (0.5 ≤ pTM < 0.7) models(47,823 pdb files) LQ:Low-Quality (pTM < 0.5) models (82,018 pdb files) foldseek_results.tar.gz Description: You will find three folders (HQ, MQ & LQ). Each one contains three files: AF2.tblout (unfiltered hits to AlphaFoldDB) CATH.tblout (unfiltered hits to CATH) PDB.tblout (unfiltered hits to PDB) NMPFAMSDB2_MODELS_SCORES.txt Description: A txt file with pTM and pLDDT score for each family model. Columns: (1) Family name Column (2) pTm score Column (3)pLDDT score

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

In the context of cloud computing, risks associated with underlying technologies, risks involving service models and outsourcing, and enterprise readiness have been recognized as potential barriers for the adoption. To accelerate cloud adoption, the concrete barriers negatively influencing the adoption decision need to be identified. Our study aims at understanding the impact of technical and security-related barriers on the organizational decision to adopt the cloud. We analyzed data collected through a web survey of 352 individuals working for enterprises consisting of decision makers as well as employees from other levels within an organization. The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability. The result from our logistic regression analysis confirms the criticality of the security concern, which results in an up to 26-fold increase in the non-adoption likelihood. Our study underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

To compete in this age of disruption, large companies cannot rely on cost efficiency, lead time reduction and quality improvement. They are now looking for ways to innovate like startups. Meanwhile, the awareness and use of the Lean startup approach have grown rapidly amongst the software startup community in recent years. This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors. A multiple case study approach is followed in the investigation. Two software product innovation projects from two large companies are examined, using a conceptual framework that is based on the method-in-action framework and extended with the previously developed Lean-Internal Corporate Venture model. Seven face-to-face in-depth interviews of the employees with different roles are conducted. Within-case analysis and cross-case comparison are applied to draw the findings from the cases. A generic process flow summarises the common key processes of Lean internal startups. The findings suggest that an internal startup that is initiated management or employees faces different challenges. A list of enablers of applying Lean startup in large companies are identified, including top management support and cross-functional team. Both cases face different inhibitors due to the different process of inception, objective of the team and type of the product. Our contributions are threefold. First, this study is one of the first attempt to investigate the use of Lean startup approach in large companies empirically. Second, the study shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context. The third is a general process of Lean internal startup and the evidence of the enablers and inhibitors of implementing it, which are both theory-informed and empirically grounded.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

Affects--emotions and moods--have an impact on cognitive processing activities and the working performance of individuals. It has been established that software development tasks are undertaken through cognitive processing activities. Therefore, we have proposed to employ psychology theory and measurements in software engineering (SE) research. We have called it "psychoempirical software engineering". However, we found out that existing SE research has often fallen into misconceptions about the affect of developers, lacking in background theory and how to successfully employ psychological measurements in studies. The contribution of this paper is threefold. (1) It highlights the challenges to conduct proper affect-related studies with psychology; (2) it provides a comprehensive literature review in affect theory; and (3) it proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

It is essential for startups to quickly experiment business ideas by building tangible prototypes and collecting user feedback on them. As prototyping is an inevitable part of learning for early stage software startups, how fast startups can learn depends on how fast they can prototype. Despite of the importance, there is a lack of research about prototyping in software startups. In this study, we aimed at understanding what are factors influencing different types of prototyping activities. We conducted a multiple case study on twenty European software startups. The results are two folds; firstly we propose a prototype-centric learning model in early stage software startups. Secondly, we identify factors occur as barriers but also facilitators for prototyping in early stage software startups. The factors are grouped into (1) artifacts, (2) team competence, (3) collaboration, (4) customer and (5) process dimensions. To speed up a startup’s progress at the early stage, it is important to incorporate the learning objective into a well-defined collaborative approach of prototyping.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#human-computer interacti... Open access May 2025

Linker-free PROTACs efficiently induce the degradation of oncoproteins

Proteolysis-targeting chimeras (PROTACs) present a potentially effective strategy against various diseases via selective proteolysis. How to increase the efficacy of PROTACs remains challenging. Here, we explore the necessity of the linker, which has been deemed as an integral part of heterobifunctional PROTACs. Adopting single amino acid-based degradation signals, we find that the linker is not a required feature of the PROTACs. Notably, the linker-free PROTAC, Pro-BA, exhibits superior efficacy over its linker-bearing counterparts in degrading EML4-ALK and inhibiting lung cancer cell growth, as Pro-BA induces a stronger interaction between the target and the E3 ubiquitin ligase. Pro-BA is a water-soluble, orally administered degrader that significantly inhibits the tumor growth in a xenograft mouse model. The broad applicability of this linker-free PROTAC strategy is further validated through the development of BCR-ABL degrader. Our study introduces a design paradigm for PROTACs, potentially facilitating the advancement of more efficient therapeutic degraders. Linkers are traditionally seen as important for PROTAC activity. Here, the authors demonstrate that linker-free PROTACs can outperform traditional designs, marking a paradigm shift in PROTAC development for targeted protein degradation.

Jianchao Zhang, Congli Chen, Xiao Chen et al. · 41 citations
#machine learning Open access Nov 2025

mRNABERT: advancing mRNA sequence design with a universal language model and comprehensive dataset

Designing effective mRNA sequences for therapeutics remains a formidable challenge. Inspired by successes in protein design, language models (LMs) are now being applied to RNA, but progress is often impeded by the lack of comprehensive training data. Existing models are frequently limited to UTR or CDS regions, restricting their application for complete mRNA sequences. We introduce mRNABERT, a robust, all-in-one mRNA designer pre-trained on the largest available mRNA dataset. To enhance performance, we propose a dual tokenization scheme with a cross-modality contrastive learning framework to integrate semantic information from protein sequences. On a comprehensive benchmark, mRNABERT demonstrates state-of-the-art performance, outperforming previous models in the majority of tasks for 5’ UTR and CDS design, RNA-binding protein (RBP) site prediction, and full-length mRNA property prediction. It also surpasses large protein models in several related tasks. In conclusion, mRNABERT’s superior performance across these diverse tasks signifies a substantial leap forward in mRNA research and therapeutic development. Designing complete mRNA sequences for new vaccines and therapies is a complex challenge. Here, the authors develop mRNABERT, a foundational AI model that designs entire mRNA sequences and demonstrates superior performance across comprehensive benchmarks.

Ying Xiong, Aowen Wang, Yu Kang et al. · 22 citations · ⚡1

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.