Skip to content

Author

Abdulmujeeb T. Onawole

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Trust-Aware Sequence-to-Function Modelling in Regulatory Genomics

Objective Sequence-to-function models increasingly predict regulatory activity, such as chromatin accessibility, directly from DNA sequence, and are used to interpret non-coding genetic variation. Standard accuracy metrics, computed over a held-out set of genomic regions, do not establish whether an individual prediction remains reliable once the input sequence departs from that set, nor whether a model’s attribution-based explanation is biologically grounded rather than coincidental. We develop and evaluate RegTrust-XAI, a trust-aware framework separating these questions using three inference-time signals: ensemble consensus, motif-grounded attribution coherence, and applicability-domain distance. Methods A five-model convolutional ensemble was trained on 517,790 K562 ATAC-seq windows and evaluated on a held-out chromosome test set (chr8/chr9, n = 42,844). Consensus, coherence, and applicability-domain distance were each tested against prediction error, alongside complementary sequence-novelty analyses and validation against an independent lentiMPRA reporter assay and saturation-mutagenesis MPRA data at the PKLR promoter. Results The ensemble reached Spearman ρ = 0.782, with skill of 0.328 over a constant-value null predictor. High-consensus predictions (Scenarios A+B) were consistently enriched for lower error than low-consensus predictions (Scenarios C+D), and attribution coherence further separated error within the high-consensus population (mean absolute error 0.396 versus 0.435, p = 9.6e-10). Applicability-domain distance showed a monotonic error gradient across six distance bands. A 4-mer composition-divergence metric was negatively associated with error and anti-correlated with applicability-domain distance, so composition-based and model-relevant novelty are not equivalent. Attribution transfer to lentiMPRA was assay- and subgroup-dependent, and predicted allele-substitution effects correlated with measured saturation-mutagenesis effects at the PKLR promoter at both 24 h and 48 h (ρ = 0.227 and 0.235). Motif-specific perturbation further showed that regulatory attributions were strongly context-dependent, with more than 90% of multi-instance motif modules exhibiting superadditive joint effects. Conclusions Prediction reliability, explanation validity, and sequence novelty are related but distinct properties of a sequence-to-function model. Evaluating each explicitly gives a more complete basis for deciding when to act on a prediction than accuracy alone.

Abdulmujeeb T. Onawole, Sulaimon Basiru, M. Sanni et al. · 0 citations
Jul 2026

tmGNN-XAI: An Explainable Graph Neural Network Tool for Predicting Electronic Properties of Transition Metal Complexes from SMILES.

Predicting the electronic properties of transition metal complexes (TMCs) from 2D molecular graphs remains challenging; organic-trained property models lack TMC transferability, universal interatomic potentials require 3D coordinates rather than SMILES, and tools providing holistic electronic property prediction with atom-level explainability and calibrated uncertainty remain limited. We present tmGNN-XAI, a multitask relational graph convolutional network that predicts seven quantum-chemical properties of TMCs directly from SMILES strings and produces perturbation-based atom-level attributions for each prediction. The model encodes dative coordination bonds as a dedicated edge type distinct from covalent bonds and is trained on 100,703 complexes from the tmQM data set spanning 30 transition metals. Test-set performance is competitive with a Chemprop D-MPNN baseline, achieving R2 = 0.979 for metal partial charge and R2 = 0.964 and 0.949 for HOMO and LUMO energies. Across all 100,703 complexes, donor atoms (N, O, S, P) appear among the top-five most important atoms in more than 99.8% of complexes for every property, a large-scale data-driven result consistent with ligand field theory. A trust framework combining ensemble agreement with attribution direction separates predictions into four reliability scenarios; confident predictions achieve 1.6 to 2.5 times lower mean absolute error than uncertain ones for five of seven properties. The framework generalizes to cross-level DFT validation, phototherapy candidate screening (area under the ROC curve (AUC) = 0.735), and indirect redox prediction via Koopmans' theorem. An interactive web application makes property predictions, atom-level attributions, and trust labels accessible without programming or DFT expertise. tmGNN-XAI is designed as an explainable, first-tier screening tool for TMC electronic property estimation.

Abdulmujeeb T. Onawole · 2 citations