Skip to content

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Generic Vision and Cross-Attention for Reaction Yield Prediction

Traditional reaction yield prediction is constrained by 1D quantum descriptors that lack explicit spatial information. To address this gap, a dual-modal Vision Cross-Attention architecture is proposed, fusing tabular physical-organic data with 2D molecular topologies. Notably, it is demonstrated that a generic computer vision backbone processing simple 2D skeletal structures independently outperforms purely quantum-based baselines. By synergizing both modalities, superior predictive accuracy compared to traditional methodologies is achieved by the optimal cross-attention framework (Test RMSE = 5.27%). Through mechanistic probing, active, descriptor-guided spatial querying is observed, effectively offloading macroscopic steric identification to the visual pathway. Furthermore, a dynamic chemical hierarchy is learned by the network to heavily prioritize critical steric bottlenecks, such as the aryl halide. Concurrently, residual skip connections are utilized to protect non-spatial electronic parameters from destructive attenuation during fusion. Collectively, a scalable and highly interpretable blueprint is provided for augmenting physical chemistry with deep visual learning.

Qiwei Han, Chi Zhou · 0 citations
Preprint Jul 2026

ChemFusion: A Multimodal Cross-Attention Network for Reaction Yield Prediction

ChemFusion is presented, a hybrid neural network that fuses conventional electronic features with explicit 3D atomic coordinates and reveals that the architecture autonomously learns to identify and penalize restrictive steric hindrances, demonstrating that spatially aware networks can navigate complex reaction sterics that standard statistical models typically miss.

Qiwei Han, Chi Zhou · 1 citation
Preprint Jul 2026

Multimodal Molecular Representation Learning with Graph Neural Networks, Deep&Cross Networks, and SMILES Embeddings

This work introduces a parameter-efficient Tri-Branch Modular Fusion Neural Network that synthesizes three orthogonal modalities: 3D spatial geometry, discrete topological grammar, and explicit macroscopic physicochemical descriptors that offers a highly efficient alternative to brute-force parameter scaling.

Qiwei Han, Chi Zhou, Ruo-Yuan Wang et al. · 0 citations