Results demonstrate that hierarchical modeling of 3D facial geometry enables interpretable, ontology-linked phenotype classification, though performance on rare leaf terms remains limited.
Abstract
FaceMesh2HPO is a framework for classifying facial phenotypic descriptors aligned with the Human Phenotype Ontology (HPO) to support clinical diagnosis. Using annotations from 124 clinicians across 10 disorders (107 HPO terms) combined with non-syndromic controls, we generated 3D facial meshes (478 landmarks) from 2D images and trained a hierarchical PointNet-based pipeline with cascading classification and feature elimination. The best models, incorporating 3D meshes, facial outline, and demographic metadata, achieved AUROCs between ~0.55 and ~0.89, with higher performance at parent nodes than leaf terms. External validation showed variable generalizability across disorders. Results demonstrate that hierarchical modeling of 3D facial geometry enables interpretable, ontology-linked phenotype classification, though performance on rare leaf terms remains limited. Improved data diversity and feature selection strategies are needed to enhance robustness and clinical utility.
Background: Human facial genetics remains challenging due to the complexity of facial morphology, its highly polygenic basis, and the combined influence of genetics, environment, age, and lifestyle on facial appearance.
Objective: To develop predictors of human facial appearance that could aid facial reconstruction in forensic anthropology and predict facial appearance using DNA-based methods alone.
Subjects and methods: Data from ~800 individuals, including Illumina GSA genotypes, EPIC methylation profiles, metadata, and 3D facial scans, were analyzed. Candidate facial SNPs were evaluated using MeshMonk with facial segmentation and landmark-based phenotyping. Associations were tested between 747 DNA variants and facial traits in 658 individuals. In addition, BMI-predictive CpG sites were validated, and a compact BMI prediction model was developed using Elastic Net regression on methylation data from 624 individuals.
Results: Using global-to-local segmentation approach we found 29 SNPs to be significantly associated with 22 segments of the faces. The facial landmarking approach identified 27 candidate SNPs to be significantly associated with 60 facial distances. Elastic Net regression yielded 30 CpG BMI predictors. In the validation set, the model achieved a mean absolute error of 2.9 kg/m². The current results represent a step forward in our research on facial prediction.
W. Branicki, Maria Wróbel, B. Subramanian et al.· Anthropologiai Közlemények· 0 citations
Autism Spectrum Disorder (ASD) is a neurodevelopmental condition that affects communication skills, social interaction, and behavioral patterns. Early detection is essential for timely intervention; however, conventional diagnostic methods remain time-consuming and subjective, as they rely heavily on clinical observations and expert judgment. This limitation highlights the need for an automated and objective approach to support early ASD screening. This study aims to analyze the performance, stability, and generalization of ASD classification using geometric features extracted from distances between facial landmarks. By representing facial morphology in terms of quantitative spatial relationships, this approach provides a more interpretable alternative to raw image-based methods. This study contributes by proposing a geometric feature representation based on facial landmark distances, providing a comparative analysis between linear and nonlinear classifiers, and ensuring robust evaluation through cross-validation. The dataset consists of 2,032 facial images, evenly distributed between children with ASD and those with typical development. A total of 68 facial landmark points were detected and used to compute pairwise Euclidean distances as classification features. Two classification algorithms, Logistic Regression and Extra Trees Classifier, were evaluated using 5-fold cross-validation to ensure reliable and unbiased performance estimation. The results show that Logistic Regression achieved an average accuracy of 89.91%, precision of 91.04%, recall of 88.56%, and F1-score of 89.76%. Meanwhile, the Extra Trees Classifier outperformed the linear model, achieving an average accuracy of 91.88%, precision of 92.69%, recall of 90.89%, and F1-score of 91.77%. Overall, both models demonstrated stable and consistent performance across validation folds, with the Extra Trees Classifier showing superior ability to capture nonlinear patterns in the data. These findings indicate that geometric feature extraction based on facial landmark distances is effective for ASD detection and has strong potential to be developed as an objective, interpretable, and efficient early screening tool using children’s facial images.
Y. Nurdin, Syifa Anzella, Melinda Melinda et al.· Jurnal Teknokes· 0 citations
AI-assisted facial phenotyping supports rare genetic disorder prioritization by retrieving visually similar diagnosed cases from facial image reference databases such as the GestaltMatcher Database (GMDB). Existing GestaltMatcher-based retrieval frameworks compare each test image with individual gallery images in a facial phenotype embedding space. However, this pointwise formulation does not fully exploit available evidence, because patients may have multiple images and disorders may be represented by multiple diagnosed gallery patients. We propose an inference-time multi-level evidence aggregation framework that improves facial phenotype retrieval without modifying the underlying GestaltMatcher-Arc encoder. The framework combines embedding-level patient aggregation of multiple images from the same individual, patient-weighted disorder centroids, and hybrid individual-centroid scoring to integrate test-patient observations, disorder-level gallery evidence, and local nearest-neighbor evidence. We evaluated the approach on GMDB v1.1.4 across disorders represented during training (GMDB-Freq), unseen disorders (GMDB-Rare), and multi-image patient subsets, using a unified gallery containing both GMDB-Freq and GMDB-Rare disorders. Multi-level evidence aggregation improved mean per-disorder top-$N$ retrieval accuracy across all evaluation subsets. Top-1 accuracy increased from 38.52% to 48.82% on GMDB-Freq and from 19.38% to 23.79% on GMDB-Rare. On multi-image subsets, top-1 accuracy increased from 46.12% to 60.94% on GMDB-Multi-Freq and from 18.54% to 26.71% on GMDB-Multi-Rare. These findings show that inference-time aggregation can improve next-generation facial phenotype retrieval without retraining the encoder, supporting a shift from isolated single-image matching toward multi-level aggregation of patient and disorder evidence for rare-disorder prioritization.
Alexander Hustinx, Carolin Kaffiné, Behnam Javanmardi et al.· 0 citations
Face2Gene is a clinical tool that leverages facial features to aid genetic diagnosis. The DeepGestalt application suggests potential diagnoses based on facial similarity, while the D-Score evaluates likelihood of an individual having dysmorphic features suggestive of a possible genetic diagnosis. Given performance variability across populations and limited data from South Africa, this study assessed clinical utility in South African children with neurodevelopmental disorders (NDDs). Facial photographs from 301 children were analysed. The cohort comprised three groups: 36 children with NDDs with confirmed molecular diagnoses, 176 with NDDs without molecular diagnoses and 89 unaffected children. Diagnostic (recognition) accuracy was measured by whether the confirmed diagnosis appeared in the top-1 or top-10 ranked algorithm-generated suggestions (DeepGestalt). D-Scores were extracted to calculate group differences. Among children with confirmed molecular diagnoses, accuracy was 19% (95% CI: 9-35%) (top-1) and 34% (95% CI: 20-52%) (top-10), improving to 33% (95% CI: 16-56%) and 61% (95% CI: 39-80%) when limited to conditions included in the DeepGestalt training set. One-way ANOVA revealed differences between participants with and without significant dysmorphic features, as assessed by clinicians. The D-Score demonstrated moderate sensitivity (78%, [95% CI: 0.64, 0.88]) and low specificity (42%, [95% CI: 0.38, 0.50]), but high negative predictive value (91%, [95% CI: 0.84, 0.95]), suggesting it may be more useful for ruling out dysmorphism; however, the low specificity indicates a high rate of false positives, even among clinically non-dysmorphic children. These findings suggest that under-representation of African populations may limit clinical performance and equity of AI-based facial phenotyping tools.
Z. Bruwer, Hendrike Mc Donald, Michal R. Zieff et al.· European Journal of Human Ge...· 0 citations
Down syndrome is associated with characteristic craniofacial features that have motivated the development of computer-vision-based facial-image-based screening systems. Recent studies have increasingly relied on deep computer vision and deep-learning approaches, but many provide limited interpretability, while earlier landmark-based methods offered transparent geometric and texture-based measurements. This creates a gap between interpretable handcrafted features and high-performing deep representations. To address this gap, this study proposes a hybrid interpretable–deep framework that combines landmark-derived geometry features, landmark-guided local binary pattern (LBP) texture descriptors, and frozen EfficientNetB0 convolutional neural network (CNN) deep features. The primary contribution of this study is the systematic integration and comprehensive evaluation of complementary interpretable and deep-feature representations within a unified facial-image screening framework. Feature fusion is followed by random forest feature ranking and SVM-RBF classification. Experiments were conducted on 2979 successfully processed facial images from an original dataset of 2999 images. Geometry-only, texture-enhanced, deep-feature, and hybrid fusion models were evaluated using repeated stratified train–test splits. The final RF Top-800 fusion model achieved strong facial-image classification screening performance, with F1 = 0.9045 +/− 0.0134 and AUC = 0.9675 +/− 0.0073 across repeated stratified train–test splits. Ablation analysis showed that removing geometry features, removing LBP features, or using only EfficientNetB0 reduced performance, supporting complementary contributions from interpretable geometry and texture feature components and deep-feature representations. Statistical comparisons and duplicate-sensitivity analyses further supported the robustness of the results. The findings demonstrate that landmark-derived geometry and texture descriptors remain valuable when integrated with modern deep representations, providing feature-level interpretability while improving screening performance through a hybrid framework that combines interpretable handcrafted features with high-performing deep representations.
Meshal Alfuraydi, H. Mathkour· Electronics· 0 citations
A comprehensive machine learning framework to classify ASD severity (mild, moderate, severe) is developed and validates by investigating the differential impact of feature engineering and selection, revealing a critical “evaluation paradox” where radical, unguided feature reduction improved geometric cluster cohesion but degraded clinical accuracy.
Arazo Mohammed Mustafa· ARID International Journal f...· 0 citations
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.