Aug 2026· Molecular breeding· Vol 46· 0 citations· 44 references
Medicine
TL;DR
SNoptimizer, a user-friendly Shiny application that uses a genetic algorithm–based framework to optimally select discriminatory SNPs from large-scale genotyping datasets, provides a flexible solution for researchers seeking to reduce genotyping costs while maintaining high discriminative power.
Abstract
The ability to efficiently discriminate genotypes is a critical step in genomics-assisted breeding, population genomics, biodiversity studies, traceability along food chains, and germplasm management. However, identifying the minimal and most informative subset of SNPs capable of uniquely distinguishing a large set of individuals remains a computationally challenging task. Here, we present SNPoptimizer, a user-friendly Shiny application that uses a genetic algorithm–based framework to optimally select discriminatory SNPs from large-scale genotyping datasets. By leveraging the evolutionary principles of selection, mutation, and crossover, SNPoptimizer iteratively identifies compact SNP panels that maximize genotype resolution. The application supports HapMap-formatted and VCF genotype files and includes an optional second-round optimization for resolving putative duplicates. We benchmarked SNPoptimizer across three independent datasets, including a tomato diversity panel, 820 Cauliflower genotypes, and a soybean diversity panel comprising 30 million variants across 1,511 samples. Across the three datasets, panels of 17–22 SNPs yielded R-VDP values ranging from 0.8744 to 0.9973, with complete discrimination obtained in Dataset III, demonstrating robust performance across different datasets. Cross-tool comparisons revealed complementary trade-offs among discriminatory power, panel size, runtime, and run-to-run reliability. SNPoptimizer provides a flexible solution for researchers seeking to reduce genotyping costs while maintaining high discriminative power.
This review compares SNP detection programs such as GATK, BCFtools, FreeBayes, SAMtools, SAMtools, and DeepVariant and their algorithmic structures, namely pileup- based, haplotype-based, and machine-learning approaches and suggests that no single tool is the best.
Shikhi Baruri, Sunita Khanal· Nepal Journal of Biotechnolo...· 0 citations
SUMMARY
Genome-wide association studies (GWASs) have identified thousands of genetic variants associated with complex traits and diseases. However, explaining the mechanisms underlying phenotypic variation remains challenging. Here, we introduce SNPannotator, an automated post-GWAS analysis software package designed to streamline the interpretation of GWAS findings. Our pipeline implements a multi-step process that identifies proxy variants in high linkage disequilibrium (LD) with associated lead variants, then queries comprehensive resources (including Ensembl, the GTEx Portal, the eQTL Catalog, and STRING DB) for genomic position, deleteriousness, regulatory annotations, clinical significance, trait associations, expression (eQTLs) and splicing quantitative trait loci (sQTLs), and functional enrichment analyses and compiles the results into user-friendly reports. This package is implemented in the R programming language and includes auxiliary functions for variant lookup and LD exploration. SNPannotator provides a practical framework for efficiently deriving biologically meaningful insights from GWAS data and for assisting researchers in prioritizing candidate variants for functional validation.
AVAILABILITY AND IMPLEMENTATION
The SNPannotator package is available from the Comprehensive R Archive Network (CRAN) at https://cran.r-project.org/web/packages/SNPannotator. The development version and tutorial is available on GitHub (https://github.com/omicslaboratory/SNPannotator). The online version of the package is available at https://omicslab.org/snpannotator.
SUPPLEMENTARY INFORMATION
Supplementary data are available at Bioinformatics online.
Alireza Ani, I. Nolte, Zoha Kamali et al.· Bioinformatics· 0 citations
Faba bean breeding and genomics have seen steady progress in recent years, supported by genome sequences and high-density genotyping platforms. These tools have been valuable for trait mapping, diversity assessment, and genomic research, but they have limited routine use in breeding programs due to their relatively high cost. Recent progress in establishing an optimized, cost-efficient genotyping-by-sequencing protocol tailored to the large and complex faba bean genome has created the foundation for a more accessible genotyping solution.
Using this approach, we explored the genetic diversity of faba bean germplasm from various panels, providing a comprehensive representation of the crop’s genetic landscape. From this dataset, we identified and selected a high-quality set of informative SNP markers that are evenly distributed across the genome. Building on these resources, we designed a breeder-friendly 10K SNP chip.
The 10K SNP chip delivers high accuracy, broad genomic coverage, and affordability. The chip was validated across diverse germplasm panels, demonstrating strong clustering performance, high reproducibility, and applicability to breeding-relevant germplasm.
This platform offers a cost-effective alternative to higher-density arrays, enabling its integration into genomic selection, marker-assisted breeding, and diversity monitoring, ultimately supporting accelerated genetic gain and the delivery of improved varieties to farmers.
Hyeonah Shim, Hailin Zhang, Thomas Groß et al.· Frontiers in Plant Science· 0 citations
As genome-wide association studies and genetic risk prediction models extend to globally diverse and admixed biobanks, accurate, scalable ancestry deconvolution, also called local ancestry inference (LAI), has become crucial. LAI assigns ancestry to each genomic segment within an individual, enabling studies of population history and ancestry-associated haplotypic effects. Existing LAI methods scale poorly to biobank-scale data, to the distant past, and to large numbers of ancestries. Here, we introduce several independent LAI methods implemented in the Gnomix software suite, achieving higher accuracy and faster computational performance than all existing approaches and with portable models that can be shared without exposing individual-level training data. Gnomix is paired with Gnofix, a swift, scalable phase correction counterpart. We demonstrate performance on worldwide whole-genome data from humans and canids, leveraging high-resolution accuracy to localise ancient New World haplotypes in the Xoloitzcuintli, dating back over 100 generations. Code is available at https://github.com/AI-sandbox/gnomix. The authors present Gnomix, a local ancestry framework that delivers leading accuracy across diverse admixed datasets on both whole-genome and array data with high efficiency, together with Gnofix, its fast phasing-error correction counterpart.
Information about segregating haplotypes and structural variation (SV) can be extremely rich for a variety of applications in population genomics but remains largely inaccessible for many non-model species. Of the available methods, linked-read sequencing is especially promising for its low cost and scalability, but its adoption remains limited. One existing linked-read method is Haplotagging, which barcodes sequencing reads to reconstruct long molecules that encode haplotype information, with the potential to generate phased whole-genome data and detect structural variants. In this study, we present BLink-seq, a novel Haplotagging method that is compatible with standard short-read next-generation sequencing platforms, is locally reproducible with low-cost reagents, and is scalable for high-throughput sample processing. We optimized library preparation parameters, explored their relationship to linked-read library metrics, and validated phasing performance and structural variant detection in two evolutionary extremes: an experimental Drosophila melanogaster cross of inbred lines carrying known inversions, and four Atlantic silverside (Menidia menidia) parent-offspring trios sourced from highly outbred, wild-caught populations. We then applied our protocol to a cohort of 376 silversides to demonstrate its scalability and potential for SV detection and genotype imputation. Using BLink-seq, we generated chromosome-scale phased blocks and identified known inversions in both validation datasets. We discovered previously uncharacterized structural complexity within a known adaptive inversion on silverside chromosome 11, demonstrating that linked-read data can refine our understanding of SV architecture beyond what short reads alone can resolve. Finally, we provide a user guide for researchers interested in using BLink-seq.
Azwad R Iqbal, Pavel V. Dimens, J. Rick et al.· bioRxiv· 0 citations
Population genomic studies in non-model organisms increasingly depend on whole-genome resequencing, yet translating raw reads into reliable variant callsets remains a practical challenge due to the complexity of multi-step bioinformatics pipelines and the absence of species-specific best practices. Here we present a step-by-step protocol for snpArcher, a Snakemake-based workflow that takes raw sequencing reads and a reference genome as input and produces a filtered, joint-called VCF suitable for downstream population genomic analysis. We guide users through six phases: installation and environment setup, sample sheet creation, run configuration, execution on local or high-performance computing systems, quality control review using an interactive HTML dashboard, and downstream analysis, focusing on postprocessing and filtering. The QC dashboard aggregates individual-level metrics including principal component analysis, relatedness estimation, depth-missingness diagnostics, and admixture analysis to help identify batch effects, contamination, cryptic relatedness, and outlier samples before downstream analysis. We demonstrate the impact of sequential filtering steps on the site frequency spectrum and demographic inference using a dataset of 137 burrowing owl (Athene cunicularia) genomes, showing how removal of low-coverage individuals, sex-linked scaffolds, and regions of excess heterozygosity eliminates artifacts that would otherwise bias inference of population size history. This protocol is intended as a practical companion to the original snpArcher publication, enabling researchers working with non-model organisms to produce and evaluate analysis-ready variant callsets in a reproducible manner.
Cade Mirchandani, Abdelmajid Omarjee, Guillaume Achaz et al.· Molecular biology and evolut...· 0 citations