Skip to content
Open access

Measuring and removing near-duplicate contamination in alignment-free SARS-CoV-2 lineage classification benchmarks

Aug 2026 · bioRxiv · 0 citations · 25 references
Biology

Abstract

Alignment-free lineage assignment from k-mer frequency profiles is widely used for SARS-CoV-2 surveillance, and the methods that do it are ranked against each other by margins of one or two percentage points. Those rankings rest on an unchecked protocol. Public repositories hold many near-duplicate genomes, and stratified random splitting puts members of such a group on both sides of the split, so a classifier is credited for sequences it has already seen. We propose quantised profile hashing, which finds near duplicates in k-mer feature space by rounding each frequency vector and hashing it. No sequence is compared with any other, so one pass over the feature matrix suffices and no similarity threshold has to be chosen. Rounding is also what makes the groups well defined, and they are then kept whole across the training, validation and test sets. On 255,611 genomes from seven Pango lineages, random splitting leaves 5.09% of test sequences with a near duplicate in training, on a benchmark ranked by margins of one or two points. Ten update rules were trained twice, identically except for the partition. The contaminated benchmark separates one rule from the leader at 0.05; the clean one separates none. The two orderings are uncorrelated, Kendall τ = +0.022, with rules moving 3.2 positions on average and the leader of one benchmark ranking eighth on the other. A ranking obtained under contamination therefore says nothing about the ranking without it, and the quantity worth reporting beside a score is the leakage rate of the split.

Read PDF

Similar papers

Preprint Jul 2026

modelDNA: Calibrated Lineage Verification and Merge Decomposition from Sampled Weight Fingerprints

The lineage graph of open-weight language models is self-reported: Hugging Face's base_model metadata field is optional and unverified, and over 60% of Hub models document no parentage at all. Methods for detecting lineage from weights exist in the research literature, but each ships as paper code tied to one signal and one experiment; when a provenance dispute breaks, the analysis is redone by hand. This report describes modelDNA, a tool that fingerprints a model from roughly 100-300 MB of ranged HTTP reads (instead of a full 15 GB download for a 7B model), compares the fingerprint against a reference database of foundation models across four published signal families, and returns one of eight verdict classes with a calibrated probability, preferring honest abstention to confident error. On a benchmark of 15 real Hub models with org-documented parentage, judged against 8 candidate bases (13 positives, 107 hard negatives), the system achieves AUROC 1.0, zero false positives at its reporting threshold, and 13/13 correct top-1 parent attribution. The report's second contribution is merge decomposition. Every mainstream weight-merging method is (near-)linear per tensor, and fingerprint sample positions are deterministic functions of tensor identity, so a merged model's fingerprint is the same linear combination of its parents'fingerprints. Mixture weights can therefore be recovered from fingerprints alone by sum-to-one constrained least squares. Against merges with published mergekit configurations as ground truth, the method recovers a slerp merge's layer-interpolation curves at r = 0.999 and a dare_ties merge's mixture weights to within 0.011 of the published values, without downloading any weights beyond the fingerprints. All fingerprints, benchmarks, and the inferred lineage graph of 55 models are public and reproducible offline.

Muhammad Adil, Saad Aamir · 0 citations
Preprint Jul 2026

A Deterministic Binary Fingerprinting Framework with Zero-Trained Feature Extraction for Sparse Count Matrices

Sparse count matrices from single-cell transcriptomes to k-mer profiles and document-term frequencies are conventionally analyzed via PCA-reduced graph clustering or iterative optimization in continuous embedding spaces. We introduce MMTB, a deterministic non-learned binary representation framework that requires no label supervision, model fitting, or gradient-based optimization. Column-wise Min-Max normalization followed by fixed cutoffs maps each sample to a thermometer fingerprint whose Hamming distances show empirical correspondence with normalized L1 distances, with Pearson correlation approximately 0.92 on single-cell RNA-seq pairs. In a favorable three-cell-line mixture, the 3-threshold fingerprint achieves NMI of 0.99 at 188 bytes per cell. Under a fair Hamming nearest-neighbor graph plus Leiden readout, MMTB approaches PCA plus Leiden on this coarse task. On challenging tissue-like annotations, continuous pipelines often lead; PBMC Seurat NMI is 0.39 for MMTB versus 0.49 for Scanpy, underscoring that MMTB is suited for coarse-grained separation and resource-constrained deployments rather than fine-grained subtype discovery or as a general replacement for continuous embeddings. Relative to dense float32 representations, MMTB fingerprints reduce memory by approximately 10-fold while providing fixed-width Hamming-indexable codes. PCA30 embeddings and sparse CSR may be smaller; we do not claim universal compression. A label-free suitability score is provided as a deployment guideline, not a performance predictor.

Lei Zhao, Fujin Huang, Ling Kang et al. · 0 citations
Open access Aug 2026

scFair: Geometry-Aware Gene Budgets and Same-Rank Extension for Highly Variable Gene Selection

Background Highly variable gene (HVG) selection begins almost every single-cell RNA-seq analysis. While ranking formulas have been compared extensively, the integer gene budget at which any ranking must be truncated is typically left to the user and habitually fixed near 2,000. Relying on such a convention carries hidden costs—lists that are too short erase subtle structure, whereas lists that are too long add noise and computational overhead. Moreover, because global rankings measure variance across all cells, markers for rare populations often lose the “variance vote count” to dominant bulk variation, leading to an unfair feature allocation at the hard cutoff. Whether this convention is defensible, and whether the budget and tail can be set from data without disturbing the ranking, has not been examined systematically. Results Under a frozen seurat_v3 ranking, k-sweeps across 18 labeled datasets show that n = 2,000 is ARI-optimal on 1 of 18 datasets and that the best available budget is worth a mean ARI gain of +0.033 over it, establishing cardinality as a real and largely unexploited design axis. We present scFair, a Scanpy-compatible HVG layer that automates list length alone: geometry-aware auto_n sets a base size k from multi-seed density and stability features of an intermediate embedding (trading a modest, intentional compute increase for a safer data-driven default), and a same-rank append step acts as a conservative safeguard against cutoff unfairness by adding a short near-miss tail. The ranking is never recomputed or reweighted. On the 18-dataset panel, the default path improved Leiden–label agreement over HVG@2000 (median ΔARI = +0.016; 13/5; Wilcoxon P = 0.0077) and outperformed the neighborhood-based selector triku at author defaults on 15/18 datasets (median +0.024; P = 0.004), while triku did not improve on HVG@2000. Controls locate the effect: cell-number-only rules do not beat HVG@2000, an FDR-chosen length imposed on the frozen ranking is flat, and a fixed HVG@2200 default is not a general substitute because it cannot produce the short lists that compact matrices call for. Conclusions A fixed budget near 2,000 HVGs is frequently suboptimal, and list cardinality is a separable design axis that can be automated without changing the ranking formula. Effect sizes are modest, the short-list branch rests on four datasets, and rule thresholds were developed with partial overlap to the evaluation panel.

Zhao Li, Aaron W. James, Shengxuan Li · 0 citations
Preprint Jul 2026

Benchmarking NACTI Species Recognition in Long-Tailed Regimes

As with most ``in the wild''collections of the natural world, the North America Camera Trap Images (NACTI) dataset exhibits long-tailed class imbalance, with the largest class covering over 50% of its 3.7M images. Building on the PyTorch Wildlife model, we systematically evaluate Long-Tail Recognition (LTR) methodologies to benchmark species recognition performance, including specialised loss functions and LTR-sensitive regularisation. Our optimised configuration achieves state-of-the-art 99.40% Top-1 accuracy on the NACTI test split, significantly outperforming standard baselines and previously reported top performances. To assess robustness under domain shifts (e.g., night-time captures, occlusion, motion-blur), we extend our evaluation across three independent reduced-bias test sets (including ENA-Detection, Caltech Camera Traps and Missouri Camera Traps). Across these out-of-distribution (OOD) evaluations, our LTR-enhanced model consistently demonstrates substantially stronger generalisation capabilities compared to standard cross-entropy approaches. However, qualitative and quantitative analyses underline that current LTR optimisations cannot fully overcome representational bottlenecks, resulting in catastrophic predictive breakdown for rare `Tail'classes under severe domain shift. For maximum reproducibility, all dataset splits, key code, and network weights are published with this paper at https://github.com/ZehuaLiuY/Species-Classification.

Zehua Liu, T. Burghardt · 0 citations
Book Open access Jul 2026

Clustering-Guided Knowledge Base for Multi-Objective Rule Mining on Imbalanced Datasets

Mining classification rules for the minority class is challenging not only because positive examples are rare, but because they concentrate in small, geometrically irregular subregions of feature space that unguided search methods systematically miss. We propose a clustering-based knowledge base that captures positive-class structure before search begins and uses it to focus both initialization and neighborhood exploration in Moca-I, a multi-objective local-search algorithm that mines interpretable rule sets by trading off minority-class recall, precision, and complexity. Rather than sampling attribute conditions blindly from the full discretized space, the knowledge base clusters positive-class instances, extracts per-cluster attribute ranges, and aligns them with the algorithm's discretization—seeding the initial archive with minority-class-informed prototypes and restricting neighborhood operators to locally relevant regions. We instantiate this framework with two clustering methods: Self-Organizing Maps (MOCA-ISOM), which preserve the topological structure of the positive-class manifold, and K-Means (MOCA-IKM), a centroid-based baseline. Evaluated on 19 imbalanced benchmark datasets, MOCA-ISOM achieves statistically significant F-measure improvements on eight datasets and Hypervolume improvements on nine, with gains up to +17% absolute on high-dimensional data. MOCA-IKM shows comparable average performance but exhibits five statistically significant degradations.

Raymonde Akiki, Chantal Saad Hajjar, M. Chamoun et al. · 0 citations