Benchmark-Task Heterogeneity in Misinformation-Related and Human–AI Text Classification
Abstract
Stylometric features and frozen sentence-transformer embeddings are widely used as low-cost inputs for misinformation-related text classification, but heterogeneous benchmark tasks are rarely compared under a common protocol. We evaluated term frequency–inverse document frequency (TF-IDF), a 22-feature stylometric battery, frozen sentence-BERT, and fine-tuned RoBERTa on 15 English-language benchmarks (n = 24,906), including five human–AI provenance corpora. A descriptive partition of the raw area under the receiver operating characteristic curve (AUC) values attributed 71.8% of the total sum of squares to task category means and 13.7% to benchmarks within categories; an exact benchmark label permutation yielded p = 0.000161, but category remained confounded with benchmark construction, source, domain, label mapping, and sampling. After correcting three label orientations, the mean off-diagonal transfer AUC was 0.492 for stylometry and 0.507 for sentence-BERT. Retained-character matching reduced the length-only AUC to approximately 0.50, while low-cost methods retained AUCs of 0.779–0.999 in four estimable human–AI benchmarks; AITextPile lacked common support. Alternative LIAR/LIAR2 mappings changed individual estimates and method ordering. Random intake sensitivity was method-dependent: the MAGE length-only AUC fell from 0.936 to a median of 0.476 and AITextPile from 1.000 to 0.560. The results support benchmark- and target-domain-specific validation, not a universal misinformation or AI-text detector.