Skip to content

Author

Jiamin Xiao

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Open access Sep 2026

Benchmarking Large Language Models on Long-Tail Plant Taxonomic Knowledge with PTTB-600

Plant taxonomic knowledge contains a long tail of infrequently encountered names, diagnostic characters, and nomenclatural decisions, yet model reliability across this distribution remains unclear. We developed the Chinese-language PTTB-600, comprising 200 general, 300 ordinary specialized, and 100 long-tail fill-in questions, and evaluated 31 large language models (LLMs) or run modes under closed-book conditions without retrieval augmentation. The first author drafted the question bank and answer key; three coauthors with doctorates in plant taxonomy reviewed them independently. All models scored at least 197/200 on general questions, and 21 achieved full marks. The six highest-scoring models answered 291-297/300 ordinary specialized questions (97.0-99.0%) but achieved 63.0-90.0% accuracy on long-tail fill-in questions. Gemini 3.1 Pro Preview ranked first at 587/600; ranks two through six formed a closely spaced cluster with no significant adjacent differences after Holm correction. Across 11 within-family comparisons, thinking-mode runs yielded 20-65 additional correct answers, chiefly on specialized and fill-in tasks. Factual errors were uncommon in routine undergraduate content and concentrated in the generation of rare genus names, fine diagnostic distinctions, and alternative nomenclatural treatments. Top-performing LLMs can provide reliable support for routine teaching under instructor oversight, whereas long-tail identifications and nomenclatural decisions require verification against authoritative sources.

Jian He, Jiamin Xiao, Hong Qu et al. · 0 citations