Skip to content

Author

Xiangyu Zhao

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Open access Sep 2026

Chinese-SkillSpan: A Benchmark for Competency Span Extraction from Chinese Job Advertisements

Online job advertisements can reveal changing skill demand only when competency mentions are recoverable as auditable text spans. Chinese lacks a large resource that combines explicit span-boundary rules with coverage of different recruitment genres. We present Chinese-SkillSpan, a corpus of 22,840 sentences from four Chinese recruitment sources. It uses a flat, ESCO-derived inventory of language skills and knowledge, knowledge, skills, and transversal skills and competences (LSKT), with Chinese-specific rules for minimal-complete, non-overlapping spans. Language models propose annotation drafts, but human reviewers retain authority over accepted offsets and types under a shared handbook. The benchmark uses identifier-strict scoring, exact and overlap-tolerant metrics, and source- and length-based analyses. Its evaluation reference is undergoing final human adjudication, so the current results are provisional estimates. The baseline study shows why the resource is challenging: exact extraction changes with Chinese boundary conventions, and accuracy varies across recruitment sources. Taken together, Chinese-SkillSpan contributes a multi-source span resource, a Chinese-specific annotation and evaluation protocol, and a reproducible baseline and diagnostic suite that exposes category, source, and boundary effects. Internal artifact names are confined to the supplementary reproducibility record. Project materials are maintained at https://github.com/AlfredJamesLi/chinese-skillspan-benchmark. The pretrained JobBERT-zh model is available at https://huggingface.co/AlfredJames/jobbert-zh.

Guojing Li, Zichuan Fu, Junyi Li et al. · 0 citations
#large language models Open access Sep 2026

Chinese-SkillSpan: A Benchmark for Competency Span Extraction from Chinese Job Advertisements

Online job advertisements can reveal changing skill demand only when competency mentions are recoverable as auditable text spans. Chinese lacks a large resource that combines explicit span-boundary rules with coverage of different recruitment genres. We present Chinese-SkillSpan, a corpus of 22,840 sentences from four Chinese recruitment sources. It uses a flat, ESCO-derived inventory of language skills and knowledge, knowledge, skills, and transversal skills and competences (LSKT), with Chinese-specific rules for minimal-complete, non-overlapping spans. Language models propose annotation drafts, but human reviewers retain authority over accepted offsets and types under a shared handbook. The benchmark uses identifier-strict scoring, exact and overlap-tolerant metrics, and source- and length-based analyses. Its evaluation reference is undergoing final human adjudication, so the current results are provisional estimates. The baseline study shows why the resource is challenging: exact extraction changes with Chinese boundary conventions, and accuracy varies across recruitment sources. Taken together, Chinese-SkillSpan contributes a multi-source span resource, a Chinese-specific annotation and evaluation protocol, and a reproducible baseline and diagnostic suite that exposes category, source, and boundary effects. Internal artifact names are confined to the supplementary reproducibility record. Project materials are maintained at https://github.com/AlfredJamesLi/chinese-skillspan-benchmark. The pretrained JobBERT-zh model is available at https://huggingface.co/AlfredJames/jobbert-zh.

Guojing Li, Zichuan Fu, Junyi Li et al. · 0 citations