Skip to content

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

Aug 2026 · 0 citations · 14 references
Computer Science

Abstract

When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an examination uses a non-additive grading scheme. The 2025 reform of Vietnam's National High School Graduation Examination demonstrates the cost of this substitution. In Part II of the exam, candidates evaluate four true/false statements per question. The grading is convex: the number of correct statements earns 0, 0.10, 0.25, 0.50, or 1.00 points. Identifying three statements correctly pays 0.50 points, not the 0.75 points that standard accuracy metrics would award. Because Part II accounts for 4.00 of the exam's 10.00 points, reporting accuracy inflates the score by rewarding partial knowledge that the state explicitly penalizes. We introduce THPT-Ladder, a benchmark of 632 items from 21 official exams across 11 subjects, graded exactly as the ministry grades its students. The ministry publishes the marks of over a million candidates, allowing us to place models directly into the human cohort. Across eight models, the official rubric pays 0.020 to 0.159 points less per Part II question than proportional credit. This shortfall changes a model's apparent competence. For Qwen3.5-27B on the 2025 History exam, a 0.042-point shortfall drops its standing from the 90th to the 77th percentile among 481,293 candidates. A model's accuracy does not predict this penalty. At Claude Sonnet 5's accuracy level, different distributions of errors yield scores varying from 0.869 to 0.932 points per question. Official marks depend on how correct statements are grouped, meaning standard benchmarks report a competence the institution would not certify.

View source

Similar papers

Preprint Jul 2026

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rates, and target prices often admit multiple reasonable answers. Existing benchmarks nevertheless tend to grade such outputs against a single expert reference. Using independently built analyst models for the same companies, we find that across 108 directed pairs covering 65 companies, the median single-reference score is 0.33, 92.6% score below 0.70, and no same-vintage pair agrees on implied price within 10%. Point-tolerance grading can therefore penalize disagreement already present among professionals. We introduce GAUGE, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer. GAUGE uses 1,001 vendor-classified analyst workbooks and a 196-task evaluation set, with a three-layer observed-practice envelope, 56 auditable facets, eight validity gates, and deterministic structural checks. We validate the benchmark with a 55-participant known-groups study, company-grouped cross-fitting, and judge-stability audits. On the failure-aware score $\phi_0$, senior analysts average 88.3, juniors 66.0, and finance students 43.2. Across 24 agents and 1,011 scored generations, the best agent scores 53.4, above the student mean but below every senior and most juniors. It passes 93% of mechanical facets and 78% of judgment facets, with a fleet-median gap of 26 points. Current agents are substantially stronger at model construction than valuation judgment. We release the methodology, a gated de-identified data tier, a controlled training split, a versioned 48-task evaluation core, and a withheld refresh pool.

Jiacheng Lu, Sinuo Wang, Wentao Zhao et al. · 1 citation
Preprint Jul 2026

Effort Matters in Score-Based Admissions: How Retaking and Aggregation Shape Test Scores

Observed standardized test scores are the result of an endogenous process: students strategically allocate effort across multiple retake attempts to improve their outcomes. Because students differ in their ability to make these investments, the interaction between applicant strategy and institutional scoring rules---such as the widely used Single-Sitting and Superscoring policies---can disparately distort observed scores. We develop a strategic framework where students allocate effort in response to different scoring policies. We show that Superscoring---the practice of combining the best section scores across attempts---introduces systematic score inflation through order-statistic selection over noise draws. This degrades signal accuracy and amplifies wealth-based disparities by disproportionately rewarding applicants who can afford repeated testing. Conversely, Single-Sitting---which keeps the best overall score rather than section-level scores---preserves signal fidelity but excludes high-ability students who lack the resources to prepare for all subjects simultaneously. Neither rule uniformly dominates; instead, they force a structural trade-off between statistical precision and fair outcomes. Finally, to address this, we propose three algorithmic interventions which either modify how scores from multiple attempts are combined, or apply a post-hoc correction to observed scores. Using simulations calibrated to 2025 College Board data, we compare standard scoring rules against these proposed interventions.

Christine Ling, Diptangshu Sen, Juba Ziani · 0 citations
Open access Aug 2026

When a tax claim enters the accounts: materiality and financial-statement diagnostics after Brazil's Tema 69

We examine whether the accounting importance of a favorable tax claim maps into cash realization. Brazil's Tema 69 excluded ICMS from the federal PIS/Cofins base and generated recoveries with heterogeneous legal, accounting, and administrative paths. Hand-verified disclosures show 13 high-confidence gross pre-tax rows with median values equal to 5.35% of lagged assets and 112.96% of positive prior-year net income; three reconcilable after-tax disclosures have medians of 2.28% and 52.20%. Using 8,173 CVM issuer-years for 1,072 issuers, we identify 38 first high-confidence auditor-visible episodes. Fixed-effects event-year associations for ROA, operating cash flow/assets, and their difference are statistically insignificant, whereas a broader 65-event cohort produces a positive gap. Within-treated medians rise for ROA and the gap, but optimal matching does not corroborate these shifts. The evidence is diagnostic, not causal. We conclude that documented Tema 69 claims can be material yet nonrecurring and should be analyzed through accounting-classification, recurring-earnings, ownership, and utilization bridges.

Vanessa Janiszewski · 0 citations
Review Aug 2026

The Price of Permission: Classification Uncertainty in Constrained Capital Markets

Shariah-compliant equity screening provides a transparent setting in which institutional rules determine who may own a stock. A binary label identifies current eligibility but not whether the feasible investor base is fragmented across standards or close to changing. We define this instability as classification uncertainty and formalize its investor-base consequence through permitted investor mass. In a 1999-2024 CRSP-Compustat panel of 13,188 securities classified under seven researcher-emulated Shariah rulebooks, screening-rule disagreement and proximity to active boundaries rank next-month screen-implied transitions. U.S. Fama-MacBeth diagnostics do not support an unconditional equal-weighted permission premium, and a September 2023 DJIM/S&P methodology change produces no robust matched repricing. The central event evidence uses 25 official Securities Commission Malaysia lists. The 410 inclusions already trading before the preceding review have positive but imprecise matched returns. Applying the pre-event turnover floor yields 295 inclusions with 1.76 percentage points over $[0,10]$ trading days ($p_{\mathrm{date}}=0.008$; $p_{\mathrm{wild}}=0.017$) and 2.25 points over $[0,20]$ ($p_{\mathrm{date}}=0.018$; $p_{\mathrm{wild}}=0.035$). Leave-one-date-out, first-inclusion-only, and mid-review placebo checks are supportive, although a joint 20-day pre-event test rejects. Ownership and demand-pressure diagnostics do not identify a unique marginal buyer or clean causal demand shock. The evidence supports treating classification risk as a portfolio-monitoring state. Official Shariah permission is associated with price effects in a recognized local market among sufficiently tradable securities; formal eligibility alone is insufficient.

Abdulrahman Qadi, A. Sharma, Francesca Medda · 0 citations
Preprint Aug 2026

Does a Structural Model Add Anything to the Closing Price? Calibrated forecasting, incremental information, and match leverage in the Italian Serie A

Studies of association-football forecasting routinely report three-way accuracy in the low fifties and present it as competitive with the betting market. Accuracy against a uniform benchmark answers the wrong question; the question worth asking is whether a model carries information a margin-free closing price has not already absorbed. We formalise that test as the fitted weight in a logarithmic opinion pool and apply it to nineteen complete Serie A seasons (7,220 matches). The answer is negative and stable. A Dixon-Coles model with tuned exponential decay attains 53.4% accuracy and a Ranked Probability Score of 0.1972 against the market's 0.1905; the paired difference is +0.0067 (95% CI [0.0046, 0.0088]) and the market wins in all seven test seasons. The fitted pooling weight on the structural model is 0.000, and the log-loss profile is monotone increasing in that weight on validation and test alike, so this is a boundary solution, not an optimisation artefact. Refitting the same machinery to shots on target yields a variant earning weight 0.35 against the goals model -- it carries information the goals model lacks -- and 0.000 against the market. Two structural signals, each informative about the other, both priced. The structural model is better calibrated than the market on the home-win margin (slope 0.995 versus 1.103) while clearly less sharp: the market's advantage is discrimination rather than honesty, which accuracy alone cannot distinguish. Value lies not in a better forecast but in what is built on a calibrated one. We define match leverage, the change in a club's probability of achieving a season objective between winning and losing a fixture, and compute it for ACF Fiorentina: an away fixture against a relegation rival carried 2.25x the leverage of hosting the eventual champions. The paper also documents and corrects errors in an earlier study of our own.

Y. Pitcan · 0 citations

SUERF Policy Brief

Unknown authors · 0 citations

Related blog posts