Skip to content

Author

Dylan Nguyen

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#edge computing Open access Aug 2026

SESTINA v1.0: Exact Inference, Robust Evaluation, and Adversarial Testing in Six-Player Canadian Fish

Six-player Canadian Fish is a decentralised imperfect-information team game in which every action is public, so the hidden state reduces to the initial deal and the posterior over that deal can be computed exactly. We develop and evaluate SESTINA v1.0, the strongest configuration produced in this project’s FishBot lineage. It combines an approximate Sinkhorn fit to that posterior, started from a fitted policy prior, with a linear ask and declaration policy, a public-history tie-breaking rule that preserves common knowledge among teammates, a half-suit contestation weighting, a deduction-state stall detector in place of an event-count termination rule, and a guarded determinized test-time search. Evaluation follows a protocol registered in advance and run on sealed holdout material, using duplicate deal blocks, deal-clustered bootstrap confidence intervals, replication across two disjoint deal banks as an advance-specified criterion, calibrated detection floors, and mechanical side-channel controls. Against F-cheap, the cheapest configuration genuinely on the v0.6 frontier and the registered comparison target, SESTINA v1.0 achieves a +3.33 percentage-point win-rate edge (95% CI [+2.88, +3.78]) over 48,000 sealed games, with the sign replicating on both banks. The advantage persists under cross-play between independently trained runs and eight rule dialects. Under partner substitution against a v0.5 opponent 6 of 7 changed-partner rows stay positive, but the worst is −0.19 [−1.04, +0.67], which that battery does not resolve against the registered −1.00 collapse threshold. Separately, none of eight independently constructed adversarial searches found a positive edge at the tested budgets. SESTINA v1.0 does not, however, measurably outperform a composite configuration assembled earlier in the same programme (+0.15 pp, 95% CI [−0.29, +0.59]), indicating that the architecture work which followed added no measurable strength. Over a shared 31-member opponent panel SESTINA v1.0’s worst cell is −0.04 pp [−1.41, +1.33], which does not replicate in sign, and it is 3rd of four on minimax regret. Four candidate mechanisms failed to produce measurable improvement at this resolution. At the calibrated resolution of this evaluation the tested policy class appears locally flat, and we state what evidence a near-optimality claim would require.

Dylan Nguyen · 0 citations