Silicon Sample Benchmark — Tier 3 (direct effect forecast) submission (team team_15)
Tier 3 (direct effect forecast) entry to the Silicon Sample Benchmark from team_15. We predict all 208 average treatment effects (16 message interventions x 13 preregistered outcomes) of a sealed ~18,000-person US megastudy on trust in climate scientists, blind and before any human data were released. The approach uses no synthetic respondents: a single large language model (gpt-5.6-sol) is prompted as a social-science expert and asked to forecast effect sizes directly. Queries are decomposed one outcome per call, with all 16 interventions compared within that outcome (Mode A), so the model ranks messages against each other on a shared scale rather than scoring them in isolation. Each outcome is queried under an ensemble of 10 instruction variants and the predictions are aggregated. Prompts state each outcome's own response scale explicitly, with orientation warnings for the reverse-coded items and dollar/binary scale warnings for the donation and newsletter outcomes. This record archives the prediction file, the generation code, and the completed method-registration form. No team member accessed, solicited, or was shown any human outcome data from the megastudy before the prediction lock.