Preprint
Aug 2026
ChemDIRT: A Diversified Instruction, Representation, and Task Benchmark for Robust Chemistry-LLM Evaluation
This work introduces ChemDIRT (Diversified Instruction, Representation, and Task Benchmark), a comprehensive evaluation framework designed to assess the robustness of chemical reasoning in LLMs and benchmark a diverse set of open- and closed-source LLMs.
Eric Inae, Tim Gunn, Chris Bond et al.
· 0 citations