An eval harness found what qualitative review couldn’t: AI models are most confident when wrongAugust 15, 2026