Work
Long-form research and writing.
- Does It Know It Can't? Capability and Self-Assessment Dissociate in Language Models July 30, 2026
**Start with the author's failure.** Before any frontier model touched this benchmark, I pre-registered predictions: Sonnet 4.6 would pass ~70% of the hard agentic tasks, Haiku 4.5 ~45%, GPT-4o-mini ~25%. The actual pass rates were 0%, 0%, and 0%.
continue →