Work
Long-form research and writing.
- Does It Know It Can't? Capability and Self-Assessment Dissociate in Language Models July 30, 2026
Start with the author's failure. Before any frontier model touched this benchmark, I pre-registered predictions: Sonnet 4.6 would pass ~70% of the hard agentic tasks, Haiku 4.5 ~45%, GPT-4o-mini ~25%. The actual pass rates were 0%, 0%, and 0%.
continue → - The Handoff Study: What Survives When an Agent Stops July 30, 2026
This study is told in the order it happened, because the order is part of the finding.
continue → - The Handoff Study, Part II: Asymmetry and Trust July 31, 2026
Part I established that a structured 400-token briefing lets a fresh agent match uninterrupted work. Part II asks the two questions that result immediately raises: does it matter who writes the briefing versus who receives it — and what happens when the briefing lies?
continue →