#model-evaluation
- LLM as a Judge
The heavy lifting nature of agentic evaluation, and the humans keeping it honest
- The Benchmark With Teeth
Why asking AI to draw Gary Busey tells us more than it probably should
- SWE-Bench and the Machinery of Doubt
The operational work behind a SWE-bench score that can survive scrutiny