#pipeline
- LLM as a Judge
The heavy lifting nature of agentic evaluation, and the humans keeping it honest
- SWE-Bench and the Machinery of Doubt
The operational work behind a SWE-bench score that can survive scrutiny
- Annotation Has a Control Room
The hidden operating layer behind the most ordinary annotation task.
- Holding the Signal
What it takes to keep a Human Data Programme coherent