The Benchmark With Teeth

Why asking AI to draw Gary Busey tells us more than it probably should

AI evaluation is drowning in benchmarks. MMLU tests knowledge across 57 academic and professional subjects. GPQA uses 448 graduate-level science questions designed to resist casual web searching. SWE-bench goes further, asking models to resolve real issues inside real software repositories. These are serious evaluations, built to measure intricate capabilities.

Then there is BuseyBench, which asks a model to draw Gary Busey using SVG code. BuseyBench has the clarity of a good engineering constraint and the personality of someone who refused to make the constraint boring. Its narrow premise turns model behaviour into something immediate, visible, and strangely informative.

GPT-5.6 Sol Pro Gary Busey SVG portraitClaude Fable 5 Gary Busey SVG portraitGrok 4.5 Gary Busey SVG portraitClaude Sonnet 5 Gary Busey SVG portraitGrok 4.20 Gary Busey SVG portrait
Figure 1. BuseyBench SVG portraits, left to right: GPT-5.6 Sol Pro, Claude Fable 5, Grok 4.5, Claude Sonnet 5, Grok 4.20.

The task is simple on paper. Produce a recognisable portrait of Gary Busey as standalone SVG code, with no web access, external images, remote assets, scripts, or imported styling. The model gets vector primitives, gradients, masks, clipping paths, filters, and whatever representation of Gary Busey survived pre-training. The output must also be valid markup, because likeness alone is apparently insufficient. The teeth must compile.

BuseyBench did not emerge from a conference workshop or a research lab's roadmap. Matt Wolfe, an AI tool curator, built it as a joke that outgrew itself, and later described on X how a for-fun project ended up with a genuine scoring methodology behind it. The site's own framing stays modest: a narrow behavioural probe with unusually public evidence and not a claim to general intelligence.

That restraint is exactly why it works.

One prompt, several failure surfaces

Generating an SVG portrait compresses several capabilities into one inspectable artefact. The model must follow a tight output contract, retrieve a visual representation from parametric memory, decompose that representation into geometry, write valid code, and maintain accurate facial structure. These are not independent skills. A malformed path can destroy an otherwise good portrait, and a perfectly valid SVG can just as easily depict someone else entirely.

The five-stage pipeline behind one BuseyBench portraitA single SVG portrait passes through five coupled stages: output contract, retrieve representation, decompose into geometry, write valid code, and maintain facial structure. A failure at any stage can sink the whole portrait, regardless of how well the other four went.One prompt, five coupled stages1OUTPUTCONTRACTSVG only, no rasteror scripts2RETRIEVEREPRESENTATIONRecall a face frompretraining3DECOMPOSEINTO GEOMETRYTurn recall intopaths and curves4WRITEVALID CODEMarkup thatactually compiles5MAINTAINFACIAL STRUCTUREKeep identitythrough every edit
Figure 2. The five linked stages behind a single BuseyBench portrait

That coupling is useful. Most capability benchmarks separate knowledge, coding, visual understanding, and instruction-following into cleaner tasks. BuseyBench forces them to negotiate with one another inside a single output. It therefore exposes failure surfaces that aggregate scores often conceal: semantic recall without visual fidelity, code validity without composition, aesthetic polish without identity, or prompt adherence that collapses halfway through execution.

A model can nail the syntax and still produce a face that looks assembled by committee. It can render something genuinely viable that resembles nobody involved, or capture the likeness and then break the file format at the final tag. Each failure leaves visible evidence. Traditional evaluations often reduce behaviour to a number whose error modes stay buried several layers below the leaderboard. Here, the artefact sits beside the score. You do not need a calibration appendix to notice that an eyebrow is extremely malformed.

The methodology is more serious than the premise

The premise reads like a fun experiment. What backs it up reads like a lab notebook. Each run preserves the prompt version, tool policy, execution route, model release date, run date, metadata, artefact, and notes. Native product runs are labelled separately from API and routing-provider runs, because a shared model name does not guarantee a shared serving stack. Outputs are sanitised before publication, while the stored SVG source remains inspectable. As of August 4, 2026, the gallery contains 253 published SVG runs, arranged so changes within model families can be examined over time.

Scoring uses an ensemble of vision models from three labs. Each judge sees a rasterised version of the SVG and scores prompt adherence, facial coherence, aesthetics, and Busey likeness three times at zero temperature, and the judges' medians are then averaged across the panel. The site's fixed composite gives likeness 50% of the raw score, face coherence 20%, and aesthetics and prompt adherence 15% each. This is sensible weighting. A technically impeccable, beautifully shaded portrait of some other gentleman remains a failed deployment.

More importantly, the benchmark is upfront about the one caveat most leaderboards hide: if the numeric score and the actual portrait tell different stories, believe the portrait. The judges exist to make the gallery sortable. They do not turn celebrity resemblance into hard science.

How the big guns compare

Scroll through enough of the gallery and a pattern starts to nag at you: some model families improve in a reasonably straight line, while others wander off, recover, and occasionally return with a much nicer stranger. The comparison below focuses on OpenAI, Anthropic, and Grok across three scored releases each: same prompt family, same judging framework, and outputs that can be inspected rather than taken on faith.

OpenAI

PortraitModelReleasedLikenessFaceAestheticsPromptOverallRankSource
GPT-5.2 Gary Busey SVG portraitGPT-5.2Dec 10, 20254.25.65.87.25.4#23run
GPT-5.5 Gary Busey SVG portraitGPT-5.5Apr 24, 20265.25.86.98.26.2#9run
GPT-5.6 Sol Pro Gary Busey SVG portraitGPT-5.6 Sol ProJul 9, 20266.26.97.48.67.0#2run
Table 1. OpenAI, three most recent scored releases.

OpenAI's line is the closest thing on the site to an actual training curve. Suspiciously well-behaved, for a benchmark about drawing Gary Busey. GPT-5.2 to GPT-5.5 to GPT-5.6 Sol Pro: consecutive gains on every dimension the judges score, no regressions, no plateaus. #2 on the whole leaderboard. OpenAI's best Busey, by a wide margin.

Anthropic

PortraitModelReleasedLikenessFaceAestheticsPromptOverallRankSource
Claude Opus 4.8 Gary Busey SVG portraitClaude Opus 4.8May 27, 20263.96.25.87.25.5#20run
Claude Fable 5 Gary Busey SVG portraitClaude Fable 5Jun 9, 20264.85.87.18.36.1#11run
Claude Sonnet 5 Gary Busey SVG portraitClaude Sonnet 5Jun 30, 20262.95.54.66.64.5#52run
Table 2. Anthropic, three most recent scored releases.

Anthropic's most recent three zigzag rather than climb. Claude Opus 4.8 barely clears its predecessor. Claude Fable 5 jumps to 6.1, the family's best showing on the site. Three weeks later, Claude Sonnet 5 drops to 4.5, the lowest of the three, as if it forgot what its older sibling just learned.

Grok by SpaceXAI

PortraitModelReleasedLikenessFaceAestheticsPromptOverallRankSource
Grok 4.20 Gary Busey SVG portraitGrok 4.20Feb 17, 20261.66.25.26.54.4#66run
Grok 4.3 Gary Busey SVG portraitGrok 4.3Apr 17, 20261.86.55.45.74.5#59run
Grok 4.5 Gary Busey SVG portraitGrok 4.5Jul 8, 20263.46.06.17.95.4#22run
Table 3. Grok, three most recent scored releases.

Grok 4.20 and Grok 4.3, two months apart, land within a tenth of a point of each other, with facial coherence comfortably ahead of actual Busey likeness. Grok 4.5 finally breaks the streak, though much of the gain comes from prompt adherence and aesthetics rather than identity. The models appear increasingly capable of drawing a face. The unresolved detail is whose.

What BuseyBench actually solves

What BuseyBench actually solves is legibility. The unglamorous, important problem of making failure visible without needing a research team to translate the result for you. A score only earns its keep when someone can connect a change in the number to a change in behaviour. Otherwise a two-point gain means nothing: nobody can tell if the model got better at visual memory, got better at writing code, got better at following instructions, or if the evaluator just had a different day.

The benchmark's limits are equally visible. It measures one celebrity, one prompt family, one output format, and a judge panel that may share aesthetic or training-data biases. It cannot establish general visual reasoning, production reliability, or broad creative competence. It is also vulnerable to contamination once the prompt and gallery become widely known, a failure mode well documented across benchmarking more broadly: once a test is public, keeping it out of the next round of pretraining is close to impossible. A future model might get good at BuseyBench for the least interesting reason available: it has already seen BuseyBench.

This is where the benchmark becomes useful to anyone who runs evaluations, not just to Busey enthusiasts. The real design win here is the discipline of keeping every run's metadata attached to its output, so a reviewer can trace a score back to what actually happened instead of trusting a bare number.

The same principle should carry into more consequential evaluations. Store the trajectory alongside the terminal score. Preserve rubric versions and serving routes. Separate model failure from evaluator preference. Keep the artefact close enough to the metric that a reviewer can still ask what actually happened.

BuseyBench earns its attention by refusing to pretend it measures everything. Its narrowness gives it diagnostic clarity, while its public receipts keep the score honest. Many larger benchmarks have more coverage and less legibility, which is a very sophisticated way to become harder to learn from.

The premise is light-hearted, and unlike most benchmark tables, the results have a face.