When Polish Starts Hiding the Problem

Why polished AI artifacts demand sharper evaluation

The latest models can produce extraordinary artifacts. That makes their ordinary mistakes harder to catch.

We are entering a slightly awkward phase of AI capability.

GPT Astra, Claude, Fable 5.1, and other frontier models can now produce remarkably sophisticated work. They build animated presentations, structured reports, detailed spreadsheets, and complete deliverables in a single session. Some outputs would have required a designer, an analyst, and a project manager not very long ago.

The first impression is often excellent. The typography is consistent. The slides have movement. The workbook has a dashboard. The document has tasteful headings and an executive summary wearing its best suit.

Then you review the artifact properly.

One chart is slightly clipped. The summary figures are hardcoded. The formulas stop working when a row is added. The deck contains ten polished slides of AI slop and no discernible argument. The report uses confident language to disguise a missing conclusion.

Everything looks finished but very little feels intentional.

As an AI trainer, I find this more concerning than an obviously weak response. Visible mistakes are easy to catch. Polished mistakes arrive with good spacing and professional color palettes.

The models have reached extraordinary heights. Somewhere during the climb, a few of them have started forgetting to tie their shoelaces.

This is why artifact evaluation matters.

The artifact is the answer

When the requested output is a presentation, document, or workbook, evaluating the text response alone tells us very little.

The finished artifact carries the truth.

A presentation must communicate a coherent argument. A document must preserve the source, suit its audience, and remain readable. A spreadsheet must calculate correctly when its inputs change. All three should remain editable by someone who did not create them.

Artifact-based RLHF therefore needs to answer four questions:

  1. Did the model follow the request?
  2. Is the content accurate, complete, and useful?
  3. Is the artifact professional and usable?
  4. Would a real user choose this response over the alternatives?

That last question requires taste. The first three keep taste honest.

Start with the brief

Before evaluating the response, read the full prompt, source material, template, examples, and task-specific instructions.

A model can create a beautiful answer to a neighboring problem. This happens more often than it should.

Check whether the artifact covers every material requirement. Verify figures, terminology, conclusions, and citations against the supplied sources. Preserve the source’s level of certainty. A cautious finding should not emerge from the model wearing the confidence of a quarterly earnings call.

Outside knowledge can help identify suspicious claims. The supplied material remains the ground truth unless external verification was explicitly requested.

Length deserves no special reward. More slides, pages, formulas, or decorative rectangles do not automatically create depth.

Open the actual file

Artifact evaluation should happen in the application where the artifact will be used.

PowerPoint reveals clipped text, flattened charts, fragile positioning, and animations that have developed independent ambitions.

Word exposes broken page hierarchy, overflowing tables, stranded headings, and blank pages that appear to be taking personal time.

Excel reveals hardcoded outputs, inconsistent formulas, broken references, and dashboards whose numbers have only a ceremonial relationship with the source data.

Browser previews are helpful for quick inspection. They also hide rendering and editability problems with impressive loyalty.

Review every slide, page, sheet, chart, table, and relevant supporting tab.

File status

StatusHow to handle it
Opened normallyScore every required evaluation axis.
Repaired and usableRecord the repair and evaluate the visible artifact.
Repaired with missing contentIdentify what is missing and score what can still be judged.
CorruptedUse only when the file cannot be opened, rendered, or judged fairly after another attempt.
Table 1.1. File statuses and how to handle them during artifact evaluation.

Poor quality does not make a file corrupted. It simply makes the scoring portion of the exercise more decisive.

Evaluate the right dimensions

Different artifacts fail in different ways.

ArtifactEvaluation axesCommon hidden failures
Presentation deckContent, storytelling, writing, aesthetics, layout, editability, overallAttractive slides with no argument, weak hierarchy, clipped elements, flattened content
Presentation outlineContent, storytellingComplete-looking lists with no progression, prioritization, or takeaway
DocumentContent, aesthetics, overallUnsupported claims, generic prose, poor organization, broken pagination
SpreadsheetCorrectness, layout, formula usage, style, overallHardcoded results, fragile references, buried assumptions, decorative dashboards
Table 1.2. Evaluation dimensions and common hidden failures by artifact type.

These dimensions should be scored separately.

A deck may look excellent and tell no story. A document may be accurate and exhausting. A workbook may produce the correct answer once and collapse the moment someone adds Thursday’s data.

Assign the Overall score last. Overall quality reflects delivery readiness. It should not be calculated mechanically from the other scores simply because averaging feels scientific.

Ask:

  • Would someone genuinely present, share, or use this?
  • Can the intended audience understand it?
  • Can another person edit or extend it safely?
  • How much work remains before delivery?

A serious failure in correctness, instruction following, or usability should outweigh visual polish.

Use scoring anchors consistently

Presentations and spreadsheets use a seven-point scale.

ScoreMeaning
1Severely broken or unusable
2Major weaknesses require substantial repair
3Partly workable, with important revisions needed
4Adequate baseline with uneven quality
5Strong and usable with a few fixable issues
6Close to delivery-ready
7Ready to use with no meaningful weakness
Table 1.3. Seven-point scoring anchors for presentations and spreadsheets.

Documents use a six-point scale from zero to five.

ScoreMeaning
0Unusable or fundamentally off target
1Major problems prevent normal use
2A core requirement needs substantial revision
3Meets the basic goal with visible issues
4Strong and useful with low-impact weaknesses
5Complete, professional, and usable as delivered
Table 1.4. Six-point scoring anchors for documents.

For documents, a material instruction failure, fabricated claim, missing required citation, or blocking omission should cap the Content score at 2.

An elegant layout cannot rehabilitate an invented fact. It can only give the fact nicer accommodation.

Taste lives in the rationale

“Looks professional” provides almost no training signal.

“Bad formulas” is equally unhelpful.

A strong rationale identifies the evidence, location, consequence, and connection to the score:

The totals on the Dashboard sheet are hardcoded instead of referencing the Transactions table. The summary will therefore remain unchanged when new rows are added.

That explanation tells us what failed, where it failed, and why a user should care.

The same standard applies to strengths. Identify the slide sequence that builds a convincing argument. Name the section that handles uncertainty well. Point to the input, calculation, and output flow that makes a workbook auditable.

Taste in evaluation is disciplined specificity.

It recognizes when density carries substance and when it carries padding. It notices when a deck has visual rhythm but no intellectual movement. It can distinguish a concise artifact from one that has simply misplaced half the task.

Score first, rank second

Evaluate every response independently before comparing candidates.

Response A should receive the same score whether Response B is outstanding or appears to have been assembled during an electrical emergency.

Once the independent scoring is complete, rank the artifacts as finished deliverables.

Avoid counting how many slides, pages, or sheets each response “won.” Avoid mechanically averaging every axis. A major factual error can outweigh several smaller aesthetic advantages.

When scores are tied, rank based on the factor that matters most to actual use: correctness, audience fit, editability, clarity, or the amount of repair still required.

The ranking explanation should state the decisive tradeoff. Preference data becomes useful when we understand why one artifact deserves to win.

The better models get, the sharper evaluation must become

Before submitting an annotation, confirm that:

  • The correct artifact type and scoring scale were used.
  • Every required axis has a score.
  • Every score has specific, located evidence.
  • Claims were checked against the supplied sources.
  • Visual and structural defects were inspected in the native application.
  • The Overall score reflects delivery readiness.
  • The ranking agrees with the individual evaluations.
  • Low quality was not confused with file corruption.

Frontier models will continue producing more ambitious artifacts. The animations will improve. The layouts will become cleaner. The first impression will become increasingly persuasive.

Their mistakes will also become better dressed.

Artifact evaluation has to look past the finish and inspect the thinking underneath it. The final question remains wonderfully unfashionable:

Would you actually ship this?