Runs in Claude Code, Codex and ChatGPT · works on anything you build

Rejected · 6.92 / 10

Your AI thinks everything it makes is brilliant. These judges don't.

/perfection summons a panel of world-class expert judges matched to whatever you just built. They score it 0 to 10 with cited evidence, REJECT it with a precise fix list, and re-judge after every fix. The pass mark is computed in code, so no AI can sweet-talk its way through. Work is not done until it clears the bar.

Three editions in one zip: Claude Code, Codex, and ChatGPT. Installed in under a minute.

ROUND 1
6.92
REJECTED, no guarantee on the page
ROUND 2
7.52
REJECTED, proof stack too thin
ROUND 3
7.84
4 of 5 judges vote SHIP
RESULT
LIVE
a real sales page, measurably stronger

Exhibit zero · seventy four seconds

The whole idea, told as a story. Sound on.

1 min 14 sec

Every maker owns a magic mirror that calls everything perfect. This is what happens when five judges who cannot be flattered walk in, and how a build climbs from three past eight, the pass mark, toward ten.

Case file 01 · The problem

AI grades its own homework.

Ask any AI to build you a page, an app, an email, and it hands the work back glowing about what it just made. First drafts ship because nobody in the loop has a reason to say no. That is how you end up with a folder of finished-looking things that never quite convert, never quite feel premium, and never get touched again.

The missing ingredient isn't a better builder. It's a judge with the power to reject.

Case file 02 · How it works

Summon. Judge. Gate.

I

The right experts show up

Say /perfection and the skill reads what you built, then convenes the matching panel. A sales page gets a direct-response master, a scarred buyer, a funnel mechanic, a design director and a compliance skeptic. An email gets three specialists. A small asset gets one elite judge. You never pick a rubric.

II

Evidence only, no vibes

Before anyone judges, the skill gathers hard evidence: screenshots, link checks, the real fonts and colors in use, working demos actually clicked. Judges may only cite that evidence. A complaint with no citation gets struck. An invented problem gets caught by a mechanical text check and downgraded.

III

The gate is math

PASS requires every judge at 8.0 or higher, every scoring dimension at 7.0 or higher, and zero deal-breakers. The verdict is computed in plain code from the scores. Not the judges, not the chairman, not even the AI running the loop can wave work through. On REJECT you get a ranked fix list, the fixes get made, and the panel sits again.

Case file 03 · The bench

Seven panels, one for every kind of work.

What you builtWho judges it
Sales page / funnelFull 5-judge panel plus a chairman. Money pages never get the small treatment.
Website / landing pageConversion critic, design taste judge, technical QA judge, target buyer
App / product UIProduct design judge, UX flow judge, technical QA judge, microcopy judge
Marketing emailEmail conversion judge, brand compliance judge, skeptical subscriber
Video scriptRetention judge, voice judge, conversion judge
Deck / presentationStory judge, slide design judge, the person in the back row
Anything elseThe single best person alive at that exact craft, a ruthless skeptic, and your real audience

The Direct-Response Master

Halbert energy. Does the hook stop the scroll, is the offer unmistakable, is the proof credible, does every section earn the next one. Impatient with pretty pages that don't convert.

The Scarred Buyer

Has bought 100 products and finished none. Hunts the exact line where trust breaks: the overclaim, the fake-feeling scarcity, the guarantee that reads like a trap. The hardest judge to please, on purpose.

The Design Taste Judge

A design director who has rejected work over a single wrong font. Typography, spacing discipline, palette restraint, premium feel, and the AI-slop test: would a stranger instantly believe a machine made this?

The Compliance Skeptic

Protects YOU. Flags every claim that could burn you: income claims without proof, scarcity you can't substantiate, delivery promises the back office can't keep. Never touches your offer terms, only warns you.

Case file 04 · Ways to run it

Four speeds. From a quick look to relentless.

ModeSayWhat happens
Speed Scan"speed scan"The low-cost first look for smaller plans. One judge takes every expert's seat for a single fast read and hands you the top five fixes, at about a fifteenth of a full run's usage. A scan never marks work as cleared.
Default/perfectionUp to 3 judge-fix-rejudge rounds, then an honest scorecard of anything still short of the bar. You decide what happens next.
Relentless"relentless"Up to 5 rounds, and the later rounds escalate to a more powerful judging model. For work that has to clear the bar.
Until it passes"until it passes"No round cap, with two built-in safety brakes: a hard ceiling, and a stall detector that stops the loop the moment scores stop genuinely improving.

Panel size, auto-scaled

  • Small asset, one email or banner: 1 elite judge
  • Standard build, a page or feature: 3 judges + chairman
  • Money page or full build: full panel + chairman
  • Force it either way: "one judge" or "full panel"

What can never happen

  • A judge approving work while flagging a deal-breaker: auto-rejected by the gate
  • An invented finding forcing a rejection: downgraded for review
  • A polite judge rubber-stamping: every pass still names the 3 things closest to failing
  • An infinite loop: ceilings and a stall detector, computed in code
Case file 05 · Exhibits

What a run actually looks like.

Every exhibit below is recreated from a real /perfection run on a real, live sales page. Nothing staged.

Exhibit AYou say one word. The court convenes.
Claude Code
> /perfection

Judging this as a SALES PAGE with a 5-judge panel, default mode.
Say "one judge" or "full panel" to change.

Gathering evidence: screenshots (desktop + mobile), link checks,
live demo clicked like a buyer, fonts and colors read from the page.
Cost heads-up: 5 judges + chairman across up to 3 rounds,
or say "one judge" for about a quarter of that,
or "speed scan" for a fast first look at a fraction of it.
Exhibit BA judge files a verdict. Every gap cites evidence.

The Scarred Buyer

4 / 10 on guarantee
"There is no guarantee, refund window, or safety net of any kind written anywhere on the page, on a $397 one-time payment aimed at people who have been burned by past purchases. A buyer who has been burned before will notice the missing safety net right at the moment they're deciding whether to pay."

EVIDENCE: index.html lines 576-603 and all 7 FAQ entries read in full · VERDICT: REJECT

Exhibit CThe gate overrules a soft judge. This is the whole point.
round 3 of 3
Email Conversion Judge: verdict SHIP, overall 7.6

GATE: REJECT  (7.6 < 8.0, the bar is not negotiable)

The judge went soft. The math didn't.
Round cap reached. Writing your honest scorecard.
Exhibit DThe scorecard you actually read. Plain English, real trajectory.
Perfection Scorecard · StepMate sales page · 3 rounds Not cleared · 4 of 5 ship
JudgeR1R2R3Final verdict
Funnel Mechanic7.27.68.6SHIP
Design Taste Judge8.18.08.1SHIP
Compliance Skeptic7.07.58.1SHIP
Direct-Response Master5.87.38.0SHIP
Scarred Buyer6.57.26.4REJECT

Between those rounds the judges forced a real guarantee onto the page, caught hero buttons cut off on the exact screen webinars run at, proved the buy button lands on the correct checkout by clicking it, and killed two contradictions about what buyers actually get. The page that came out the other side is live and selling.

Case file 06 · The receipts

Proof it can't be flattered.

Self-verdict · Not cleared

It judged ITSELF and refused to pass.

4 rounds

The skill's own judges tore into its own code and instructions across four full rounds, found 25+ real defects, caught two fabricated findings and struck them, and still refused to give itself a pass against its own 8.0 bar. A quality gate that won't flatter its own maker won't flatter your work either.

Field result · Deployed

It rebuilt a live money page.

6.92 → 7.84

Three rounds on a real $397 sales page: a double-money-back guarantee forced into existence, the buy path click-proven to the checkout, roughly 24 judged fixes applied, and 4 of 5 world-class judges moved from REJECT to SHIP. The stronger page is live right now.

Case file 07 · Using it yourself

Installed in a minute. Judged in one word.

You need Claude Code, the version of Claude that runs on your computer. That's it. The skill runs on the Claude plan you already pay for.

Unzip the skill folder

Drop the perfection folder into ~/.claude/skills/ on Mac, or your Claude skills folder on Windows. No build, no install, no accounts.

Open Claude Code on any project

Finish something first. The judges grade FINISHED work: a page, an email, an app screen, a script. /perfection is the last step before you call it done, not the first.

Say the word

Type /perfection or just say "send in the judges." It detects what you built, names the panel and the mode in one line, and starts gathering evidence.

Read the scorecard, make the call

Rejections come with a ranked fix list and Claude applies the fixes between rounds. When it clears, ship. When it caps out, you get the honest gap list and YOU decide: push on, ship as is, or park it.

Words the skill listens for

  • /perfection · start a run in default mode
  • send in the judges · same thing, said like a human
  • speed scan · one fast, low-cost first look, top 5 fixes
  • relentless · 5 rounds, escalating judges
  • until it passes · uncapped, with safety brakes
  • one judge / full panel · force the panel size

Good to know

  • Runs on YOUR Claude plan. A solo-judge round is light; full money-page panels cost more and warn you first, with a cheaper option offered.
  • Judges never touch your offer, your price, or your terms. Anything in that territory gets flagged to you, never changed.
  • Everything is logged to a state file in your project, so you can stop and resume any time.
Case file 08 · Questions

Asked and answered.

Do I need to know how to code?

No. You talk to it in plain English and it reports back in plain English. The scorecard is a table a non-developer reads in ten seconds. The judges do the technical reading for you.

What do I need for it to work?

Claude Code on Mac or Windows, on any paid Claude plan. If you already build things with Claude Code, you have everything.

Can Claude talk the judges into passing bad work?

No, and this is the whole design. The pass/fail decision is computed in plain code from the scores: every judge 8.0+, every dimension 7.0+, zero deal-breakers. A judge who says SHIP while flagging a serious problem gets auto-rejected for contradicting itself. We know the gate holds because the skill judged its own code for four rounds and never once passed itself.

What if my work never clears the bar?

Default mode stops after 3 rounds and hands you an honest scorecard: exactly what's still short, how bad each gap is, and what a 10 would look like. You stay in charge. Plenty of real work ships at 7.8 with its remaining gaps known and accepted. What never happens is a fake pass.

What does a run cost to execute?

It runs on your existing Claude subscription, not an extra API bill. A single-judge round is light. A full 5-judge money-page round is the heavyweight option, and the skill tells you the estimated size before it starts, every time, with a leaner option offered.

Will it work on things that aren't websites?

Yes. Emails, sales pages, funnels, app screens, video scripts, decks, and a general panel for anything else. The judges change to match the work; the gate never changes.

Case file 09 · Get the skill

Own the judge your AI answers to.

  • The complete /perfection skill folder, yours to keep
  • Seven expert panels covering everything you build
  • Four run modes, from a low-cost Speed Scan to relentless
  • The code-enforced gate with its 63-test proof suite included
  • This page as your permanent instruction manual
$997

One time. Runs forever on your own Claude plan.

Get /perfection now

Delivery by email

You're going to keep building with AI either way. The only question is whether anything in the loop has the authority to say not good enough yet.