Lesson 28 of 32

The Mirror of Self-Critique

Reflection Loops

Build the checking into the prompt itself. Give a short checklist of what the finished work must achieve - four to six things a careful reader could answer yes or no - then tell the model to draft, score itself against each point while quoting the line that falls short, and rewrite. Two rounds, not ten. It works because "make it better" has nothing to measure against, while five checkable criteria do.

Reach for it whenWhen you can say precisely what good looks like - a cover letter, a plan, a story for a particular class - and you would rather receive a draft that has already been through an editor than one you must edit yourself.

Tag the parts and the loop runs cleanly: <rubric> for the standard, <creator> and <critic> for the two characters, and the instruction to score, quote and rewrite placed last. Claude takes to this chapter more readily than the others because it will criticise its own text without being coaxed - ask what is weakest in a draft and you generally get a real answer rather than a compliment. Put the final version in an Artifact while the scores and the critic's notes stay in the chat, so the clean text is separate from the argument about it. A rubric you use often belongs in a Project's instructions, where every draft in that project is measured the same way.

From the book Put your favourite rubrics in a Project's instructions so every draft in that project is scored the same way. Ask Claude to show the final version in an Artifact while the scores and critique stay in the chat, so the clean text is easy to copy or download.
Worth knowingWillingness to criticise is not the ability to verify, and the difference matters. Claude will cheerfully tell you a sentence is weak, vague or repetitive, and it cannot tell you that the figure inside the sentence is wrong - a rubric checks your draft against your criteria, never against the world. On top of that, it inflates scores once it has praised a draft, and it is agreeable: tell it the story is lovely and the critic softens for the rest of the conversation. Keep your praise until the loop is finished, insist on one real weakness per round, and do your own factual read at the end.
  • <rubric>, <creator>, <critic> as tags; scoring instruction last.
  • Ask what is weakest - Claude answers that honestly.
  • Hold back your praise: it softens the critic for the whole chat.
  • Final version in an Artifact, scores and notes in the chat.
  • A rubric judges wording, not truth. Check facts yourself.
The Vague Polish vs. The Scored Rubric

Instead of

Improve my cover letter for this job.

Claude version

<job_advertisement>
[Paste the advertisement here, in full. Long material first, my instruction at the end.]
</job_advertisement>

<me>
B.Com, 2025. Tally and Excel. Six months interning at a CA firm in Bhubaneswar, where my actual work was [write the two or three things you really did]. Applying for junior accounts assistant at a logistics firm.
</me>

<rubric>
Answer each of these with yes or no before you put a mark out of 5 against it:
- Is the letter under 200 words?
- Does its first sentence say which post I am applying for?
- Does it pick up at least three phrases from the advertisement above?
- Does it contain one thing I actually did, as against where I happened to be?
- Is it free of the sentences every fresher writes?
</rubric>

<scoring_discipline>
Score each criterion out of 5. Quote the line responsible for any score below 5. Every round must name at least one genuine weakness, the final round included - a clean sheet of fives tells me the critic has gone to sleep, not that the letter is finished.

And be exact about the boundary of what you are checking: the rubric tests whether the letter is well made. It cannot test whether the letter is true. If you have written anything about my internship that I did not give you above, flag it under a heading <invented> and ask me, rather than scoring it as satisfied.
</scoring_discipline>

<task>
Write the letter, score it, rewrite it, score it again. Two rounds. Then put the final letter in an Artifact and leave the scores and your notes here in the chat.
</task>
Open Claude 1,543 characters

Then loop it

  1. Your second-round scores are all fives. Which criterion is the final letter weakest on, and why? Quote the line. Then fix that one thing and nothing else.
  2. Your <invented> section was right to flag the ledger line - I did that twice, not every month. Correct it. Then go through the whole letter and mark every clause that makes a claim about my experience you could not have drawn from what I told you.
  3. Settled. Now give me the rubric and the scoring discipline on their own, so I can paste them into a Project called "Job applications" and have every letter measured the same way.

Why it worksThe rubric turns "make it better" into five checkable standards, and the scoring loop forces the AI to mend its own shortfalls before she sees the letter.

The Cheerful Plan vs. The Plan That Argued With Itself

Instead of

Should I grow vegetables instead of wheat?

Claude version

<my_situation>
Ten acres of wheat and paddy near Moga. I am considering putting 2 acres into tomato and capsicum next season to raise my income. I have no experience of vegetables, and no cold storage of my own.
</my_situation>

<sequence>
First a plan: the season's steps, where I would sell, what I would have to learn. Brief.

Then turn on it. Give me the five strongest reasons this fails for a farmer in my position - price collapse, harvest labour, water, storage, transport, whatever you judge strongest. Order them by how likely they are to actually sink me. No "however, with good planning" anywhere in this section; the whole value of it is that it does not reassure me.

Then revise the plan so it answers each objection, and state openly which objections cannot be answered and must simply be carried as risk.

Then questions for the Krishi Vigyan Kendra.
</sequence>

<what_you_cannot_know>
Your objections will be right in kind. Your numbers will not be current - this season's rates, my mandi's behaviour, this year's input costs are all outside what you have. So mark every figure you produce with (unverified) and gather them at the end under <take_to_kvk>. Do not let an unverified number sit inside the revised plan as though it were established, because that is precisely the error self-critique is blind to: you will score the plan sound and the arithmetic underneath it may be two years stale.
</what_you_cannot_know>

<task>
Work through the sequence above.
</task>
Open Claude 1,488 characters

Then loop it

  1. The price-collapse objection is the one I cannot sleep on. <critic>You are a tough mandi trader in Moga with no interest in encouraging me.</critic> Attack the revised plan on selling alone - timing, grading, who sets the rate, what happens in a glut week.
  2. Now do the thing the critique cannot do for itself: list everything under <take_to_kvk> in one place, with the question I should ask about each, since those are the numbers the whole plan rests on.
  3. Write the final plan as one page of simple Punjabi I can set before my family. Keep the accepted risks visible in it - do not let the revision quietly drop them to make the page encouraging.

Why it worksMaking the AI argue against itself exposes risks a cheerful first answer hides, and the revised plan must survive those objections before he trusts it.

The Lone Writer vs. The Critic and the Creator

Instead of

Write a story for kids about saving water.

Claude version

<roles>
<creator>A writer of stories for small children, with energy and some mischief.</creator>
<critic>A primary-reading specialist who has taught Class 3 for twenty years and has no interest in sparing the writer's feelings.</critic>
</roles>

<brief>
A 250-word story for Class 3 children in Kerala, about a girl who saves water at home. Simple English, a little humour, a clear ending. I teach this class and will read it aloud.
</brief>

<rubric>
The Critic answers each of these before scoring anything:
- Which words here would stop an 8-year-old reader?
- Which sentences run past twelve words?
- Is the lesson made once, or does it keep returning?
- Does the child in the story work the point out, or is the point announced to the reader?
- Where is the moment the class will repeat to each other at lunch? If there is none, say so.
</rubric>

<how_the_two_work>
Each fault must come with the exact words attached. A verdict with nothing quoted beneath it is no use to someone who has to read this out tomorrow.

If you reach a pass where you genuinely cannot fault anything, do not manufacture a fault to satisfy me. Say instead which of the five questions you are least confident you applied properly, and why - that is the more useful admission, and it is the one a self-marking loop almost never makes.

I am withholding my own praise until the whole thing is finished. Approval softens a critic, and I would rather keep the critic.
</how_the_two_work>

<task>
Two passes. Final story in an Artifact; the notes stay down here; the discarded drafts go nowhere.
</task>
Open Claude 1,582 characters

Then loop it

  1. The Critic let the preaching rule slide. The closing paragraph still lectures. Critic: take the ending alone, quote the sentences that preach. Creator: rewrite the ending and leave the rest untouched.
  2. Now the part neither role can settle: mark every word in the final story that an 8-year-old in Kerala may not have met in English, and say where you are guessing rather than knowing. My class is the real test and I will read it to them.
  3. Good. Add 5 simple comprehension questions for the class, and a Malayalam version of the story in a second Artifact - translated faithfully rather than improved along the way.

Why it worksSplitting creator and critic gives the self-check a strict, named judge and a rubric, so each rewrite answers real problems rather than vague praise.

Loop it

Read the scores before the story. If every criterion scores a perfect 5, the loop is flattering itself - reply, "Be stricter: quote the weakest line for each point and rewrite." If the final version improved one point but broke another, add that point to the rubric and run one more round. Stop after two or three rounds; more rarely helps. Then do your own human read, checking facts the rubric cannot catch. When a rubric works, save it and paste it into future prompts as a standing measure of quality.