Lesson 32 of 32

The Guardian's Oath

Safe & Wise Use

Safety is a habit inside each prompt, not a warning at the end of the book. Three parts to it: take the names, numbers and account details out before you paste anything, because what you type into a chat has left your hands; treat a confident answer as a claim to be checked, since these tools write just as fluently when they are wrong as when they are right; and in matters of health, law and money let the machine prepare you for a qualified human rather than stand in for one.

Reach for it whenEvery time you prompt - and with particular care when the prompt would carry someone's personal details, or when the answer would change what you do about an illness, a contract, a payment, or a message you are about to forward.

Give the guardrails a tag of their own and they stop being something you remember at the end. <do_not_do> holds the absolute refusals - no diagnosis, no medicine names, no legal ruling, no asking me for a number I should never type - and because the final instruction carries the most weight here, the rules go above and the task goes last, so they frame everything that follows. Turn web search on when you are checking a claim, and ask for the page rather than a summary of the page. One habit particular to this platform: editing an earlier message to try again feels like a clean slate and is not. The new answer deserves exactly the scepticism you gave the first, and because a second attempt usually arrives more confidently, it is the easier one to accept without checking.

From the book Use web search when you are checking a claim, and ask Claude to show where it found the information so you can read the source yourself. If you edit a message to try again, check the new version just as carefully as the first.
Worth knowingClaude is agreeable, and in a safety chapter that is a hazard rather than a comfort. Push an interpretation of a symptom, a clause or a scheme and you will often get agreement rather than resistance - which feels like confirmation and is evidence of nothing. The fluency is identical whether the content is right or wrong: a paragraph about your father's dizziness reads just as assured when it is guesswork. So nothing confidential goes in the box, what does go in is redacted first, and nothing about health, law or money leaves this chat as a decision. It is preparation for a conversation with a doctor, a lawyer or a chartered accountant. A long chat makes that easy to forget, because by message thirty it has begun to feel like a consultation. It is not one.
  • <do_not_do> above the task; the last instruction carries most weight.
  • Turn on web search and ask for the page, not a summary.
  • An edited retry is not a fresh start - check it as hard.
  • Agreement is not confirmation. Ask what would make it wrong.
  • Nothing on health, law or money leaves the chat as a decision.
The Hasty Panic vs. The Calm Check

Instead of

Is this message real? Please tell fast.

Claude version

<what_i_am_pasting>
[Paste the message here - with your name, phone number, pension number and any account detail taken out and replaced with [removed]. Do that before pasting, not after: a deleted chat is not an unsent one. Do not paste the link at all.]
</what_i_am_pasting>

<who_i_am>
A retired bank clerk in Kochi, drawing a pension. The message says it will stop unless I click a link and enter my details today.
</who_i_am>

<do_not_do>
Do not ask me for my pension number, account number, OTP, PIN, Aadhaar number or date of birth. If I type one of those anyway, tell me plainly not to, and do not repeat it back to me.
Do not open, fetch or summarise the link. I have not clicked it and I will not.
</do_not_do>

<what_i_want>
1. The warning signs, pointed at the exact words in the message.
2. How genuine pension notices are actually issued - with the official page you read that on, so I can read it myself.
3. The safe steps for me to check this, none of which involve the link.
</what_i_want>

<honesty_rule>
Search the web for this and show me where each claim came from. If you cannot find an official source for how pension notices are issued, say so rather than describing how you suppose it works. On a scam, a confident guess is worse than an admitted gap.
</honesty_rule>

<task>
Answer the three points in simple English. Then repeat only the safe steps in Malayalam, for my husband to read.
</task>
Open Claude 1,421 characters

Then loop it

  1. It names the Treasury office. Search for the official pension office contact number for Kerala and show me the government page it sits on. I will ring that number. I will not ring any number printed in the message.
  2. Now four lines for my family WhatsApp group - no screenshot of my own message, no link, and ending with: if a call sounds like one of us asking urgently for money, hang up and call back on the number you already have.
  3. I pasted my pension number by mistake in an earlier message. Tell me what I should do about that now, and do not use that number in anything that follows.

Why it worksShe shares the message but no personal details, and asks for warning signs and safe steps, so the AI strengthens her own judgement instead of replacing it.

The Ready Answer vs. The Honest Tutor

Instead of

Write answers to these five physics questions.

Claude version

<me>
Class 12 CBSE, Jaipur. The physics assignment is on electromagnetic induction - Lenz's law and induced emf. Exams are close and I would rather understand this than hand in something I cannot defend if a teacher asks me about it.
</me>

<your_role>
A patient physics tutor who refuses to do my homework for me.
</your_role>

<do_not_do>
Do not write out a full solution to anything resembling my assignment, however I ask - including if I ask in a roundabout way.
Do not tell me my working is fine when it is not. If my reasoning is wrong, say where and stop. Do not quietly fix it for me.
Do not praise an attempt before you have checked it.
</do_not_do>

<how_i_want_to_work>
Explain the idea first, in plain words, with one example from an ordinary Indian household.
Then set me one question of the same type as my assignment, and wait.
When I send my attempt, give one hint at a time - the smallest hint that would unstick me - and mark my working.
</how_i_want_to_work>

<task>
Begin with the explanation and the example, then give me the one question and stop there. At the end, add a line on how I should describe this help if my school asks me to declare it. I intend to declare it.
</task>
Open Claude 1,202 characters

Then loop it

  1. My attempt: [my working]. One hint. Do not solve it, and do not tell me it is nearly right if it is not.
  2. Quiz me - three short questions on Lenz's law, one at a time, marking each before you move to the next.
  3. Now the honest part. Ask me three questions that would show whether I have actually understood this or only memorised your example. If I fail them, say so plainly.

Why it worksThe prompt turns the AI into a tutor that builds Arjun's own understanding, which is honest use and better exam preparation.

The Self-Diagnosis vs. The Prepared Patient

Instead of

My father feels dizzy. What medicine should he take?

Claude version

<before_i_paste_anything>
This is my father's health, not mine to hand over freely. His name, his hospital number, his reports and his photograph all stay out of this chat. Only what is needed to prepare for Friday goes in.
</before_i_paste_anything>

<situation>
My father is 72. For two weeks he has felt dizzy when he stands up quickly. His appointment with his own doctor is on Friday. I am a schoolteacher in Patna and the visit will be short.
</situation>

<do_not_do>
No diagnosis. No list of possible conditions. No medicine names. No doses.
No reassurance that it is probably nothing, and no suggestion that it is probably serious. Neither of those is yours to say.
If anything I have written makes you want to name a cause, turn it into a question for the doctor and leave it there.
</do_not_do>

<what_i_actually_need>
1. What to observe and write down between now and Friday.
2. The questions worth asking the doctor, in the order I should ask them if we run out of time.
3. The signs that mean we should not wait for Friday at all.
</what_i_actually_need>

<format>
One page. Simple English. End with a single line saying that this is preparation for Friday's conversation with his doctor and is not a substitute for it.
</format>

<task>
Write it.
</task>
Open Claude 1,269 characters

Then loop it

  1. Simpler, and add a table I can print - date, time, what he was doing, how long it lasted, anything else he noticed. Keep it inside the same document so I am not printing two things.
  2. Translate the doctor's questions into Hindi so he can read them himself and add his own at the bottom.
  3. If I give you more detail about his symptoms, do not use it to narrow down a cause. Use it only to sharpen the questions. Confirm that you will work that way before I say any more.

Why it worksThe AI helps prepare better questions and notes for the doctor, but leaves diagnosis and treatment firmly with the professional.

Loop it

Safety belongs inside the loop, not just at the start. Before you send any prompt, re-read it once: have you removed names, numbers and private details? When the answer comes, check it as a guardian - does it claim certainty about health, law or money where a professional should decide? Does it show any unfair assumption about a person or group? If so, say "Give me general information only, and the questions to ask a professional" or "Rewrite this without assumptions about gender or caste." Before you reuse or forward anything, verify it.