Lesson 30 of 32

The Senses of the Machine

Files, Images & Voice

These assistants can now look at a photograph, read a document, work through a sheet of figures and listen to your voice - so stop labouring to describe what you could simply show. But a picture without a purpose is a ship without a port: name what the thing is, supply the context the camera could not capture, point to the exact page or column that matters, and say what you want out of it.

Reach for it whenWhen the problem is in front of you rather than in your head - a spotted leaf, a bill you cannot read, a forty-page circular handed over at five o'clock, a column of figures - or when typing is itself the obstacle.

Photographs, PDFs, Word files and spreadsheets all go in through the plus button, and the habit worth building is to attach first and then write the prompt around the attachment - what it is, where it came from, what you want back. Two things are particular to this platform. Given a file of figures it can write and run code over the data rather than reading it by eye, so ask for the calculation and ask to see the working. And Voice mode is a true spoken conversation: you talk, it answers aloud, you interrupt it, you switch language mid-sentence or say “slower” and it adjusts - which is this whole chapter for a reader who does not type comfortably.

From the book Upload images and files with the plus button and ask about them as you would a person. For speaking instead of typing, tap the Voice mode button and talk naturally in your own language; you can say "speak more slowly" or "answer in Hindi" at any time.
Worth knowingThe eye is confident and occasionally wrong. A blurred photo, a handwritten note, a faded bill or a smudged figure tends to be read as something plausible rather than reported as unreadable, so ask it to name what it cannot make out, and check any figure it quotes back at you. Voice has its own gaps: spoken English and Hindi come back well, and other Indian languages vary with the accent and with how much noise is around you, so repeat a number or a date back and have it confirm before you act on it. And crop before you upload. Covering a consumer number with your thumb as you take the photograph is a perfectly good method.
  • Attach first, then write what it is and what you want out.
  • For a sheet of figures, ask for the calculation and the working.
  • Ask it to name what it cannot read instead of guessing.
  • In Voice mode say “slower” or “in Hindi” - it adjusts mid-answer.
  • Crop or cover names, numbers and addresses before uploading.
The Bare Photo vs. The Guided Eye

Instead of

What is wrong with this leaf?

ChatGPT version

[Attach the photo of the leaf. If you can, add a second of its underside and a third taken standing back from the whole plant.]

These are from my tomato field near Kolar, Karnataka. The monsoon has just begun. About 20 plants out of 300 look like this, mostly the older lower leaves, and it has spread over roughly a week.

1. Tell me only what you can actually see in the photographs - the colour, the pattern, the edges, where on the leaf the spots sit. No diagnosis yet.
2. Then the two or three most likely causes, and for each one the single thing I should look for in the field to rule it in or out.
3. Then what I can do this week without buying anything.
4. Then the questions to put to my Krishi Vigyan Kendra or the agriculture officer.

If a photo is too blurred or too dark to judge, say so and tell me which photograph to take instead - closer, from behind, in different light. I would rather take another picture than act on a guess.
Open ChatGPT 948 characters

Then loop it

  1. Here is the underside of the leaf and one of the stem near the soil. Does either change your answer, and if so, which of your causes moves up the list?
  2. Good. Now the questions for the agriculture officer in Kannada, short enough to read out over a phone call, with the English underneath so I can check them.

Why it worksThe photo shows the problem, but the words supply what the camera cannot - place, season, spread - and ask the AI to admit when it cannot see enough.

The Whole Document vs. The Pointed Page

Instead of

Can you summarise this PDF for me?

ChatGPT version

[Attach the 40-page circular.]

Using only this attached PDF, and nothing from outside it, help me work the new pension procedure.

Read Section 4 and the annexures at the back first. Look at the rest of the circular only to check whether anything there modifies Section 4 - and if it does, tell me where.

Then give me, in this order:
1. The steps an employee must follow, numbered in the order they actually happen.
2. Every document required, set against the step it is needed at.
3. Every deadline or time limit named anywhere in the circular.

Put the page number in brackets after each point. Where the circular is unclear, self-contradictory or silent, write “Not stated in the circular” instead of filling the gap with what such circulars usually say.

Before any of that, tell me how many pages you were able to read, and whether any page came through as a scan you could not make out.
Open ChatGPT 894 characters

Then loop it

  1. On page 18 the circular mentions a form. Quote that paragraph exactly as printed, and tell me which annexure it points to.
  2. Your step 5 and the annexure on page 33 seem to disagree about who signs. Quote both and tell me which one the circular gives effect to.
  3. Now a one-page checklist table I can print and keep on my desk - step, document, deadline, page - and nothing else on the sheet.

Why it works"Using only the attached PDF" anchors the AI to the document, and page numbers let Fatima verify every answer herself.

The Hesitant Typist vs. The Clear Speaker

Instead of

(spoken) Bill… why so much?

ChatGPT version

(Spoken, in Voice mode, after photographing the bill with her thumb over the consumer number and the address.)

“I am showing you my electricity bill. Please speak slowly, in simple Tamil, as you would to your own grandmother.

First tell me two things only: how much I must pay, and the last date. Say the amount in rupees and the date in words. Then stop and wait for me.”

(After it answers, keep speaking.)

“Now look at the units used this month and the units last month, and tell me in two or three short sentences why the bill has gone up. Then one or two simple ways to use less current at home - nothing that means buying a new machine.”

“If any number on the bill is not clear in the photograph, tell me which one and I will take the picture again.”
Open ChatGPT 760 characters

Then loop it

  1. (Spoken) Say the last date to pay once more, slowly, and then tell me which day of the week that is.
  2. (Spoken) Repeat the amount back to me and say the rupees digit by digit, so I know you read the figure and not the one beside it.
  3. (Spoken) Is there any late fee or extra charge on this bill? Answer yes or no first, then explain in one sentence.

Why it worksVoice removes the barrier of typing and reading, but the prompt still states the aim, the language and the format - and personal details are hidden first.

Loop it

When you share a photo or file, the first answer tells you what the AI actually saw. Check that first: did it read the right page, the right column, the right figure? If it misread something, correct it - "The amount in row 7 is ₹1,250, not ₹125" - and ask it to redo the answer. If a photo was blurred, send a clearer one or a close-up. Then narrow the focus, ask for page numbers or quotes to verify, and turn the result into something you can reuse: a checklist, a table or a spoken summary.