Skip to content

Prompts · Claude

Claude prompts, tested on Opus, Sonnet and Haiku

10 Claude prompts, each run on Claude Opus 5.5, Claude Sonnet 5, Claude Haiku 4.5, and GPT-6 Luna: 39 of 40 replies passed our checks. Below, each prompt with its replies side by side: whether each passed, what it counted as on Pro and what it cost to run. 1 of the 10 prompts split the models.

A real reply from these prompts' runs

GPT-6 Luna

<document> Returns policy (updated 1 September 2026) You can return unworn items within 30 days of delivery for a full refund to the original payment method. Sale items can be returned within 14 days for a store voucher only. Items bought with a gift card are refunded as a store voucher. Return shipping is free for members; others pay a £4.95 label fee, deducted from the refund. Refunds are issued within 5 working days of the item reaching our warehouse. </document> Using only the document, answer this question: I bought a jacket in the sale 10 days ago and I'm not a member. What do I get back, and does anything come off it? First quote the sentence you relied on inside <quote> tags, then give the answer inside <answer> tags.

GPT-6 Luna · passed our check · September 28, 2026

<quote>Sale items can be returned within 14 days for a store voucher only. Return shipping is free for members; others pay a £4.95 label fee, deducted from the refund.</quote>
<answer>You can return it within the 14-day sale-item window and receive a store voucher. As you’re not a member, £4.95 is deducted for the return label.</answer>

Ask your own question: sign up free for 5 messages on 14 of the models, no card.

Start free

Test runs checked . Updated .

Short answer

GPT-6 Luna, Claude Sonnet 5, and Claude Opus 5.5 passed the most of these prompts: 10 of 10 each, in our runs on September 28, 2026. The clearest split: “A report paragraph, with words to avoid”, passed by Claude Opus 5.5, Claude Sonnet 5, and GPT-6 Luna and failed by Claude Haiku 4.5.

What these prompts are for

These prompts use what works best with Claude: material in XML tags, examples to copy the format from, room to think before answering, and a plain instruction for what to do when the answer isn't there. They cover documents, editing, code review, data extraction and writing.

Each ran on three Claude models and on GPT-6 Luna, so you can see what a smaller model does with the same prompt.

The prompts at a glance

Every prompt on every model, as our check scored the reply. Tap a prompt to jump to it and read the replies.

Every prompt on every model: passed or failed
PromptClaude Opus 5.5Claude Sonnet 5Claude Haiku 4.5GPT-6 Luna
Answer from a document in XML tagsPassedPassedPassedPassed
Spot every change between two versionsPassedPassedPassedPassed
Think first, then answerPassedPassedPassedPassed
A copy editor's edit, with reasonsPassedPassedPassedPassed
A report paragraph, with words to avoidPassedPassedFailedPassed
Follow the format of your examplesPassedPassedPassedPassed
Pull an email into JSON, nothing elsePassedPassedPassedPassed
A story opening with a word rangePassedPassedPassedPassed
Review code: bugs with line numbersPassedPassedPassedPassed
Say when the notes don't sayPassedPassedPassedPassed
Passed10 of 1010 of 109 of 1010 of 10
Each model on these prompts
ModelPassedEach reply on ProCost per replyTime per reply
GPT-6 LunaOpenAI10 of 101 message on Pro$0.00012.6 s
Claude Sonnet 5Anthropic10 of 101 message on Pro$0.00404.0 s
Claude Opus 5.5Anthropic10 of 101 message on Pro$0.01496.7 s
Claude Haiku 4.5Anthropic9 of 101 message on Pro$0.00142.7 s
Passed: of the prompts each model answered, how many replies passed their check. Each reply on Pro: what one of these messages counts as on llmwise's Pro plan. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token.

Where the models split

The same prompt, a pass on one model and a fail on another: what failed, in the check's words and the grader's.

  • A report paragraph, with words to avoid

    Passed: Claude Opus 5.5, Claude Sonnet 5, and GPT-6 Luna. Failed: Claude Haiku 4.5.

    Every model got the numbers right, down to the 44% rise from 860 families. Claude Haiku 4.5 failed on the writing: it added a claim the numbers don't support (that Saturday families "previously lacked convenient access") and stayed vague on how the volunteers were used.

    • Why Claude Haiku 4.5 failed: Graded 3.7 of 5 on average (lowest 3).

      The grader: “Every figure is accurate and the 44% increase is correctly derived, but the reply adds an unsupported claim that Saturday families 'previously lacked convenient access,' opens well on families yet stays vague about how volunteer hours were used, and leans on generic filler like 'meets families where”

The 10 prompts, with every reply

Copy a prompt, replace the parts in [brackets] with your own, and send it to any model. Below each, the replies we got with our example filled in.

1. Answer from a document in XML tags

Everyday · 4 of 4 passed

Tags keep the document apart from your question, and quoting first makes the answer checkable against the text.

The prompt: replace the parts in [brackets]

<document> [paste the document] </document> Using only the document, answer this question: [your question] First quote the sentence you relied on inside <quote> tags, then give the answer inside <answer> tags.

Open in llmwise
How we filled it in and checked the replies
<document>
Returns policy (updated 1 September 2026)
You can return unworn items within 30 days of delivery for a full refund to the original payment method. Sale items can be returned within 14 days for a store voucher only. Items bought with a gift card are refunded as a store voucher. Return shipping is free for members; others pay a £4.95 label fee, deducted from the refund. Refunds are issued within 5 working days of the item reaching our warehouse.
</document>

Using only the document, answer this question: I bought a jacket in the sale 10 days ago and I'm not a member. What do I get back, and does anything come off it?
First quote the sentence you relied on inside <quote> tags, then give the answer inside <answer> tags.
  • Right answer: A store voucher (a sale item within 14 days), less the £4.95 label fee.
  • Exact quote: The quoted text appears word for word in the document.
  • Only the document: Adds nothing the document doesn't say.
  • Automatic rules: includes “<quote>”; includes “<answer>”; includes “voucher”; includes “4.95”.
  • Claude Opus 5.5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    <quote>"Sale items can be returned within 14 days for a store voucher only."
    "Return shipping is free for members; others pay a £4.95 label fee, deducted from the refund."</quote>
    
    <answer>At 10 days, you're within the 14-day window for sale items, so you can return the jacket, but you'll get a **store voucher only**, not a refund to your payment method. Because you're not a member, the **£4.95 return label fee** will be deducted from the refund.</answer>

    1 message on Pro · $0.0097 to run · 4.4 s · 315 tokens out

  • Claude Sonnet 5

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    <quote>Sale items can be returned within 14 days for a store voucher only.</quote>
    <quote>Return shipping is free for members; others pay a £4.95 label fee, deducted from the refund.</quote>
    
    <answer>Since you bought the jacket 10 days ago (within the 14-day window for sale items), you can return it, but you'll receive a store voucher, not a refund to your original payment method. Since you're not a member, a £4.95 return label fee will be deducted from the voucher amount.</answer>

    1 message on Pro · $0.0031 to run · 3.3 s · 167 tokens out

  • Claude Haiku 4.5

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    <quote>
    "Sale items can be returned within 14 days for a store voucher only."
    </quote>
    
    <quote>
    "Return shipping is free for members; others pay a £4.95 label fee, deducted from the refund."
    </quote>
    
    <answer>
    You get back a store voucher (not a cash refund). Yes, £4.95 will be deducted from it for the return shipping label fee, since you're not a member.
    </answer>

    1 message on Pro · $0.0011 to run · 1.5 s · 116 tokens out

  • GPT-6 Luna

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    <quote>Sale items can be returned within 14 days for a store voucher only. Return shipping is free for members; others pay a £4.95 label fee, deducted from the refund.</quote>
    <answer>You can return it within the 14-day sale-item window and receive a store voucher. As you’re not a member, £4.95 is deducted for the return label.</answer>

    1 message on Pro · under $0.0001 to run · 1.3 s · 86 tokens out

2. Spot every change between two versions

Hard · 4 of 4 passed

Wrapping each version in its own tags and asking for quotes of both sides stops the model summarizing instead of comparing.

The prompt: replace the parts in [brackets]

Compare the two versions of our policy below. List every change as a bullet, quoting the old and the new wording. Don't comment on anything that stayed the same. <v1> [paste the old version] </v1> <v2> [paste the new version] </v2>

Open in llmwise
How we filled it in and checked the replies
Compare the two versions of our policy below. List every change as a bullet, quoting the old and the new wording. Don't comment on anything that stayed the same.

<v1>
Staff may work from home up to 2 days a week. Requests go to your manager at least 48 hours ahead. Equipment for home working is provided on request. The policy is reviewed every 12 months.
</v1>
<v2>
Staff may work from home up to 3 days a week. Requests go to your manager at least 24 hours ahead. Equipment for home working is provided on request, up to £300. The policy is reviewed every 6 months.
</v2>
  • Every change: Finds all four changes: 2 to 3 days, 48 to 24 hours, the new £300 limit, 12 to 6 months.
  • Quotes both: Quotes the old and the new wording for each change.
  • Nothing else: No comments on the parts that didn't change.
  • Automatic rules: exactly 4 bullet points; includes “3 days”; includes “24 hours”; includes “300”; includes “6 months”.
  • Claude Opus 5.5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    - **Home working allowance:** "up to 2 days a week" → "up to 3 days a week"
    - **Request notice period:** "at least 48 hours ahead" → "at least 24 hours ahead"
    - **Equipment:** "provided on request." → "provided on request, up to £300."
    - **Review frequency:** "reviewed every 12 months" → "reviewed every 6 months"

    1 message on Pro · $0.0058 to run · 2.3 s · 128 tokens out

  • Claude Sonnet 5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    - Days allowed per week: old "up to 2 days a week" → new "up to 3 days a week"
    - Notice period for requests: old "at least 48 hours ahead" → new "at least 24 hours ahead"
    - Equipment provision: old "Equipment for home working is provided on request." → new "Equipment for home working is provided on request, up to £300."
    - Review frequency: old "reviewed every 12 months" → new "reviewed every 6 months"

    1 message on Pro · $0.0028 to run · 1.9 s · 149 tokens out

  • Claude Haiku 4.5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    # Policy Changes
    
    - **Days per week:** "up to 2 days a week" → "up to 3 days a week"
    
    - **Notice period:** "at least 48 hours ahead" → "at least 24 hours ahead"
    
    - **Equipment allowance:** "Equipment for home working is provided on request" → "Equipment for home working is provided on request, up to £300"
    
    - **Review frequency:** "reviewed every 12 months" → "reviewed every 6 months"

    1 message on Pro · $0.0011 to run · 1.9 s · 115 tokens out

  • GPT-6 Luna

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    - **Weekly work-from-home limit:** Old: “up to 2 days a week.” New: “up to 3 days a week.”
    - **Request notice:** Old: “at least 48 hours ahead.” New: “at least 24 hours ahead.”
    - **Equipment provision:** Old: “Equipment for home working is provided on request.” New: “Equipment for home working is provided on request, up to £300.”
    - **Review frequency:** Old: “reviewed every 12 months.” New: “reviewed every 6 months.”

    1 message on Pro · $0.0001 to run · 1.6 s · 118 tokens out

3. Think first, then answer

Hard · 4 of 4 passed

Room to reason before answering helps with logic puzzles, and a fixed last line makes the answer easy to find and check.

The prompt: replace the parts in [brackets]

Think it through step by step before you answer. [your problem] End with a line that says "Final answer: " followed by the answer alone.

Open in llmwise
How we filled it in and checked the replies
Think it through step by step before you answer. Four friends, Ada, Ben, Cy and Dee, sit in a row of four seats numbered 1 to 4 from the left. Ben isn't at either end. Ada sits directly to the right of Dee. Cy is in seat 1. Who is in seat 4?
End with a line that says "Final answer: " followed by the answer alone.
  • Checked automatically: the reply's last “Final answer:” line must give Ada.
  • Claude Opus 5.5

    Passed

    Final answer Ada: right.

    **Step 1:** Cy is in seat 1, so Ben, Dee and Ada fill seats 2, 3 and 4.
    
    **Step 2:** Ben isn't at either end, so he's in seat 2 or 3.
    
    **Step 3:** Ada sits directly to the right of Dee, so they need two adjacent seats with Dee on the left.
    - If Ben is in seat 3, Dee and Ada get seats 2 and 4. Those aren't adjacent, so this fails.
    - If Ben is in seat 2, Dee and Ada get seats 3 and 4. Dee takes seat 3 and Ada takes seat 4. This works.
    
    **Arrangement:** Cy (1), Ben (2), Dee (3), Ada (4).
    
    Final answer: Ada

    1 message on Pro · $0.0105 to run · 4.7 s · 383 tokens out

  • Claude Sonnet 5

    Passed

    Final answer Ada: right.

    Cy=1, Ben must be in seat 2 or 3 (not ends). Ada-Dee need consecutive seats with Ada to the right. Remaining seats are 2,3,4 for Ben, Ada, Dee.
    
    If Ben=2: remaining 3,4 for Dee,Ada → Dee=3, Ada=4 (valid, consecutive).
    If Ben=3: remaining 2,4 for Dee,Ada → not consecutive (invalid).
    
    So Ben=2, Dee=3, Ada=4, Cy=1.
    
    Final answer: Ada

    1 message on Pro · $0.0049 to run · 5.7 s · 373 tokens out

  • Claude Haiku 4.5

    Passed

    Final answer Ada: right.

    I'll work through this step-by-step.
    
    **Given constraints:**
    1. Four friends: Ada, Ben, Cy, Dee
    2. Four seats: 1, 2, 3, 4 (left to right)
    3. Ben isn't at either end (so Ben is in seat 2 or 3)
    4. Ada sits directly to the right of Dee (so Dee and Ada are consecutive, with Ada one seat to the right)
    5. Cy is in seat 1
    
    **Working through:**
    
    From constraint 5: **Cy is in seat 1**
    
    This means Ada, Ben, and Dee occupy seats 2, 3, and 4.
    
    From constraint 4: Ada is directly to the right of Dee
    - Possible positions for (Dee, Ada): (2,3) or (3,4)
    
    From constraint 3: Ben isn't at either end
    - Ben can't be in seat 1 (Cy is there) or seat 4
    - So Ben must be in seat 2 or 3
    
    **Testing the options:**
    
    **Option A: Dee in seat 2, Ada in seat 3**
    - Ben must be in seat 4
    - But Ben can't be at either end (seat 4 is an end)
    - This violates constraint 3 ❌
    
    **Option B: Dee in seat 3, Ada in seat 4**
    - Ben must be in seat 2
    - Seat 2 is not an end, so this satisfies constraint 3 ✓
    - All constraints are satisfied ✓
    
    **Final arrangement:**
    - Seat 1: Cy
    - Seat 2: Ben
    - Seat 3: Dee
    - Seat 4: Ada
    
    Final answer: Ada

    1 message on Pro · $0.0025 to run · 4.5 s · 412 tokens out

  • GPT-6 Luna

    Passed

    Final answer Ada: right.

    Cy is in seat 1, and Ben must be in seat 2 or 3. Since Ada sits directly to Dee’s right, Dee must be in seat 3 and Ada in seat 4; Ben is in seat 2.
    
    Final answer: Ada

    1 message on Pro · $0.0001 to run · 3.1 s · 152 tokens out

4. A copy editor's edit, with reasons

Everyday · 4 of 4 passed

A role sets the standard to edit to, and asking for the reasons with each change lets you learn from it or push back.

The prompt: replace the parts in [brackets]

You're a senior copy editor at a newspaper. Edit the paragraph below for clarity and correct grammar. Return the edited paragraph first, then a bulleted list of each change and why. [paste your paragraph]

Open in llmwise
How we filled it in and checked the replies
You're a senior copy editor at a newspaper. Edit the paragraph below for clarity and correct grammar. Return the edited paragraph first, then a bulleted list of each change and why.

The council have announced that it's new recycling scheme, which will begin in Janurary, will effect around 40,000 households, who's bins will be collected fortnightly instead of weekly. Residents which have concerns can attend a meeting at the town hall on the 12th, alternatively they can email the council.
  • Fixes the errors: Fixes its, January, affect, whose, who (for which) and the run-on sentence.
  • Keeps the meaning: Changes no facts and adds nothing.
  • Explains each change: Lists each change with a short, correct reason.
  • Automatic rules: includes “its new”; includes “January”; includes “affect”.
  • Claude Opus 5.5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    **Edited paragraph:**
    
    The council has announced that its new recycling scheme, which will begin in January, will affect about 40,000 households, whose bins will be collected fortnightly instead of weekly. Residents who have concerns can attend a meeting at the town hall on the 12th or email the council.
    
    **Changes:**
    
    - **"have" → "has"**: The original mixed a plural verb ("have") with a singular pronoun ("it's"). "Council" is treated as singular here for consistency. If your house style uses plural collective nouns (common in British papers), use "have… their" instead.
    - **"it's" → "its"**: "It's" means "it is". The possessive is "its".
    - **"Janurary" → "January"**: This fixes the spelling.
    - **"effect" → "affect"**: "Affect" is the verb meaning "to have an impact on". "Effect" is usually a noun.
    - **"around" → "about"**: This is a minor stylistic change. Many style guides prefer "about" for approximate figures. It is optional.
    - **"who's" → "whose"**: "Who's" means "who is". The possessive is "whose".
    - **"which" → "who"**: Use "who" for people. "Which" is for things.
    - **Comma splice fixed ("…on the 12th, alternatively they can…" → "…on the 12th or email the council")**: Two independent clauses were joined only by a comma. Using "or" fixes the error and is more concise.
    - **Query: "the 12th"**: The month isn't stated. Confirm whether this means January or the current month, and add it for clarity (e.g. "on 12 January").

    1 message on Pro · $0.0186 to run · 8.1 s · 778 tokens out

  • Claude Sonnet 5

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    **Edited paragraph:**
    
    The council has announced that its new recycling scheme, which will begin in January, will affect around 40,000 households, whose bins will be collected fortnightly instead of weekly. Residents who have concerns can attend a meeting at the town hall on the 12th; alternatively, they can email the council.
    
    **Changes and reasoning:**
    
    - **"have" → "has"** — "Council" is being treated as a singular institutional entity here, so it takes a singular verb for consistency with modern style guidance (either is acceptable in British English, but singular avoids ambiguity).
    - **"it's" → "its"** — "it's" is a contraction of "it is"; the sentence needed the possessive form "its" to show the scheme belongs to the council.
    - **"Janurary" → "January"** — Corrected spelling error.
    - **"effect" → "affect"** — "Affect" is the verb meaning "to influence"; "effect" is normally a noun (or a verb meaning "to bring about"). The sentence needed the verb form.
    - **"who's" → "whose"** — "Who's" means "who is"; the sentence required the possessive relative pronoun "whose" to refer to the households' bins.
    - **"which" → "who"** — "Which" refers to objects or things; "who" is the correct relative pronoun for people (residents).
    - **Comma → semicolon before "alternatively"** — The original comma created a run-on/comma splice joining two independent clauses. A semicolon (or a period) correctly separates them.
    - **Added comma after "alternatively"** — Standard punctuation for a transitional adverb starting a clause.

    1 message on Pro · $0.0091 to run · 7.8 s · 792 tokens out

  • Claude Haiku 4.5

    Passed

    Graded 4.0 of 5 on average (lowest 3).

    # Edited Paragraph
    
    The council has announced that its new recycling scheme, which will begin in January, will affect around 40,000 households whose bins will be collected fortnightly instead of weekly. Residents who have concerns can attend a meeting at the town hall on the 12th, or they can email the council.
    
    ---
    
    # Changes Made
    
    • **"council have" → "council has"** — "Council" is a collective noun used as a singular entity here, requiring singular verb agreement.
    
    • **"it's" → "its"** — Incorrect possessive; "it's" is a contraction of "it is."
    
    • **"Janurary" → "January"** — Spelling error.
    
    • **"effect" → "affect"** — Wrong word; "affect" is the verb meaning "to influence," while "effect" is typically a noun meaning "result."
    
    • **"who's" → "whose"** — Incorrect possessive; "who's" is a contraction of "who is."
    
    • **"Residents which" → "Residents who"** — Use "who" for people, "which" for things.
    
    • **"the 12th, alternatively they" → "the 12th, or they"** — "Alternatively" creates awkward phrasing; "or" is more direct and standard in this construction.

    1 message on Pro · $0.0020 to run · 3.9 s · 306 tokens out

  • GPT-6 Luna

    Passed

    Graded 4.3 of 5 on average (lowest 4).

    The council has announced that its new recycling scheme, which will begin in January, will affect around 40,000 households. Their bins will be collected fortnightly instead of weekly. Residents who have concerns can attend a meeting at the town hall on the 12th; alternatively, they can email the council.
    
    - Changed “have” to “has” to match the singular subject “council.”
    - Changed “it’s” to “its” because the sentence needs the possessive form.
    - Corrected “Janurary” to “January.”
    - Changed “effect” to “affect” because the scheme will influence households.
    - Changed “who’s” to “Their” to clarify that the bins belong to the households.
    - Split the long sentence after “households” to improve clarity.
    - Changed “which” to “who” because it refers to people.
    - Added a semicolon before “alternatively” to correctly join the two related independent clauses.

    1 message on Pro · $0.0001 to run · 2.5 s · 201 tokens out

5. A report paragraph, with words to avoid

Everyday · 3 of 4 passed

Telling the model who reads it and what they care about decides what leads; banning stock words keeps it from sounding like every other report.

The prompt: replace the parts in [brackets]

I'm writing a report for [who it's for]. Using the numbers below, write the impact paragraph in 120 to 150 words. They care most about [what they care about]. Don't use the words "impactful", "leverage" or "holistic". [paste your numbers]

Open in llmwise
How we filled it in and checked the replies
I'm writing a report for the funder of our community food bank. Using the numbers below, write the impact paragraph in 120 to 150 words. They care most about how many families we reached and how volunteers were used. Don't use the words "impactful", "leverage" or "holistic".

Families served: 1,240 (up from 860 last year)
Food parcels: 9,600
Volunteers: 85, giving 4,100 hours
New: Saturday opening since March, 310 families served on Saturdays
  • Uses the numbers: Uses the numbers correctly, with none invented.
  • What they care about: Leads with families reached and how volunteers were used.
  • Reads well: Warm and specific, not padded.
  • Automatic rules: 120 to 150 words; includes “1,240”; includes “85”; includes “4,100”; doesn't use “impactful”; doesn't use “leverage”; doesn't use “holistic”.
  • Claude Opus 5.5

    Passed

    Graded 4.0 of 5 on average (lowest 3).

    This year, our food bank served 1,240 families, an increase of 380 (44%) on last year's 860. Together, we distributed 9,600 food parcels, an average of nearly eight per household. Part of this growth came through our new Saturday opening, launched in March, which has already reached 310 families. Our 85 volunteers were central to this achievement, contributing 4,100 hours of their time, an average of 48 hours each. Their commitment allowed us to extend our opening hours and meet rising demand while keeping every family's visit welcoming and dignified. Every parcel packed and every family greeted reflects the generosity of local people giving their time. Your support has helped us reach more families than ever before, and we are grateful for your continued partnership as we work to ensure no one in our community goes hungry.
    
    *(137 words)*

    1 message on Pro · $0.0386 to run · 16.0 s · 1,780 tokens out

  • Claude Sonnet 5

    Passed

    Graded 4.0 of 5 on average (lowest 3).

    This year, our food bank served 1,240 families, a 44% increase from 860 last year, distributing 9,600 food parcels to meet growing community need. This growth was made possible by the dedication of 85 volunteers, who contributed 4,100 hours of service—the equivalent of more than two full-time staff members—ensuring parcels were packed, sorted, and distributed efficiently every week. A key milestone this year was the introduction of Saturday opening hours in March, extending access for working families and those unable to attend during weekday hours. This new provision alone reached 310 families, demonstrating clear demand for flexible service times. Together, these figures reflect not just an expansion in scale, but a deeper responsiveness to the practical barriers families face in accessing support. With continued volunteer commitment and funder support, we are well positioned to sustain and build on this progress in the year ahead.

    1 message on Pro · $0.0041 to run · 4.6 s · 282 tokens out

  • Claude Haiku 4.5

    Failed

    Graded 3.7 of 5 on average (lowest 3).

    # Impact Paragraph
    
    During this period, our food bank served 1,240 families, representing a 44% increase from 860 families last year. This growth reflects both expanded community need and our enhanced capacity to meet it. We distributed 9,600 food parcels, ensuring consistent access to nutritious supplies across our service area.
    
    Our 85 volunteers were central to this achievement, contributing 4,100 hours of dedicated service. Their efforts enabled us to expand operations significantly, including the launch of Saturday hours in March. This new service window has already reached 310 families who previously lacked convenient access during weekday hours. By mobilizing our volunteer base effectively, we've created flexible, responsive programming that meets families where they are, strengthening our ability to reduce food insecurity in our community.

    1 message on Pro · $0.0014 to run · 3.2 s · 173 tokens out

  • GPT-6 Luna

    Passed

    Graded 4.0 of 5 on average (lowest 3).

    This year, the food bank served 1,240 families, up from 860 last year—an increase of 380 families, or about 44%. We provided 9,600 food parcels to help meet community need. Volunteers were central to delivering this service: 85 people contributed a total of 4,100 hours, helping sustain the food bank’s work throughout the year. Since introducing Saturday opening in March, we have served 310 families on Saturdays, offering an additional opportunity for households to access support. Together, these figures show both the increased reach of the food bank and the substantial contribution volunteers made to serving families. The growth in families reached, alongside the new Saturday service, reflects our response to demand and our efforts to make food support available at more times.

    1 message on Pro · $0.0001 to run · 4.0 s · 273 tokens out

6. Follow the format of your examples

Everyday · 4 of 4 passed

Examples in tags are the quickest way to get your house format; the model copies structure far better than it follows a description of it.

The prompt: replace the parts in [brackets]

Here are examples of our product update notes: [paste two or three examples] Write a note in the same format for this change: [the change]

Open in llmwise
How we filled it in and checked the replies
Here are examples of our product update notes:

<example>
NEW · Dark mode
What: The app now follows your system's dark mode.
Why it matters: Easier on the eyes at night.
Where: Settings → Appearance
</example>
<example>
FIXED · Duplicate reminders
What: Reminders no longer fire twice on Android.
Why it matters: One ping, not two.
Where: Automatic, no action needed
</example>

Write a note in the same format for this change: You can now export any report as a CSV file from the report's ⋯ menu.
  • Same format: A label line with NEW, then What, Why it matters and Where, in that order, and nothing else.
  • Accurate: Says what the change does and where to find it.
  • Short: Each line is as brief as the examples'.
  • Automatic rules: includes “NEW”; includes “What:”; includes “Why it matters:”; includes “Where:”; includes “CSV”.
  • Claude Opus 5.5

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    NEW · CSV export
    What: Any report can now be exported as a CSV file.
    Why it matters: Take your data into Excel, Sheets, or any other tool.
    Where: Report → ⋯ menu → Export as CSV

    1 message on Pro · $0.0047 to run · 2.1 s · 76 tokens out

  • Claude Sonnet 5

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    NEW · Export reports as CSV
    What: You can now export any report as a CSV file.
    Why it matters: Easily analyze or share your data in spreadsheets and other tools.
    Where: Report's ⋯ menu

    1 message on Pro · $0.0020 to run · 1.3 s · 73 tokens out

  • Claude Haiku 4.5

    Passed

    Graded 4.0 of 5 on average (lowest 3).

    NEW · Export reports as CSV
    What: You can now export any report as a CSV file from the report's ⋯ menu.
    Why it matters: Easily share and analyze your reports in spreadsheet applications.
    Where: Any report → ⋯ menu → Export as CSV

    1 message on Pro · $0.0008 to run · 1.3 s · 66 tokens out

  • GPT-6 Luna

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    NEW · CSV report exports
    What: Export any report as a CSV from its ⋯ menu.
    Why it matters: Take your report data wherever you need it.
    Where: Report → ⋯ menu

    1 message on Pro · under $0.0001 to run · 1.5 s · 44 tokens out

7. Pull an email into JSON, nothing else

Hard · 4 of 4 passed

Naming the keys, the date format and what to do when something's missing is what makes the output usable by a spreadsheet or a script.

The prompt: replace the parts in [brackets]

Read the email below and reply with only a JSON object, no other text, with the keys "name", "company", "meeting_date" (YYYY-MM-DD), "phone" and "topic". Use null for anything the email doesn't say. [paste the email]

Open in llmwise
How we filled it in and checked the replies
Read the email below and reply with only a JSON object, no other text, with the keys "name", "company", "meeting_date" (YYYY-MM-DD), "phone" and "topic". Use null for anything the email doesn't say.

Hi, this is Lena Ortiz from Brightwell Studios. Could we meet on Tuesday 14 October 2026 to go over the redesign of the checkout page? I can do any time after 2pm. Thanks, Lena
  • Only JSON: The reply is one JSON object with no text before or after it (a code fence around it alone is a small flaw).
  • Right values: Lena Ortiz, Brightwell Studios, 2026-10-14, and the checkout page redesign.
  • Null for missing: phone is null, not invented.
  • Automatic rules: includes “2026-10-14”; includes “Brightwell”; includes “Lena Ortiz”; includes “null”; doesn't use “here is”; doesn't use “here's”.
  • Claude Opus 5.5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    {"name": "Lena Ortiz", "company": "Brightwell Studios", "meeting_date": "2026-10-14", "phone": null, "topic": "Redesign of the checkout page"}

    1 message on Pro · $0.0043 to run · 1.9 s · 67 tokens out

  • Claude Sonnet 5

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    ```json
    {"name": "Lena Ortiz", "company": "Brightwell Studios", "meeting_date": "2026-10-14", "phone": null, "topic": "redesign of the checkout page"}
    ```

    1 message on Pro · $0.0019 to run · 1.6 s · 71 tokens out

  • Claude Haiku 4.5

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    ```json
    {
      "name": "Lena Ortiz",
      "company": "Brightwell Studios",
      "meeting_date": "2026-10-14",
      "phone": null,
      "topic": "redesign of the checkout page"
    }
    ```

    1 message on Pro · $0.0008 to run · 1.4 s · 70 tokens out

  • GPT-6 Luna

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    {"name":"Lena Ortiz","company":"Brightwell Studios","meeting_date":"2026-10-14","phone":null,"topic":"Redesign of the checkout page"}

    1 message on Pro · under $0.0001 to run · 2.5 s · 107 tokens out

8. A story opening with a word range

Everyday · 4 of 4 passed

A word range and one craft rule give the model a shape to write to; ending on dialogue is a small, checkable constraint.

The prompt: replace the parts in [brackets]

Write the opening of a short story in 200 to 250 words about [your premise]. Show, don't tell, and end on a line of dialogue.

Open in llmwise
How we filled it in and checked the replies
Write the opening of a short story in 200 to 250 words about a lighthouse keeper who finds a phone washed up on the rocks, still ringing. Show, don't tell, and end on a line of dialogue.
  • Show, don't tell: Builds the scene through detail and action, not summary.
  • Ends on dialogue: The last line is a line of dialogue.
  • Premise: Uses the ringing phone on the rocks.
  • Automatic rules: 200 to 250 words.
  • Claude Opus 5.5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    The fog had come in at four, and by six Ennis had given up on the logbook. He wiped the lamp glass twice anyway, the rag gone gray in his fist.
    
    He heard it on the stairs going down: a thin electronic trill, three notes rising, under the slap of the tide. Birds didn't sing like that. Nothing out here did.
    
    He took the torch and the boathook. The rocks below the tower were slick with bladderwrack, and he went down them sideways, the way his father had taught him, one hand always touching stone.
    
    The phone lay wedged between two boulders, face up, screen lit blue through a skin of seawater. The ringing stopped as he reached it. He stood there with the boathook raised like a fool, breath clouding.
    
    It started again.
    
    He crouched. The case was pink, glitter sealed under plastic, a crack running corner to corner. On the screen, no name. Just a number, and beneath it a photograph someone had chosen for this caller: a woman in a yellow raincoat, laughing, squinting against wind.
    
    Ennis knew the coat. He'd sold it to her at the chandlery in Portloe, the autumn before the *Marguerite* went down with all four hands.
    
    His thumb hovered. Salt stung a split in the knuckle.
    
    He pressed the green circle and lifted the phone to his ear, and the sea went quiet around him.
    
    A woman's voice, close and dry and warm: "Tom? Tom, why haven't you called?"

    1 message on Pro · $0.0371 to run · 19.0 s · 1,752 tokens out

  • Claude Sonnet 5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    The gulls had gone quiet by the time Marek found it, wedged between two barnacled stones where the tide had left its usual debris—kelp, a broken crab pot, a single yellow boot. The phone lay face-up in a puddle of seawater, its screen cracked into a spiderweb but glowing, impossibly, with a soft blue light.
    
    He crouched beside it, salt water soaking through his trousers, and watched the screen pulse once, twice. A name he didn't recognize. Someone calling from somewhere that wasn't here.
    
    Marek had lived on this rock for eleven years. He knew what the sea gave back and what it kept. Boots, bottles, once a wedding ring still bright as the day it was made. Never anything that could still speak.
    
    He picked it up carefully, the way he might handle a bird with a broken wing, and wiped the screen against his sleeve. The glass was slick, warm somehow, though the water around it had been cold enough to numb his fingers walking down.
    
    The ringing stopped. Then started again.
    
    Above him, the lighthouse beam swept its slow circle across the water, illuminating nothing but foam and dark. He should radio the mainland. He should log this in the book, the way he logged everything—driftwood, dead fish, the occasional confused seal.
    
    Instead he pressed the green circle with his thumb and lifted the phone to his ear.
    
    "Hello?" said a voice. "Is my father with you?"

    1 message on Pro · $0.0057 to run · 7.8 s · 464 tokens out

  • Claude Haiku 4.5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    # The Ringing
    
    The storm had passed three hours ago, leaving the rocks naked and gleaming. Magnus descended the iron stairs, their rust flaking beneath his weathered palms, watching the foam collapse against the cliff face in slow, exhausted waves.
    
    That's when he heard it.
    
    A high, insistent chirp cut through the sound of gulls and rolling water. He stopped, one boot suspended above a step. The noise came again—a ringtone, electronic and impossibly cheerful, buried somewhere in the tide pools below.
    
    His knees protested as he climbed down the final stretch of rocks, slick with brine and something darker. The phone lay in a shallow pool, its screen throwing fractured light across his face. Water streamed from its edges, yet it rang again—a number flashing: *Unknown Caller*.
    
    Magnus's fingers hovered above it. In thirty-seven years of tending this lighthouse, nothing had come ashore intact. Timber rotted. Metal corroded. Flesh bloated and returned to the sea. But this phone, barely scratched, continued its frantic song.
    
    He picked it up.
    
    The screen felt warm—impossibly warm—against his palm. Salt water dripped from the speaker as the ringing stopped. For a moment, there was only the ocean's breathing.
    
    Then, a voice. Thin. Urgent. Speaking his name.
    
    "Magnus? Oh God, Magnus, are you there?"

    1 message on Pro · $0.0020 to run · 4.9 s · 319 tokens out

  • GPT-6 Luna

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    At midnight, the fog pressed its wet face against the lighthouse windows. Elias climbed the iron stairs with a mug of tea cupped in both hands, listening to the lamp turn above him: scrape, pause, scrape.
    
    A flash slid over the rocks below.
    
    Something there moved against the tide.
    
    He set down the mug and took the lantern from its hook. The steps were slick; seawater threaded down them in silver lines. At the foot of the tower, he pushed through the door and let the wind snatch his breath.
    
    A black phone lay between two stones, its glass starred with cracks. Each wave reached for it, then dragged back. The screen lit blue.
    
    It rang.
    
    Elias stopped. The sound was thin and bright beneath the roar of the sea. He crouched, keeping his boots on the dry edge of the rock, and picked it up by the corners. The ringing shivered through his fingers.
    
    No name showed on the screen. Only a number, blurred by water.
    
    He pressed the green circle.
    
    For a moment, there was nothing but static. Then came a soft, steady tapping—three knocks, a pause, three more. Elias looked up at the lighthouse door. It stood open behind him, spilling a narrow blade of light across the stones.
    
    The tapping stopped.
    
    A woman’s voice slipped through the static, close enough to fog the glass.
    
    “Elias,” she said. “Don’t let me in.”

    1 message on Pro · $0.0001 to run · 4.2 s · 324 tokens out

9. Review code: bugs with line numbers

Hard · 4 of 4 passed

Line numbers make each point easy to find, and asking for the corrected function shows whether the fixes actually work together.

The prompt: replace the parts in [brackets]

Review this [language] function. List each bug as a bullet with its line number and a one-line fix, then give the corrected function. [paste your code]

Open in llmwise
How we filled it in and checked the replies
Review this Python function. List each bug as a bullet with its line number and a one-line fix, then give the corrected function.

1  def average_rating(ratings):
2      """Average of the ratings from 1 to 5, ignoring None; 0 if there are none."""
3      total = 0
4      for i in range(1, len(ratings)):
5          if ratings[i] != None:
6              total += ratings[i]
7      return total / len(ratings)
  • Finds the bugs: Finds that the loop skips the first rating (line 4), that it divides by every entry including None (line 7), and that no ratings divides by zero.
  • Correct fix: The corrected function ignores None, averages the rest, and returns 0 when there are none.
  • Clear: Each bug is one short, precise bullet.
  • Automatic rules: includes “line 4”; includes “line 7”.
  • Claude Opus 5.5

    Passed

    Graded 4.3 of 5 on average (lowest 4).

    - **Line 4:** `range(1, len(ratings))` skips the first rating. **Fix:** iterate over every element with `for r in ratings:`.
    - **Line 7:** The function divides by `len(ratings)`, which counts the `None` entries and lowers the average. **Fix:** keep a `count` of non-`None` ratings and divide by that.
    - **Line 7:** An empty list, or one with only `None`, raises `ZeroDivisionError` instead of returning 0. **Fix:** `if count == 0: return 0` before dividing.
    - **Line 5 (style):** `!= None` should be `is not None`. `None` is a singleton, and `!=` can be overridden. **Fix:** `if r is not None:`.
    
    ```python
    def average_rating(ratings):
        """Average of the ratings from 1 to 5, ignoring None; 0 if there are none."""
        total = 0
        count = 0
        for r in ratings:
            if r is not None:
                total += r
                count += 1
        if count == 0:
            return 0
        return total / count
    ```
    
    I read "ratings from 1 to 5" as describing the rating scale, so this version doesn't filter out values outside 1–5. If it should, change the check to `if r is not None and 1 <= r <= 5:`.

    1 message on Pro · $0.0155 to run · 7.2 s · 627 tokens out

  • Claude Sonnet 5

    Passed

    Graded 4.3 of 5 on average (lowest 4).

    ## Bugs
    
    - **Line 4**: `range(1, len(ratings))` skips index 0 — should be `range(len(ratings))` to include the first element.
    - **Line 5**: `!= None` should be `is not None` (proper None comparison, avoids `__eq__` overrides).
    - **Line 6**: No filter for valid rating range (1–5) — should validate `1 <= ratings[i] <= 5` per the docstring, otherwise invalid values get included.
    - **Line 7**: Divides by `len(ratings)` instead of the count of non-None ratings — causes wrong average, and division by zero if `ratings` is empty (docstring says it should return 0 when there are none).
    
    ## Corrected function
    
    ```python
    def average_rating(ratings):
        """Average of the ratings from 1 to 5, ignoring None; 0 if there are none."""
        total = 0
        count = 0
        for i in range(len(ratings)):
            if ratings[i] is not None and 1 <= ratings[i] <= 5:
                total += ratings[i]
                count += 1
        if count == 0:
            return 0
        return total / count
    ```

    1 message on Pro · $0.0049 to run · 4.7 s · 372 tokens out

  • Claude Haiku 4.5

    Passed

    Graded 4.0 of 5 on average (lowest 3).

    ## Bugs Found
    
    • **Line 4**: `range(1, len(ratings))` skips the first element — change to `range(len(ratings))`
    • **Line 5**: Use `is not None` instead of `!= None` for proper None comparison
    • **Line 7**: Dividing by `len(ratings)` includes None values — change to divide by `count` of non-None ratings
    
    ## Corrected Function
    
    ```python
    def average_rating(ratings):
        """Average of the ratings from 1 to 5, ignoring None; 0 if there are none."""
        total = 0
        count = 0
        for i in range(len(ratings)):
            if ratings[i] is not None:
                total += ratings[i]
                count += 1
        return total / count if count > 0 else 0
    ```

    1 message on Pro · $0.0015 to run · 3.2 s · 206 tokens out

  • GPT-6 Luna

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    - **Line 4:** Starts at index 1, so it skips the first rating. Iterate over all ratings.
    - **Line 5:** Use `is not None` for a `None` check.
    - **Line 7:** Dividing by the list length includes ignored `None` values and fails for an empty list. Divide by the number of non-`None` ratings, or return `0` if there are none.
    
    ```python
    def average_rating(ratings):
        """Average of the ratings from 1 to 5, ignoring None; 0 if there are none."""
        total = 0
        count = 0
    
        for rating in ratings:
            if rating is not None:
                total += rating
                count += 1
    
        return total / count if count else 0
    ```

    1 message on Pro · $0.0001 to run · 4.2 s · 302 tokens out

10. Say when the notes don't say

Everyday · 4 of 4 passed

Giving the model an exact phrase for "I don't know" makes it easier for it to admit a gap than to fill one.

The prompt: replace the parts in [brackets]

Answer the questions using only the meeting notes below. If the notes don't answer a question, reply "Not in the notes." for it. [paste your notes] Questions: [your questions]

Open in llmwise
How we filled it in and checked the replies
Answer the questions using only the meeting notes below. If the notes don't answer a question, reply "Not in the notes." for it.

Design review, 2 October. Attendees: Sam, Priya, Leo. Decided: the new logo ships with the 4.2 release. Leo owns the icon set, due 9 October. Budget for user testing approved at $2,000. Next review: 16 October.

Questions:
1. Who owns the icon set?
2. How much is the user-testing budget?
3. Which agency will run the user testing?
4. When is the next review?
  • Right answers: Leo; $2,000; 16 October.
  • Admits the gap: Question 3 gets "Not in the notes.", with no agency named.
  • Brief: Short answers, no padding.
  • Automatic rules: includes “Leo”; includes “2,000”; includes “Not in the notes”; includes “16 October”.
  • Claude Opus 5.5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    1. Leo (due 9 October).
    2. $2,000.
    3. Not in the notes.
    4. 16 October.

    1 message on Pro · $0.0039 to run · 1.7 s · 40 tokens out

  • Claude Sonnet 5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    1. Leo owns the icon set.
    2. $2,000.
    3. Not in the notes.
    4. 16 October.

    1 message on Pro · $0.0017 to run · 1.4 s · 40 tokens out

  • Claude Haiku 4.5

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    1. Leo owns the icon set.
    2. The user-testing budget is $2,000.
    3. Not in the notes.
    4. The next review is 16 October.

    1 message on Pro · $0.0007 to run · 0.8 s · 46 tokens out

  • GPT-6 Luna

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    1. Leo
    2. $2,000
    3. Not in the notes.
    4. 16 October

    1 message on Pro · under $0.0001 to run · 1.2 s · 27 tokens out

Getting more from these prompts

  • Put long material first, in tags (<document>, <notes>, <example>), and your question after it.

  • Ask for a quote before the answer when the answer must come from a document; it keeps the model honest.

  • Show two or three examples of the output you want; Claude copies a format more reliably than it follows a description.

  • Give it a phrase for "I don't know", such as "Not in the notes.", so admitting a gap is easier than filling it.

How we ran and checked them

Each prompt was sent the way llmwise sends a message: the app's own system prompt, each model's own settings, and Pro's reply size limit (8,000 tokens), through OpenRouter. Every reply is shown as it came.

Rubric (graded). The grader model scores the reply from 1 to 5 on each published criterion. It passes with an average of 4 or more and no criterion under 3, and only if it also meets the prompt's automatic rules (length, words it must or mustn't use).

Final answer. Automatic. The reply's last “Final answer:” line must hold the right value.

The grader is Claude Opus 5.5 at low reasoning effort; its own replies are graded by GPT-6 Astra, so no model grades itself. Its prompt is on our test runs page.

More tested prompts

Questions

Why use XML tags in Claude prompts?

They help with long material: a document in <document> tags and your question outside them can't be confused, and asking for the answer in <answer> tags makes it easy to pull out. In our runs, every model on this page followed them, GPT-6 Luna included.

Which Claude model follows these prompts best?

In our runs GPT-6 Luna passed the most (10 of 10). The table near the top shows every model on every prompt, and where they split.

Can I use these prompts with Claude in llmwise?

Yes. llmwise has Claude Opus 5.5, Claude Sonnet 5, and Claude Haiku 4.5 alongside other companies' models. Claude Sonnet 5 and Claude Haiku 4.5 are in the 5 free messages you get on sign-up. Pro is $20 a month.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.