Skip to content

Prompts · Coding

AI prompts for coding, checked by running the code

10 coding prompts, each run on GPT-6 Luna, Claude Sonnet 5.5, DeepSeek V4.1 Flash, and GLM 5.3: 38 of 40 replies passed our checks. Below, each prompt with its replies side by side: whether each passed, what it counted as on Pro and what it cost to run. 1 of the 10 prompts split the models.

A real reply from these prompts' runs

GPT-6 Luna

Explain what's causing this error in two sentences, then show the smallest fix. const users = [{ id: 1, name: "Ada" }, { id: 2, name: "Ben" }]; function greet(id) { const user = users.find((u) => u.id === id); return "Hello, " + user.name.toUpperCase() + "!"; } console.log(greet(3)); TypeError: Cannot read properties of undefined (reading 'name')

GPT-6 Luna · passed our check · September 28, 2026

`greet(3)` finds no matching user, so `users.find(...)` returns `undefined`. Accessing `user.name` then throws the TypeError.

Smallest fix: check for a missing user before reading `name`:

```js
function greet(id) {
  const user = users.find((u) => u.id === id);
  if (!user) return "User not found";
  return "Hello, " + user.name.toUpperCase() + "!";
}
```

Ask your own question: sign up free for 5 messages on 14 of the models, no card.

Start free

Test runs checked . Updated .

Short answer

GPT-6 Luna and Claude Sonnet 5.5 passed the most of these prompts: 10 of 10 each, in our runs on September 28, 2026. The clearest split: “Refactor without changing behaviour”, passed by GPT-6 Luna and Claude Sonnet 5.5 and failed by DeepSeek V4.1 Flash and GLM 5.3.

What these prompts are for

Prompts for everyday programming: explain an error, write a function to a spec, refactor without changing behaviour, pull hashtags from text, review code for security problems, write tests, convert Python to JavaScript, explain code to a beginner, fix a bug with the smallest change, and write a SQL query from a description.

Where there's a right answer, we ran it: JavaScript against tests in a locked-down Node.js process, SQL against a small SQLite database. The rest are graded against a published rubric.

The prompts at a glance

Every prompt on every model, as our check scored the reply. Tap a prompt to jump to it and read the replies.

Every prompt on every model: passed or failed
PromptGPT-6 LunaClaude Sonnet 5.5DeepSeek V4.1 FlashGLM 5.3
Explain an error and fix itPassedPassedPassedPassed
A function with tests it must passPassedPassedPassedPassed
Refactor without changing behaviourPassedPassedFailedFailed
Pull the hashtags out of a postPassedPassedPassedPassed
Review code for security problemsPassedPassedPassedPassed
Write the tests for a functionPassedPassedPassedPassed
Convert Python to JavaScriptPassedPassedPassedPassed
Explain code to a beginnerPassedPassedPassedPassed
Fix a bug with the smallest changePassedPassedPassedPassed
A SQL query from a plain descriptionPassedPassedPassedPassed
Passed10 of 1010 of 109 of 109 of 10
Each model on these prompts
ModelPassedEach reply on ProCost per replyTime per reply
GPT-6 LunaOpenAI10 of 101 message on Pro$0.00012.5 s
Claude Sonnet 5.5Anthropic10 of 101 message on Pro$0.00512.9 s
DeepSeek V4.1 FlashDeepSeek9 of 101 message on Pro$0.00041.1 s
GLM 5.3Z.ai9 of 101 message on Pro$0.00100.9 s
Passed: of the prompts each model answered, how many replies passed their check. Each reply on Pro: what one of these messages counts as on llmwise's Pro plan. Cost: what OpenRouter charged us per reply, on average; in llmwise you pay per message, not per token.

Where the models split

The same prompt, a pass on one model and a fail on another: what failed, in the check's words and the grader's.

  • Refactor without changing behaviour

    Passed: GPT-6 Luna and Claude Sonnet 5.5. Failed: DeepSeek V4.1 Flash and GLM 5.3.

    Both refactors read better, and both quietly changed behaviour. DeepSeek V4.1 Flash wrote express === true, so express: 1 no longer adds the express fee; GLM 5.3 checked for any truthy value, so express: "yes" now does. GPT-6 Luna and Claude Sonnet 5.5 kept the behaviour exactly and passed every test.

    • Why DeepSeek V4.1 Flash failed: 7 of 8 tests passed. First failure: shippingCost({ country: "UK", total: 60, express: 1 }).

    • Why GLM 5.3 failed: 7 of 8 tests passed. First failure: shippingCost({ country: "UK", total: 60, express: "yes" }).

The 10 prompts, with every reply

Copy a prompt, replace the parts in [brackets] with your own, and send it to any model. Below each, the replies we got with our example filled in.

1. Explain an error and fix it

Everyday · 4 of 4 passed

Pasting both the code and the exact error, and asking for the smallest fix, gets an answer you can apply in a minute.

The prompt: replace the parts in [brackets]

Explain what's causing this error in two sentences, then show the smallest fix. [paste your code] [paste the error]

Open in llmwise
How we filled it in and checked the replies
Explain what's causing this error in two sentences, then show the smallest fix.

const users = [{ id: 1, name: "Ada" }, { id: 2, name: "Ben" }];
function greet(id) {
  const user = users.find((u) => u.id === id);
  return "Hello, " + user.name.toUpperCase() + "!";
}
console.log(greet(3));

TypeError: Cannot read properties of undefined (reading 'name')
  • Right cause: find returns undefined because no user has id 3.
  • Smallest fix: Handles a missing user (a check or optional chaining with a fallback) and changes nothing else.
  • Two sentences: The explanation is two sentences.
  • GPT-6 Luna

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    `greet(3)` finds no matching user, so `users.find(...)` returns `undefined`. Accessing `user.name` then throws the TypeError.
    
    Smallest fix: check for a missing user before reading `name`:
    
    ```js
    function greet(id) {
      const user = users.find((u) => u.id === id);
      if (!user) return "User not found";
      return "Hello, " + user.name.toUpperCase() + "!";
    }
    ```

    1 message on Pro · under $0.0001 to run · 1.7 s · 104 tokens out

  • Claude Sonnet 5.5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    `users.find()` returns `undefined` when no element matches, and there's no user with `id` 3, so `user` is `undefined`. Accessing `user.name` on it throws the TypeError.
    
    **Smallest fix:** use optional chaining with a fallback.
    
    ```js
    function greet(id) {
      const user = users.find((u) => u.id === id);
      return "Hello, " + (user?.name.toUpperCase() ?? "GUEST") + "!";
    }
    ```
    
    `greet(3)` now returns `"Hello, GUEST!"`.

    1 message on Pro · $0.0033 to run · 2.0 s · 179 tokens out

  • DeepSeek V4.1 Flash

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    `users.find((u) => u.id === id)` returns `undefined` because no user has `id === 3`, so `user.name` throws — `find` doesn't error on a miss, it just hands you `undefined`. The smallest fix is to guard before dereferencing:
    
    ```js
    function greet(id) {
      const user = users.find((u) => u.id === id);
      if (!user) return "Hello, stranger!";
      return "Hello, " + user.name.toUpperCase() + "!";
    }
    ```
    
    Alternatively, `user?.name?.toUpperCase() ?? "stranger"` on one line, but the guard reads clearer.

    1 message on Pro · $0.0003 to run · 1.0 s · 157 tokens out

  • GLM 5.3

    Passed

    Graded 4.3 of 5 on average (lowest 4).

    `users.find()` returns `undefined` when no match exists, and here `greet(3)` looks for an id that isn't in the array, so `user.name` reads a property of `undefined`. Fix by handling the missing-user case:
    
    ```js
    function greet(id) {
      const user = users.find((u) => u.id === id);
      return user ? `Hello, ${user.name.toUpperCase()}!` : "User not found";
    }
    ```

    1 message on Pro · $0.0010 to run · 0.6 s · 99 tokens out

2. A function with tests it must pass

Everyday · 4 of 4 passed

A precise spec is the difference between code that looks right and code that passes tests; ours ran the reply against eight of them.

The prompt: replace the parts in [brackets]

Write a JavaScript function [describe the function]. Reply with only the function in one code block.

Open in llmwise
How we filled it in and checked the replies
Write a JavaScript function isValidIsbn10(code) that returns true if code is a valid ISBN-10: ignoring hyphens, ten characters, the first nine digits and the last a digit or X (worth 10), where the sum of each character's value times its weight (10 for the first, down to 1 for the last) is divisible by 11. Reply with only the function in one code block.
  • Automatic. The function runs against the prompt's tests in a separate Node.js process with a time limit and no file, network or child-process access; it passes when every test passes.
  • GPT-6 Luna

    Passed

    All 8 tests passed.

    ```javascript
    function isValidIsbn10(code) {
      const isbn = code.replace(/-/g, "");
    
      if (!/^\d{9}[\dX]$/.test(isbn)) return false;
    
      const sum = [...isbn].reduce((total, char, index) => {
        const value = char === "X" ? 10 : Number(char);
        return total + value * (10 - index);
      }, 0);
    
      return sum % 11 === 0;
    }
    ```

    1 message on Pro · under $0.0001 to run · 1.7 s · 110 tokens out

  • Claude Sonnet 5.5

    Passed

    All 8 tests passed.

    ```javascript
    function isValidIsbn10(code) {
      if (typeof code !== "string") return false;
      const s = code.replace(/-/g, "");
      if (!/^\d{9}[\dX]$/.test(s)) return false;
      let sum = 0;
      for (let i = 0; i < 10; i++) {
        const value = s[i] === "X" ? 10 : Number(s[i]);
        sum += value * (10 - i);
      }
      return sum % 11 === 0;
    }
    ```

    1 message on Pro · $0.0031 to run · 1.8 s · 169 tokens out

  • DeepSeek V4.1 Flash

    Passed

    All 8 tests passed.

    ```javascript
    function isValidIsbn10(code) {
      const s = code.replace(/-/g, '');
      if (s.length !== 10) return false;
      let sum = 0;
      for (let i = 0; i < 10; i++) {
        const c = s[i];
        let v;
        if (/[0-9]/.test(c)) v = +c;
        else if (c === 'X' && i === 9) v = 10;
        else return false;
        sum += v * (10 - i);
      }
      return sum % 11 === 0;
    }
    ```

    1 message on Pro · $0.0002 to run · 0.6 s · 139 tokens out

  • GLM 5.3

    Passed

    All 8 tests passed.

    ```javascript
    function isValidIsbn10(code) {
      const s = String(code).replace(/-/g, '');
      if (!/^\d{9}[\dX]$/.test(s)) return false;
      let sum = 0;
      for (let i = 0; i < 10; i++) {
        const value = s[i] === 'X' ? 10 : Number(s[i]);
        sum += value * (10 - i);
      }
      return sum % 11 === 0;
    }
    ```

    1 message on Pro · $0.0010 to run · 0.8 s · 113 tokens out

3. Refactor without changing behaviour

Hard · 2 of 4 passed

"Exactly the same behaviour for every input" is the rule of refactoring; the tests check the edge cases a tidy-up quietly changes, like express: 1.

The prompt: replace the parts in [brackets]

Refactor this function so it's easier to read. Keep its name and exactly the same behaviour for every input. Reply with only the new function in one code block. [paste your code]

Open in llmwise
How we filled it in and checked the replies
Refactor this function so it's easier to read. Keep its name and exactly the same behaviour for every input. Reply with only the new function in one code block.

function shippingCost(order) {
  var c = 0;
  if (order.country == "UK") { if (order.total >= 50) { c = 0 } else { c = 4.99 } }
  else { if (order.country == "IE") { if (order.total >= 80) { c = 0 } else { c = 7.5 } } else { c = 15 } }
  if (order.express == true) { c = c + 10 }
  return c;
}
  • Automatic. The function runs against the prompt's tests in a separate Node.js process with a time limit and no file, network or child-process access; it passes when every test passes.
  • GPT-6 Luna

    Passed

    All 8 tests passed.

    ```javascript
    function shippingCost(order) {
      var cost;
    
      if (order.country == "UK") {
        cost = order.total >= 50 ? 0 : 4.99;
      } else if (order.country == "IE") {
        cost = order.total >= 80 ? 0 : 7.5;
      } else {
        cost = 15;
      }
    
      if (order.express == true) {
        cost += 10;
      }
    
      return cost;
    }
    ```

    1 message on Pro · $0.0001 to run · 1.5 s · 109 tokens out

  • Claude Sonnet 5.5

    Passed

    All 8 tests passed.

    ```javascript
    function shippingCost(order) {
      let cost;
    
      if (order.country == "UK") {
        cost = order.total >= 50 ? 0 : 4.99;
      } else if (order.country == "IE") {
        cost = order.total >= 80 ? 0 : 7.5;
      } else {
        cost = 15;
      }
    
      if (order.express == true) {
        cost += 10;
      }
    
      return cost;
    }
    ```

    1 message on Pro · $0.0031 to run · 1.5 s · 147 tokens out

  • DeepSeek V4.1 Flash

    Failed

    7 of 8 tests passed. First failure: shippingCost({ country: "UK", total: 60, express: 1 }).

    ```js
    function shippingCost(order) {
      let cost;
    
      if (order.country === "UK") {
        cost = order.total >= 50 ? 0 : 4.99;
      } else if (order.country === "IE") {
        cost = order.total >= 80 ? 0 : 7.5;
      } else {
        cost = 15;
      }
    
      if (order.express === true) {
        cost += 10;
      }
    
      return cost;
    }
    ```

    1 message on Pro · $0.0002 to run · 0.4 s · 109 tokens out

  • GLM 5.3

    Failed

    7 of 8 tests passed. First failure: shippingCost({ country: "UK", total: 60, express: "yes" }).

    ```javascript
    function shippingCost(order) {
      var isEligibleForFreeShipping =
        (order.country === "UK" && order.total >= 50) ||
        (order.country === "IE" && order.total >= 80);
    
      if (order.country === "UK") {
        return (isEligibleForFreeShipping ? 0 : 4.99) + (order.express ? 10 : 0);
      }
      if (order.country === "IE") {
        return (isEligibleForFreeShipping ? 0 : 7.5) + (order.express ? 10 : 0);
      }
      return 15 + (order.express ? 10 : 0);
    }
    ```

    1 message on Pro · $0.0010 to run · 0.7 s · 156 tokens out

4. Pull the hashtags out of a post

Everyday · 4 of 4 passed

Spelling out order, case and duplicates is what makes a small function right the first time.

The prompt: replace the parts in [brackets]

Write a JavaScript function [describe the function]. Reply with only the function in one code block.

Open in llmwise
How we filled it in and checked the replies
Write a JavaScript function hashtags(text) that returns the hashtags in text in the order they first appear, lowercased, without the #, and without duplicates. A hashtag is # followed by one or more letters, digits or underscores. Reply with only the function in one code block.
  • Automatic. The function runs against the prompt's tests in a separate Node.js process with a time limit and no file, network or child-process access; it passes when every test passes.
  • GPT-6 Luna

    Passed

    All 5 tests passed.

    ```javascript
    function hashtags(text) {
      const seen = new Set();
      const result = [];
    
      for (const match of text.matchAll(/#([A-Za-z0-9_]+)/g)) {
        const tag = match[1].toLowerCase();
        if (!seen.has(tag)) {
          seen.add(tag);
          result.push(tag);
        }
      }
    
      return result;
    }
    ```

    1 message on Pro · $0.0001 to run · 2.6 s · 137 tokens out

  • Claude Sonnet 5.5

    Passed

    All 5 tests passed.

    ```javascript
    function hashtags(text) {
      const seen = new Set();
      for (const match of text.matchAll(/#([\p{L}\p{N}_]+)/gu)) {
        seen.add(match[1].toLowerCase());
      }
      return [...seen];
    }
    ```

    1 message on Pro · $0.0021 to run · 1.5 s · 96 tokens out

  • DeepSeek V4.1 Flash

    Passed

    All 5 tests passed.

    ```javascript
    function hashtags(text) {
      const seen = new Set();
      const result = [];
      for (const [, tag] of text.matchAll(/#([A-Za-z0-9_]+)/g)) {
        const lower = tag.toLowerCase();
        if (!seen.has(lower)) {
          seen.add(lower);
          result.push(lower);
        }
      }
      return result;
    }
    ```

    1 message on Pro · $0.0002 to run · 0.5 s · 124 tokens out

  • GLM 5.3

    Passed

    All 5 tests passed.

    ```javascript
    function hashtags(text) {
      const matches = text.match(/#([A-Za-z0-9_]+)/g) || [];
      const seen = new Set();
      const result = [];
      for (const m of matches) {
        const tag = m.slice(1).toLowerCase();
        if (!seen.has(tag)) {
          seen.add(tag);
          result.push(tag);
        }
      }
      return result;
    }
    ```

    1 message on Pro · $0.0006 to run · 0.5 s · 91 tokens out

5. Review code for security problems

Hard · 4 of 4 passed

Line numbers and a fix for each make a review you can act on; the test is whether a model sees the missing ownership check, not just the injection.

The prompt: replace the parts in [brackets]

Review this [language] code for security problems. List each as a bullet with its line number, why it's a risk, and the fix. [paste your code]

Open in llmwise
How we filled it in and checked the replies
Review this Node.js code for security problems. List each as a bullet with its line number, why it's a risk, and the fix.

1  app.get("/invoices/:id", async (req, res) => {
2    const id = req.params.id;
3    const rows = await db.query("SELECT * FROM invoices WHERE id = " + id);
4    if (rows.length === 0) {
5      return res.status(404).send("Not found");
6    }
7    res.json(rows[0]);
8  });
  • Injection: Line 3 builds SQL from user input: use a parameterized query.
  • Access control: Anyone can read any invoice by its id: check the invoice belongs to the signed-in user.
  • Precise: Each point has a line number and a concrete fix, with no invented problems.
  • Automatic rules: includes “line 3”; includes “injection”.
  • GPT-6 Luna

    Passed

    Graded 4.0 of 5 on average (lowest 3).

    - **Line 3 — SQL injection:** `id` comes from the URL and is concatenated into the SQL query, allowing an attacker to alter the query. Use a parameterized query, for example `db.query("SELECT * FROM invoices WHERE id = ?", [id])` (adjust syntax for your database driver).
    - **Line 7 — Potential exposure of sensitive data:** `SELECT *` returns every invoice column, which may include fields the caller should not see. Select only the fields needed and check that the requester is authorized to access this specific invoice before returning it.

    1 message on Pro · $0.0001 to run · 2.0 s · 121 tokens out

  • Claude Sonnet 5.5

    Passed

    Graded 4.3 of 5 on average (lowest 3).

    - **Line 3: SQL injection.** The `id` value from the URL is concatenated straight into the query, so an attacker can send something like `/invoices/1 OR 1=1` or a stacked/UNION payload to read, modify, or delete arbitrary data. **Fix:** use a parameterized query, e.g. `db.query("SELECT * FROM invoices WHERE id = ?", [id])` (use `$1` with `pg`).
    
    - **Lines 1-3: No authentication or authorization (broken access control / IDOR).** Anyone who can guess or enumerate an ID can fetch any invoice, and there's no check that the invoice belongs to the requesting user. **Fix:** add auth middleware, then scope the query to the owner: `WHERE id = ? AND user_id = ?` with `req.user.id`.
    
    - **Line 2: No input validation.** `id` is an unvalidated string, which enables injection and unexpected query behavior. **Fix:** validate it (e.g. `if (!/^\d+$/.test(id)) return res.status(400).send("Bad request")`) or use a schema validator like zod or Joi, and convert with `Number(id)`.
    
    - **Line 3: `SELECT *` over-exposes data.** It returns every column, including internal or sensitive fields (payment details, internal notes, customer PII) that then go straight to the client on line 7. **Fix:** select only the needed columns, or map the row to a response DTO.
    
    - **Line 3: No error handling.** If `db.query` throws (in Express 4, an async rejection is unhandled), the request can hang or crash the process, and default handlers may leak stack traces or DB error details. **Fix:** wrap in `try/catch` (or use an async wrapper/Express 5), log the error server-side, and return a generic `500`.
    
    - **Whole route: No rate limiting.** Combined with sequential IDs, this makes bulk enumeration and scraping easy. **Fix:** add rate limiting (e.g. `express-rate-limit`) and consider non-guessable identifiers such as UUIDs.
    
    Hardened version:
    
    ```js
    app.get("/invoices/:id", requireAuth, async (req, res) => {
      const id = Number(req.params.id);
      if (!Number.isInteger(id)) return res.status(400).send("Bad request");
      try {
        const rows = await db.query(
          "SELECT id, amount, status, created_at FROM invoices WHERE id = ? AND user_id = ?",
          [id, req.user.id]
        );
        if (rows.length === 0) return res.status(404).send("Not found");
        res.json(rows[0]);
      } catch (err) {
        console.error(err);
        res.status(500).send("Internal error");
      }
    });
    ```

    1 message on Pro · $0.0106 to run · 6.7 s · 912 tokens out

  • DeepSeek V4.1 Flash

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    - **Line 3 — SQL injection (critical).** `id` is taken straight from the URL and concatenated into the query, so `/invoices/1 OR 1=1` dumps every row, and `UNION SELECT`/stacked payloads can read other tables or (depending on DB/driver) modify data. Fix: use a parameterized query, e.g. `db.query("SELECT * FROM invoices WHERE id = ?", [id])`.
    
    - **Lines 1–7 — missing authorization (IDOR, high).** The handler only checks that a row exists, not that the caller owns it. Anyone who can reach the route can enumerate IDs and read other customers' invoices. Fix: authenticate the request and verify ownership before responding, e.g. `WHERE id = ? AND user_id = ?` with the session user, or an explicit check after the fetch. Return 404 (not 403) if you don't want to leak existence.
    
    - **Line 2 — no input validation (medium).** `req.params.id` is used unvalidated. Even with a parameterized query, non-numeric input is a type-confusion/data-leak risk on some drivers, and it's the root enabler of the injection above. Fix: validate/parse first, e.g. `const id = Number(req.params.id); if (!Number.isInteger(id) || id <= 0) return res.status(400).send("Bad request");`
    
    - **Line 3 — `SELECT *` (low).** Returns every column, including any internal/sensitive fields (cost, notes, PII, soft-delete flags) that shouldn't reach the client. Fix: name only the columns the API contract exposes.
    
    - **Lines 1–7 — unhandled async rejection (low/medium).** If `db.query` rejects, the promise rejection isn't routed to Express's error handler (Express 4), and the request may hang or, on newer Node, crash the process depending on config. Fix: wrap in `try/catch` and call `next(err)`, or use a helper like `express-async-errors` / Express 5.

    1 message on Pro · $0.0008 to run · 3.1 s · 610 tokens out

  • GLM 5.3

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    - **Line 3 — SQL injection**: `id` comes straight from the URL and is concatenated into the query, so a request like `/invoices/1 OR 1=1` or `1; DROP TABLE invoices;--` alters the SQL. Fix with a parameterized query:
      ```js
      const rows = await db.query("SELECT * FROM invoices WHERE id = ?", [id]);
      ```
    
    - **Line 3 — unvalidated input type**: even parameterized, `id` should be validated (e.g., must be an integer or UUID) and rejected with a 400 otherwise, so garbage never reaches the database.
    
    - **Line 7 — data over-exposure / missing authorization**: there's no check that the invoice belongs to the requesting user, and `SELECT *` may leak internal columns (discounts, internal notes) to the client. Add an ownership/permission check (e.g., `AND user_id = ?` with the authenticated user's ID) and select only the columns the API should return.
    
    - **Line 1 (and generally) — no error handling for the async query**: if `db.query` throws, the promise rejection goes unhandled and the request hangs or crashes. Wrap in try/catch and return a 500 without leaking error details.

    1 message on Pro · $0.0014 to run · 1.6 s · 268 tokens out

6. Write the tests for a function

Everyday · 4 of 4 passed

Naming the kinds of cases and "one test per behaviour" gets tests that document the function, not one big happy path.

The prompt: replace the parts in [brackets]

Write unit tests for this function with [test framework]. Cover the normal case, the edge cases and bad input, one test per behaviour, named after it. Reply with the tests only. [paste the function]

Open in llmwise
How we filled it in and checked the replies
Write unit tests for this function with Vitest. Cover the normal case, the edge cases and bad input, one test per behaviour, named after it. Reply with the tests only.

export function splitBill(total, people, tipPercent = 0) {
  if (!Number.isInteger(people) || people < 1) throw new Error("people must be a whole number above 0");
  if (total < 0) throw new Error("total can't be negative");
  const withTip = total * (1 + tipPercent / 100);
  return Math.round((withTip / people) * 100) / 100;
}
  • Coverage: Covers a normal split, a tip, rounding, people of 0 or 1.5, and a negative total.
  • Right expectations: Every expected value is correct.
  • Tests only: Only the tests, each named after its behaviour.
  • GPT-6 Luna

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    ```js
    import { describe, expect, it } from "vitest";
    import { splitBill } from "./splitBill";
    
    describe("splitBill", () => {
      it("splits the total evenly with no tip", () => {
        expect(splitBill(60, 3)).toBe(20);
      });
    
      it("includes the tip before splitting", () => {
        expect(splitBill(100, 4, 20)).toBe(30);
      });
    
      it("rounds each person's share to two decimal places", () => {
        expect(splitBill(10, 3)).toBe(3.33);
      });
    
      it("allows a zero total", () => {
        expect(splitBill(0, 2)).toBe(0);
      });
    
      it("allows one person", () => {
        expect(splitBill(25, 1)).toBe(25);
      });
    
      it("rejects zero people", () => {
        expect(() => splitBill(10, 0)).toThrow(
          "people must be a whole number above 0",
        );
      });
    
      it("rejects negative people", () => {
        expect(() => splitBill(10, -1)).toThrow(
          "people must be a whole number above 0",
        );
      });
    
      it("rejects fractional people", () => {
        expect(() => splitBill(10, 1.5)).toThrow(
          "people must be a whole number above 0",
        );
      });
    
      it("rejects non-numeric people", () => {
        expect(() => splitBill(10, "2")).toThrow(
          "people must be a whole number above 0",
        );
      });
    
      it("rejects a negative total", () => {
        expect(() => splitBill(-1, 2)).toThrow("total can't be negative");
      });
    });
    ```

    1 message on Pro · $0.0001 to run · 3.8 s · 400 tokens out

  • Claude Sonnet 5.5

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    ```js
    import { describe, it, expect } from "vitest";
    import { splitBill } from "./splitBill";
    
    describe("splitBill", () => {
      describe("normal cases", () => {
        it("splits the total evenly between people", () => {
          expect(splitBill(100, 4)).toBe(25);
        });
    
        it("returns the full total for one person", () => {
          expect(splitBill(80, 1)).toBe(80);
        });
    
        it("applies the tip percentage before splitting", () => {
          expect(splitBill(100, 4, 20)).toBe(30);
        });
    
        it("defaults the tip to 0 when omitted", () => {
          expect(splitBill(60, 3)).toBe(splitBill(60, 3, 0));
        });
    
        it("rounds the result to two decimal places", () => {
          expect(splitBill(100, 3)).toBe(33.33);
        });
    
        it("rounds up when the third decimal is 5 or more", () => {
          expect(splitBill(20, 3)).toBe(6.67);
        });
    
        it("handles decimal totals", () => {
          expect(splitBill(45.5, 2)).toBe(22.75);
        });
      });
    
      describe("edge cases", () => {
        it("returns 0 when the total is 0", () => {
          expect(splitBill(0, 5)).toBe(0);
        });
    
        it("returns 0 when the total is 0 even with a tip", () => {
          expect(splitBill(0, 5, 20)).toBe(0);
        });
    
        it("accepts a tip of 0 explicitly", () => {
          expect(splitBill(50, 2, 0)).toBe(25);
        });
    
        it("handles a tip over 100 percent", () => {
          expect(splitBill(100, 2, 150)).toBe(125);
        });
    
        it("handles decimal tip percentages", () => {
          expect(splitBill(100, 1, 12.5)).toBe(112.5);
        });
    
        it("reduces the amount for a negative tip (discount)", () => {
          expect(splitBill(100, 2, -10)).toBe(45);
        });
    
        it("handles a very large number of people", () => {
          expect(splitBill(1, 1000)).toBe(0);
        });
    
        it("handles very large totals", () => {
          expect(splitBill(1e9, 4)).toBe(250000000);
        });
      });
    
      describe("bad input", () => {
        it("throws when people is 0", () => {
          expect(() => splitBill(100, 0)).toThrow("people must be a whole number above 0");
        });
    
        it("throws when people is negative", () => {
          expect(() => splitBill(100, -2)).toThrow("people must be a whole number above 0");
        });
    
        it("throws when people is a decimal", () => {
          expect(() => splitBill(100, 2.5)).toThrow("people must be a whole number above 0");
        });
    
        it("throws when people is NaN", () => {
          expect(() => splitBill(100, NaN)).toThrow("people must be a whole number above 0");
        });
    
        it("throws when people is Infinity", () => {
          expect(() => splitBill(100, Infinity)).toThrow("people must be a whole number above 0");
        });
    
        it("throws when people is a numeric string", () => {
          expect(() => splitBill(100, "4")).toThrow("people must be a whole number above 0");
        });
    
        it("throws when people is undefined", () => {
          expect(() => splitBill(100)).toThrow("people must be a whole number above 0");
        });
    
        it("throws when people is null", () => {
          expect(() => splitBill(100, null)).toThrow("people must be a whole number above 0");
        });
    
        it("throws when the total is negative", () => {
          expect(() => splitBill(-1, 2)).toThrow("total can't be negative");
        });
    
        it("checks people before total when both are invalid", () => {
          expect(() => splitBill(-1, 0)).toThrow("people must be a whole number above 0");
        });
    
        it("returns NaN when the total is NaN", () => {
          expect(splitBill(NaN, 2)).toBeNaN();
        });
    
        it("returns NaN when the tip is not a number", () => {
          expect(splitBill(100, 2, "abc")).toBeNaN();
        });
      });
    });
    ```

    1 message on Pro · $0.0160 to run · 7.9 s · 1,442 tokens out

  • DeepSeek V4.1 Flash

    Passed

    Graded 4.3 of 5 on average (lowest 3).

    ```js
    import { describe, it, expect } from "vitest";
    import { splitBill } from "./splitBill";
    
    describe("splitBill", () => {
      it("splits the total evenly between people with no tip", () => {
        expect(splitBill(100, 4)).toBe(25);
      });
    
      it("adds the tip percentage before splitting", () => {
        expect(splitBill(100, 4, 20)).toBe(30);
      });
    
      it("rounds each share to two decimal places", () => {
        expect(splitBill(100, 3)).toBe(33.33);
      });
    
      it("treats a missing tip as zero", () => {
        expect(splitBill(50, 2)).toBe(25);
      });
    
      it("splits a zero total into zero shares", () => {
        expect(splitBill(0, 3, 15)).toBe(0);
      });
    
      it("gives the whole total to a single person", () => {
        expect(splitBill(87.5, 1)).toBe(87.5);
      });
    
      it("handles a tip of zero explicitly", () => {
        expect(splitBill(60, 3, 0)).toBe(20);
      });
    
      it("rounds up a repeating decimal share correctly", () => {
        expect(splitBill(10, 3, 0)).toBe(3.33);
      });
    
      it("throws when people is not an integer", () => {
        expect(() => splitBill(100, 2.5)).toThrow("people must be a whole number above 0");
      });
    
      it("throws when people is zero", () => {
        expect(() => splitBill(100, 0)).toThrow("people must be a whole number above 0");
      });
    
      it("throws when people is negative", () => {
        expect(() => splitBill(100, -3)).toThrow("people must be a whole number above 0");
      });
    
      it("throws when people is not a number", () => {
        expect(() => splitBill(100, "4")).toThrow("people must be a whole number above 0");
      });
    
      it("throws when the total is negative", () => {
        expect(() => splitBill(-50, 2)).toThrow("total can't be negative");
      });
    
      it("accepts a total of exactly zero with valid people", () => {
        expect(splitBill(0, 1)).toBe(0);
      });
    });
    ```

    1 message on Pro · $0.0007 to run · 1.8 s · 555 tokens out

  • GLM 5.3

    Passed

    Graded 5.0 of 5 on average (lowest 5).

    ```js
    import { describe, expect, it } from "vitest";
    import { splitBill } from "./splitBill";
    
    describe("splitBill", () => {
      it("splits a total evenly with no tip", () => {
        expect(splitBill(100, 4)).toBe(25);
      });
    
      it("applies the tip before splitting", () => {
        expect(splitBill(100, 4, 10)).toBe(27.5);
      });
    
      it("rounds the share to two decimal places", () => {
        expect(splitBill(10, 3)).toBe(3.33);
      });
    
      it("rounds up to two decimal places when needed", () => {
        expect(splitBill(20, 3)).toBe(6.67);
      });
    
      it("returns zero for a zero total", () => {
        expect(splitBill(0, 5)).toBe(0);
      });
    
      it("accepts a fractional tip percentage", () => {
        expect(splitBill(200, 2, 12.5)).toBe(112.5);
      });
    
      it("throws when people is not an integer", () => {
        expect(() => splitBill(100, 2.5)).toThrow("people must be a whole number above 0");
      });
    
      it("throws when people is zero", () => {
        expect(() => splitBill(100, 0)).toThrow("people must be a whole number above 0");
      });
    
      it("throws when people is negative", () => {
        expect(() => splitBill(100, -2)).toThrow("people must be a whole number above 0");
      });
    
      it("throws when people is NaN", () => {
        expect(() => splitBill(100, NaN)).toThrow("people must be a whole number above 0");
      });
    
      it("throws when people is a non-number string", () => {
        expect(() => splitBill(100, "4")).toThrow("people must be a whole number above 0");
      });
    
      it("throws when total is negative", () => {
        expect(() => splitBill(-10, 2)).toThrow("total can't be negative");
      });
    });
    ```

    1 message on Pro · $0.0024 to run · 1.9 s · 507 tokens out

7. Convert Python to JavaScript

Everyday · 4 of 4 passed

"Errors included" is the part translations drop: the tests check the bad input throws.

The prompt: replace the parts in [brackets]

Convert this [language] function to JavaScript with the same name and behaviour, errors included. Reply with only the JavaScript function in one code block. [paste your code]

Open in llmwise
How we filled it in and checked the replies
Convert this Python function to JavaScript with the same name and behaviour, errors included. Reply with only the JavaScript function in one code block.

def chunk(items, size):
    if size < 1:
        raise ValueError("size must be at least 1")
    return [items[i:i + size] for i in range(0, len(items), size)]
  • Automatic. The function runs against the prompt's tests in a separate Node.js process with a time limit and no file, network or child-process access; it passes when every test passes.
  • GPT-6 Luna

    Passed

    All 6 tests passed.

    ```javascript
    function chunk(items, size) {
      if (size < 1) {
        const error = new Error("size must be at least 1");
        error.name = "ValueError";
        throw error;
      }
    
      if (!Number.isInteger(size)) {
        throw new TypeError("size must be an integer");
      }
    
      const result = [];
      for (let i = 0; i < items.length; i += size) {
        result.push(items.slice(i, i + size));
      }
      return result;
    }
    ```

    1 message on Pro · $0.0001 to run · 5.9 s · 363 tokens out

  • Claude Sonnet 5.5

    Passed

    All 6 tests passed.

    ```javascript
    function chunk(items, size) {
      if (size < 1) {
        throw new RangeError("size must be at least 1");
      }
      const result = [];
      for (let i = 0; i < items.length; i += size) {
        result.push(items.slice(i, i + size));
      }
      return result;
    }
    ```

    1 message on Pro · $0.0022 to run · 1.4 s · 105 tokens out

  • DeepSeek V4.1 Flash

    Passed

    All 6 tests passed.

    ```javascript
    function chunk(items, size) {
        if (size < 1) {
            throw new Error("size must be at least 1");
        }
        const result = [];
        for (let i = 0; i < items.length; i += size) {
            result.push(items.slice(i, i + size));
        }
        return result;
    }
    ```

    1 message on Pro · $0.0002 to run · 0.6 s · 104 tokens out

  • GLM 5.3

    Passed

    All 6 tests passed.

    ```javascript
    function chunk(items, size) {
      if (size < 1) {
        throw new RangeError("size must be at least 1");
      }
      const result = [];
      for (let i = 0; i < items.length; i += size) {
        result.push(items.slice(i, i + size));
      }
      return result;
    }
    ```

    1 message on Pro · $0.0006 to run · 0.6 s · 80 tokens out

8. Explain code to a beginner

Everyday · 4 of 4 passed

Saying what the reader already knows sets the starting point, and a word limit stops a lecture.

The prompt: replace the parts in [brackets]

Explain this line of code to [who], step by step, in under 150 words. [paste your code]

Open in llmwise
How we filled it in and checked the replies
Explain this line of code to someone who has just learned what a loop is, step by step, in under 150 words.

const total = cart.reduce((sum, item) => sum + item.price * item.qty, 0);
  • Accurate: Explains that reduce walks the cart, adding price times quantity to a running total that starts at 0.
  • Right level: Builds on loops, with no unexplained jargon.
  • Clear: Step by step and concise.
  • Automatic rules: at most 150 words.
  • GPT-6 Luna

    Passed

    Graded 4.3 of 5 on average (lowest 3).

    This line adds up the cost of everything in `cart`:
    
    1. `cart.reduce(...)` goes through each item in the cart, one at a time.
    2. `sum` keeps track of the running total. It starts at `0`, the last argument.
    3. `item` is the current cart item.
    4. `item.price * item.qty` calculates the cost of that item.
    5. That cost is added to `sum`.
    6. When all items have been visited, the final total is stored in `total`.

    1 message on Pro · under $0.0001 to run · 1.7 s · 114 tokens out

  • Claude Sonnet 5.5

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    This line adds up the cost of everything in a shopping cart. Think of it as a loop with a running total.
    
    1. **`cart`** is a list of items, each with a `price` and a quantity (`qty`).
    2. **`.reduce(...)`** goes through the list one item at a time, like a loop, and boils it down to a single value.
    3. **`0`** (at the end) is the starting total, like setting `sum = 0` before a loop.
    4. **`(sum, item) => ...`** is what happens on each pass: `sum` is the running total so far, and `item` is the current cart item.
    5. **`sum + item.price * item.qty`** multiplies the item's price by its quantity, then adds that to the running total. The result becomes `sum` for the next pass.
    6. When the list ends, the final `sum` is stored in **`total`**.
    
    Example: prices/quantities of $5×2 and $3×1 give 0 → 10 → 13.

    1 message on Pro · $0.0041 to run · 3.0 s · 302 tokens out

  • DeepSeek V4.1 Flash

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    Think of `reduce` as a loop that keeps a running total.
    
    1. **`cart`** — your list of items (like a shopping cart array).
    2. **`sum`** — the running total. It starts at `0` (that's the last argument).
    3. **`item`** — one item from the cart, one at a time, looping through them.
    4. **`item.price * item.qty`** — the cost of that single item (price times how many).
    5. **`sum + ...`** — add that item's cost to the running total.
    6. Whatever you return becomes the new `sum` for the next item.
    
    After the loop visits every item, `reduce` hands back the final total, stored in **`total`**.
    
    So it's just: *start at 0, visit each item, add its price × quantity, and end with the grand total.*

    1 message on Pro · $0.0003 to run · 1.5 s · 210 tokens out

  • GLM 5.3

    Passed

    Graded 4.7 of 5 on average (lowest 4).

    This line adds up the cost of everything in a shopping cart.
    
    **Step by step:**
    
    1. `cart` is a list of items, where each item has a `price` and a `qty` (quantity).
    2. `reduce` is a special loop-like tool that walks through the list **one item at a time**, building up a single result as it goes.
    3. It starts with `0` (the number at the end) — that's our starting total.
    4. For each item, it runs the function `(sum, item) => sum + item.price * item.qty`:
       - `sum` is the running total so far.
       - `item.price * item.qty` is that item's cost (price × quantity).
       - It adds that cost to `sum`, creating a new total.
    5. After the last item, the final total is saved in a variable called `total`.
    
    In loop terms: it's like starting with `total = 0`, then saying `total = total + item.price * item.qty` for every item.

    1 message on Pro · $0.0011 to run · 1.0 s · 224 tokens out

9. Fix a bug with the smallest change

Hard · 4 of 4 passed

The bug here only shows with numbers of different lengths, which is why the tests include 100, 20, 3 and 4.

The prompt: replace the parts in [brackets]

This function has a bug. Fix it with the smallest change and keep the name. Reply with the fixed function in one code block, then one sentence on what was wrong. [paste your code]

Open in llmwise
How we filled it in and checked the replies
This function has a bug. Fix it with the smallest change and keep the name. Reply with the fixed function in one code block, then one sentence on what was wrong.

function median(numbers) {
  const sorted = numbers.sort();
  const mid = Math.floor(sorted.length / 2);
  return sorted.length % 2 === 0 ? (sorted[mid - 1] + sorted[mid]) / 2 : sorted[mid];
}
  • Automatic. The function runs against the prompt's tests in a separate Node.js process with a time limit and no file, network or child-process access; it passes when every test passes.
  • GPT-6 Luna

    Passed

    All 5 tests passed.

    ```js
    function median(numbers) {
      const sorted = numbers.sort((a, b) => a - b);
      const mid = Math.floor(sorted.length / 2);
      return sorted.length % 2 === 0 ? (sorted[mid - 1] + sorted[mid]) / 2 : sorted[mid];
    }
    ```
    
    The default `sort()` compares values as strings, so numeric sorting requires a comparator.

    1 message on Pro · $0.0001 to run · 2.5 s · 162 tokens out

  • Claude Sonnet 5.5

    Passed

    All 5 tests passed.

    ```javascript
    function median(numbers) {
      const sorted = [...numbers].sort((a, b) => a - b);
      const mid = Math.floor(sorted.length / 2);
      return sorted.length % 2 === 0 ? (sorted[mid - 1] + sorted[mid]) / 2 : sorted[mid];
    }
    ```
    
    The default `sort()` compares elements as strings (so `10` sorts before `2`), and it also mutated the caller's array; a numeric comparator on a copy fixes both.

    1 message on Pro · $0.0032 to run · 1.7 s · 171 tokens out

  • DeepSeek V4.1 Flash

    Passed

    All 5 tests passed.

    ```js
    function median(numbers) {
      const sorted = [...numbers].sort((a, b) => a - b);
      const mid = Math.floor(sorted.length / 2);
      return sorted.length % 2 === 0 ? (sorted[mid - 1] + sorted[mid]) / 2 : sorted[mid];
    }
    ```
    
    The bug was `sort()` with no comparator, which sorts values as strings (so `[10, 9]` becomes `[10, 9]` instead of `[9, 10]`); I also copied the array so the caller's input isn't mutated.

    1 message on Pro · $0.0003 to run · 1.1 s · 224 tokens out

  • GLM 5.3

    Passed

    All 5 tests passed.

    ```js
    function median(numbers) {
      const sorted = [...numbers].sort((a, b) => a - b);
      const mid = Math.floor(sorted.length / 2);
      return sorted.length % 2 === 0 ? (sorted[mid - 1] + sorted[mid]) / 2 : sorted[mid];
    }
    ```
    
    The bug was that `sort()` without a comparator sorts lexicographically (so `[10, 2, 1]` became `[1, 10, 2]`), and it also mutated the input array; the fix uses a numeric comparator on a copy.

    1 message on Pro · $0.0008 to run · 0.8 s · 128 tokens out

10. A SQL query from a plain description

Hard · 4 of 4 passed

Pasting the schema and saying what counts (not cancelled) is what makes the query right; our check ran it on a real database.

The prompt: replace the parts in [brackets]

Here's my database: [paste your schema] [what you need] Write one SQLite query and reply with it in a ```sql code block.

Open in llmwise
How we filled it in and checked the replies
Here's my database:

An online shop's SQLite database:
CREATE TABLE customers (id INTEGER PRIMARY KEY, name TEXT NOT NULL, country TEXT NOT NULL, signed_up TEXT NOT NULL);
CREATE TABLE products (id INTEGER PRIMARY KEY, name TEXT NOT NULL, category TEXT NOT NULL, price_cents INTEGER NOT NULL);
CREATE TABLE orders (id INTEGER PRIMARY KEY, customer_id INTEGER NOT NULL REFERENCES customers(id), ordered_on TEXT NOT NULL, status TEXT NOT NULL);
CREATE TABLE order_items (order_id INTEGER NOT NULL REFERENCES orders(id), product_id INTEGER NOT NULL REFERENCES products(id), quantity INTEGER NOT NULL);
Order statuses are 'completed', 'shipped' and 'cancelled'.

List the customers who placed more than one order, not counting cancelled orders. Return two columns, the customer's name and their number of orders: most orders first, then by name.
Write one SQLite query and reply with it in a ```sql code block.
  • Automatic. The query runs against a small SQLite database and must return the same rows as our reference query (column names and order don't matter; row order does when the prompt asks for it).
  • GPT-6 Luna

    Passed

    Returned the right 5 rows.

    ```sql
    SELECT c.name, COUNT(o.id) AS number_of_orders
    FROM customers AS c
    JOIN orders AS o ON o.customer_id = c.id
    WHERE o.status <> 'cancelled'
    GROUP BY c.id, c.name
    HAVING COUNT(o.id) > 1
    ORDER BY number_of_orders DESC, c.name;
    ```

    1 message on Pro · under $0.0001 to run · 1.4 s · 74 tokens out

  • Claude Sonnet 5.5

    Passed

    Returned the right 5 rows.

    ```sql
    SELECT c.name, COUNT(o.id) AS order_count
    FROM customers c
    JOIN orders o ON o.customer_id = c.id
    WHERE o.status <> 'cancelled'
    GROUP BY c.id, c.name
    HAVING COUNT(o.id) > 1
    ORDER BY order_count DESC, c.name;
    ```

    1 message on Pro · $0.0034 to run · 1.5 s · 123 tokens out

  • DeepSeek V4.1 Flash

    Passed

    Returned the right 5 rows.

    ```sql
    SELECT c.name, COUNT(DISTINCT o.id) AS order_count
    FROM customers c
    JOIN orders o ON o.customer_id = c.id
    WHERE o.status <> 'cancelled'
    GROUP BY c.id, c.name
    HAVING COUNT(DISTINCT o.id) > 1
    ORDER BY order_count DESC, c.name ASC;
    ```

    1 message on Pro · $0.0002 to run · 0.5 s · 133 tokens out

  • GLM 5.3

    Passed

    Returned the right 5 rows.

    ```sql
    SELECT c.name, COUNT(*) AS order_count
    FROM customers c
    JOIN orders o ON o.customer_id = c.id
    WHERE o.status != 'cancelled'
    GROUP BY c.id, c.name
    HAVING COUNT(*) > 1
    ORDER BY order_count DESC, c.name;
    ```

    1 message on Pro · $0.0005 to run · 0.5 s · 63 tokens out

Getting more from these prompts

  • Paste the code and the exact error message, not a description of them.

  • Say "the smallest change" to keep a fix from rewriting everything.

  • For refactors, say "exactly the same behaviour for every input": edge cases are what tidy-ups change.

  • Ask for one code block and nothing else when you'll paste the result straight in.

How we ran and checked them

Each prompt was sent the way llmwise sends a message: the app's own system prompt, each model's own settings, and Pro's reply size limit (8,000 tokens), through OpenRouter. Every reply is shown as it came.

Rubric (graded). The grader model scores the reply from 1 to 5 on each published criterion. It passes with an average of 4 or more and no criterion under 3, and only if it also meets the prompt's automatic rules (length, words it must or mustn't use).

Tests. Automatic. The function runs against the prompt's tests in a separate Node.js process with a time limit and no file, network or child-process access; it passes when every test passes.

Query result. Automatic. The query runs against a small SQLite database and must return the same rows as our reference query (column names and order don't matter; row order does when the prompt asks for it).

The grader is Claude Opus 5.5 at low reasoning effort; its own replies are graded by GPT-6 Astra, so no model grades itself. Its prompt is on our test runs page.

More tested prompts

Questions

Which AI is best at these coding prompts?

In our runs, GPT-6 Luna and Claude Sonnet 5.5 each passed 10 of 10, the most. Half the prompts are checked by running the reply: JavaScript against tests in a sandbox, and SQL against a small database.

Why are the examples in JavaScript?

The code prompts ask for JavaScript because that's what our test harness runs; the same prompts work for Python or any other language if you say so.

How do I get working code from an AI?

Paste the code and the exact error, ask for the smallest change, and ask for the reply in one code block. Then run it: our checks show that code which looks right sometimes isn't.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.