GPT · Coding
GPT-6 for coding
llmwise has 4 of OpenAI's GPT models, from GPT-6 Luna (up to 60 messages a day on Pro) to GPT-6 Astra (up to 31 messages a month). We ran the same coding prompts on every one and published every reply: which GPT model to use, from the results, what each costs per message, and how to get more out of it.
Based on 20 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
In our test runs on September 29, 2026, all 4 models passed 5 of 5 coding prompts, so these prompts don't pick one for hard problems. For value, GPT-6.1 Sol (5 of 5), 125 a month on Pro; for everyday coding, GPT-6 Luna (5 of 5), from the daily count.
Our picks for coding
Hard problems
Shared by 4 models
All 4 models passed 5 of 5, both hard ones: GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol and GPT-6 Luna. These prompts don't tell them apart, so they share the pick.
Best value
Passed 5 of 5 coding prompts, with 125 a month on Pro.
Everyday
Passed 5 of 5 coding prompts; an everyday model, so its messages come from the daily count (60 a day on Pro), not the monthly allowance.
These picks aren't our opinion: they're what the results below give, by these rules, among the GPT models in llmwise. They change when the results do.
- Hard problems: the model that passed the most prompts and, of those, the most hard ones. Models level on both share the pick: the prompts don't tell them apart, so we don't break the tie by price or by name.
- Best value: among the models that draw on the monthly allowance, the one with the most messages on Pro that passed no more than one prompt fewer than the top model. Ties go to the one that passed more, then to the lower cost per reply.
- Everyday: among the cheapest models on the page (the everyday models, which come from the daily count, when the page has any), the one that passed the most. Ties go to the one that passed more of the hard prompts, then to the lower cost per reply.
Our coding test runs, model by model
How each model did on our 5 coding prompts, what each reply counted as on Pro, and what it cost to run.
| Model | Passed | Hard ones | Messages used on Pro | Cost per reply | Time per reply |
|---|---|---|---|---|---|
| GPT-6 AstraOpenAI | 5 of 5 | 2 of 2 | 1 each, of 31 a month on Pro | $0.0223 | 6.8 s |
| GPT-6.1 SolOpenAI | 5 of 5 | 2 of 2 | 1 each, of 125 a month on Pro | $0.0023 | 4.9 s |
| GPT-6 SolOpenAI | 5 of 5 | 2 of 2 | 1 each, of 125 a month on Pro | $0.0052 | 5.7 s |
| GPT-6 LunaOpenAI | 5 of 5 | 2 of 2 | 1 each, of 60 a day on Pro | $0.0003 | 4.9 s |
One prompt, side by side
Each model's reply to “Evaluate an arithmetic expression, no eval”, as it came, with whether it passed and what it cost.
The prompt: Evaluate an arithmetic expression, no eval (hard)
Write a JavaScript function evaluate(expression) that computes an arithmetic expression given as a string and returns the number. It supports numbers like 3, 0.5 and 12.25; the operators + - * / and ^ (power); parentheses; unary minus; and spaces anywhere. - ^ binds tighter than unary minus and is right-associative, so -2^2 is -4 and 2^3^2 is 512. A unary minus may follow ^, as in 2^-1, which is 0.5. - * and / bind tighter than + and -. Otherwise operators of the same level go left to right, so 8/4/2 is 1. - Throw an Error for input that isn't a valid expression, such as "2 +", "(1" or "1 2". Don't use eval, Function or any library. Reply with the whole function in one ```javascript code block: plain JavaScript for Node.js 22, no imports, no TypeScript.
GPT-6 Astra
Passed: All 15 tests passed.
```javascript function evaluate(expression) { if (typeof expression !== "string") { throw new Error("Expression must be a string"); } let pos = 0; function skipSpaces() { while (pos < expression.length && /\s/.test(expression[pos])) { pos++; } } function consume(character) { skipSpaces(); if (expression[pos] === character) { pos++; return true; } return false; } function fail(message) { throw new Error(`${message} at position ${pos}`); } function parsePrimary() { if (consume("(")) { const value = parseSum(); if (!consume(")")) fail("Expected ')'"); return value; } skipSpaces(); const match = /^(?:\d+(?:\.\d*)?|\.\d+)/.exec(expression.slice(pos)); if (!match) fail("Expected a number or '('"); pos += match[0].length; return Number(match[0]); } function parsePower() { const base = parsePrimary(); // Parsing the exponent as unary allows 2^-1 and makes ^ right-associative. return consume("^") ? base ** parseUnary() : base; } function parseUnary() { // Power binds more tightly than unary minus. return consume("-") ? -parseUnary() : parsePower(); } function parseProduct() { let value = parseUnary(); while (true) { if (consume("*")) value *= parseUnary(); else if (consume("/")) value /= parseUnary(); else return value; } } function parseSum() { let value = parseProduct(); while (true) { if (consume("+")) value += parseProduct(); else if (consume("-")) value -= parseProduct(); else return value; } } const result = parseSum(); skipSpaces(); if (pos !== expression.length) fail("Unexpected input"); return result; } ```549 tokens in, 640 out (200 of them reasoning) · 12.0 s · $0.0375 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·
GPT-6.1 Sol
Passed: All 15 tests passed.
```javascript function evaluate(expression) { if (typeof expression !== "string") { throw new Error("Expression must be a string"); } let pos = 0; function skipSpaces() { while (pos < expression.length && /\s/.test(expression[pos])) { pos++; } } function take(char) { skipSpaces(); if (expression[pos] === char) { pos++; return true; } return false; } function invalid() { throw new Error(`Invalid expression at position ${pos}`); } function primary() { skipSpaces(); if (take("(")) { const value = additive(); if (!take(")")) invalid(); return value; } const match = /^(?:\d+(?:\.\d*)?|\.\d+)/.exec(expression.slice(pos)); if (!match) invalid(); pos += match[0].length; return Number(match[0]); } function power() { const base = primary(); // Parsing the exponent as unary permits 2^-1 and right-associative powers. return take("^") ? base ** unary() : base; } function unary() { return take("-") ? -unary() : power(); } function multiplicative() { let value = unary(); while (true) { if (take("*")) value *= unary(); else if (take("/")) value /= unary(); else return value; } } function additive() { let value = multiplicative(); while (true) { if (take("+")) value += multiplicative(); else if (take("-")) value -= multiplicative(); else return value; } } const result = additive(); skipSpaces(); if (pos !== expression.length) invalid(); return result; } ```551 tokens in, 632 out (228 of them reasoning) · 7.8 s · $0.0037 · 1 message on Pro · answered by openai/gpt-6.1-sol via OpenAI ·
GPT-6 Luna
Passed: All 15 tests passed.
```javascript function evaluate(expression) { if (typeof expression !== "string") { throw new Error("Expression must be a string"); } let index = 0; function skipSpaces() { while (index < expression.length && /\s/.test(expression[index])) { index++; } } function parsePrimary() { skipSpaces(); if (expression[index] === "(") { index++; const value = parseExpression(); skipSpaces(); if (expression[index] !== ")") throw new Error("Missing closing parenthesis"); index++; return value; } const start = index; let digitsBefore = 0; while (/\d/.test(expression[index] || "")) { index++; digitsBefore++; } let digitsAfter = 0; if (expression[index] === ".") { index++; while (/\d/.test(expression[index] || "")) { index++; digitsAfter++; } } if (digitsBefore === 0 && digitsAfter === 0) { throw new Error("Expected a number or parenthesized expression"); } return Number(expression.slice(start, index)); } // Unary minus has lower precedence than exponentiation, but is allowed // on an exponent's right-hand side. function parseUnary() { skipSpaces(); if (expression[index] === "-") { index++; return -parseUnary(); } return parsePower(); } function parsePower() { let value = parsePrimary(); skipSpaces(); if (expression[index] === "^") { index++; value **= parseUnary(); } return value; } function parseProduct() { let value = parseUnary(); while (true) { skipSpaces(); const operator = expression[index]; if (operator !== "*" && operator !== "/") break; index++; const right = parseUnary(); value = operator === "*" ? value * right : value / right; } return value; } function parseExpression() { let value = parseProduct(); while (true) { skipSpaces(); const operator = expression[index]; if (operator !== "+" && operator !== "-") break; index++; const right = parseProduct(); value = operator === "+" ? value + right : value - right; } return value; } const result = parseExpression(); skipSpaces(); if (index !== expression.length) throw new Error("Unexpected input"); return result; } ```549 tokens in, 1,007 out (455 of them reasoning) · 8.8 s · $0.0006 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·
The coding prompts, and how they're scored
Tests. Automatic. The function runs against the prompt's tests in a separate Node.js process with a time limit and no file, network or child-process access; it passes when every test passes.
Each prompt was sent the way llmwise sends a message in a side-by-side comparison, which offers no tools: the app's own system prompt, the model's own settings, and Pro's reply size limit (8,000 tokens). Read all five coding prompts and how every reply was scored.
Prompts like these to try yourself
GPT in llmwise
| Model | On Pro | On Free | Context window | Images | PDFs | Reasoning | API price per 1M, in / out |
|---|---|---|---|---|---|---|---|
| GPT-6 AstraOpenAI | 31/mo on Pro | No | 1.05M tokens | Yes | Whole file | Yes | $10.00 / $50.00 |
| GPT-6.1 SolOpenAI | 125/mo on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes | $2.00 / $10.00 |
| GPT-6 SolOpenAI | 125/mo on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes | $2.00 / $10.00 |
| GPT-6 LunaOpenAI | 60/day on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes | $0.10 / $0.50 |
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
What matters for coding
Room for your code
The more of your code a model can take in at once, the fewer details it has to guess. Context windows are listed below.
Reasoning on hard problems
Models that reason before answering tend to do better on problems with several moving parts, like a bug that crosses files.
Running the code
Code that has been run beats code that looks right. A model that can run and test its own code catches its own mistakes.
Cost per try
Coding means many small tries. A cheap model for quick questions and a stronger one for the hard ones keeps the cost down.
GPT for coding in llmwise
Code opens as a document
Ask for a script or a component and it opens beside the chat as a code document that keeps every version, so you can step back to an earlier one.
Run code (paid plans)
On a paid plan the model can run Python or Node.js in an isolated sandbox once you approve it, read the output and fix what failed. A run counts as 1 Claude Haiku 4.5 message and stops after 60 seconds.
HTML apps run in the chat
Ask for a calculator, a small tool or a page and it runs live as an HTML app in the side panel.
Switch models mid-chat
Start on a cheaper model; if the answer isn't good enough, switch models in the same chat. The next model sees the whole conversation, so you don't paste anything twice.
In llmwise, a message to GPT goes to its maker, OpenAI, or through OpenRouter when llmwise can't reach the maker directly. See the Privacy Policy.
Tips
Paste the exact error and the code that produced it, not just a description.
Say which language, version and libraries you use.
For a big change, ask for a plan first, then one step at a time.
Pick the Coder persona for code-first answers that explain their trade-offs.
If a model gets stuck, switch to another in the same chat: it sees the code so far.
Questions
Is GPT good for coding?
In our test runs on September 29, 2026, GPT-6 Astra passed 5 of 5, GPT-6.1 Sol passed 5 of 5, GPT-6 Sol passed 5 of 5, GPT-6 Luna passed 5 of 5 of our coding prompts, against a best result of 5 of 5 among all 19 models. Every reply is published on this page and the methods page, so you can judge them yourself.
Which GPT model should I use for coding?
Start with GPT-6 Luna for quick syntax questions, regexes, small scripts and boilerplate, and move up to GPT-6 Astra for large refactors, subtle bugs and design questions where a wrong answer costs you hours.
Can I use GPT for coding for free?
Yes: GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna are on the Free plan (a one-time trial of 5 messages). GPT-6 Astra needs a paid plan.
Can llmwise run the code it writes?
On a paid plan, yes. The model writes Python or Node.js, you approve the run, and it executes in an isolated sandbox with no internet access except package registries, stopping after 60 seconds. It reads the output and can fix what failed. Each run counts as 1 Claude Haiku 4.5 message.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.