GPT-6 Sol · Writing
GPT-6 Sol for writing: our test runs
GPT-6 Sol ran our 5 writing prompts in llmwise, through the same pipeline your messages take. Here's every reply as it came, whether it passed, what it cost and what it counts as on Pro, next to the other models.
Based on 5 of our test runs on , through OpenRouter with the app's own prompt and settings.
Short answer
In our test runs on September 27, 2026, GPT-6 Sol passed 4 of 5 of our writing prompts (both hard ones): tied 3rd of 15 models. Each reply counted as 1 message on Pro, where it has 125 a month, and cost $0.0023 to run on average.
Prompt by prompt
Each of the 5 writing prompts, and how GPT-6 Sol's reply did.
| Prompt | Result | Cost | Time |
|---|---|---|---|
| Announce a second bakery shop on LinkedInEveryday | Passed: Graded 4.8 of 5 on average (lowest 4). | $0.0027 | 3.5 s |
| Rewrite corporate jargon in plain wordsEveryday | Failed: Graded 3.7 of 5 on average (lowest 3). | $0.0033 | 5.0 s |
| Decline a meeting and offer two timesEveryday | Passed: Graded 5.0 of 5 on average (lowest 5). | $0.0014 | 1.3 s |
| A product announcement with five rulesHard | Passed: Graded 4.7 of 5 on average (lowest 4). | $0.0014 | 1.6 s |
| Argue both sides of free busesHard | Passed: Graded 4.3 of 5 on average (lowest 4). | $0.0025 | 4.2 s |
Every writing reply from GPT-6 Sol
Announce a second bakery shop on LinkedIn
An everyday prompt: “Announce a second bakery shop on LinkedIn”, in full.
Passed: Graded 4.8 of 5 on average (lowest 4).
We’re happy to share that Crumb & Co. is opening a second shop in Leeds on 14 November. When we opened our first bakery, we hoped to make good bread and bakes for our neighbours. The support we’ve had since then has meant more than we can say. Opening another shop is a big step for our small team, and we’re grateful to everyone who has helped us get here. If you’re in Leeds on opening day, please stop by and say hello. Coffee is on us all day, and we’d love to welcome you into the new space. Thank you for making room for Crumb & Co. in your routine. We can’t wait to meet more of our neighbours.
389 tokens in, 189 out (39 of them reasoning) · 3.5 s · $0.0027 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Rewrite corporate jargon in plain words
An everyday prompt: “Rewrite corporate jargon in plain words”, in full.
Failed: Graded 3.7 of 5 on average (lowest 3).
In Q3, teams across the company worked together to put a more customer-focused approach into practice. In the enterprise customer segment, retention measures improved by 12%, while the rate at which customers left fell by 30 basis points (0.3 percentage points). We also adjusted how we connect with customers across channels to make better use of our capacity.
417 tokens in, 246 out (169 of them reasoning) · 5.0 s · $0.0033 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Decline a meeting and offer two times
An everyday prompt: “Decline a meeting and offer two times”, in full.
Passed: Graded 5.0 of 5 on average (lowest 5).
Subject: Vendor Demo Hi Tom, Thanks for the invitation. I’m presenting to the board on Wednesday at 11:00, so I can’t make the demo then. Would Thursday at 10:00 or Friday at 14:00 work instead? Best, [Your Name]
395 tokens in, 63 out · 1.3 s · $0.0014 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
A product announcement with five rules
A hard prompt: “A product announcement with five rules”, in full.
Passed: Graded 4.7 of 5 on average (lowest 4).
Meet the Nordvik 2 desk lamp, designed to put light where you need it. Choose from five dimming steps, turn the arm through 270 degrees, and charge it via USB-C. Priced at €79, it ships on 1 December. Where will you put yours?
406 tokens in, 63 out · 1.6 s · $0.0014 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
Argue both sides of free buses
A hard prompt: “Argue both sides of free buses”, in full.
Passed: Graded 4.3 of 5 on average (lowest 4).
A city should make its buses free because a fare can keep people from jobs, classes, and medical appointments. Removing it would make every trip easier, especially for residents with the least money. Boarding would be faster without payment, and some drivers might leave their cars at home, easing traffic and pollution. Funding buses through taxes treats mobility as a public service: everyone benefits when more people can reach what they need. But free fares are not free service. Fare revenue helps pay for frequent, reliable buses; lose it without secure replacement funding, and riders may face longer waits and crowded routes. Even with new taxes, a city could spend that money on more drivers and better coverage instead. A zero-dollar fare does little for someone whose bus rarely comes. Targeted discounts for low-income riders would protect access while preserving resources to make the service worth taking.
384 tokens in, 173 out · 4.2 s · $0.0025 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·
How it compares on writing
GPT-6 Sol beside the job's picks among every model. 2 models share the hard-problems pick, and GPT-6 Sol isn't one of them.
| Model | Passed | On Pro | Cost per reply |
|---|---|---|---|
| GLM 5.3Our pick: best value | 4 of 5 | 250 a month on Pro | $0.0006 |
| GPT-6 LunaOur pick: hard problems (shared), everyday | 5 of 5 | 60 a day on Pro | $0.0001 |
| GPT-6 Sol | 4 of 5 | 125 a month on Pro | $0.0023 |
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
How these runs were done
Rubric (graded). The grader model scores the reply from 1 to 5 on each published criterion. It passes with an average of 4 or more and no criterion under 3, and only if it also meets the prompt's automatic rules (length, words it must or mustn't use).
The grader is Claude Opus 5.5 at low reasoning effort; its own replies are graded by GPT-6 Astra, so no model grades itself. Its prompt and every rubric are on the methods page.
GPT-6 Sol, and writing, elsewhere
Questions
Is GPT-6 Sol good for writing?
In our test runs it passed 4 of 5 writing prompts, tied 3rd of the 15 models in llmwise. Every reply is on this page, so you can judge them yourself.
How many of my messages does a writing reply from GPT-6 Sol use?
1 message each on Pro, where it has 125 a month on Pro. The price of a message is fixed and shown before you send it, however long the reply.
How were these runs done?
The same way for every model: each prompt sent through llmwise's own pipeline, each reply checked the same way. The methods page has every prompt and how each is scored.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.