Skip to content

GPT-6 Sol · Writing

GPT-6 Sol for writing: our test runs

GPT-6 Sol ran our 5 writing prompts in llmwise, through the same pipeline your messages take. Here's every reply as it came, whether it passed, what it cost and what it counts as on Pro, next to the other models.

Based on 5 of our test runs on , through OpenRouter with the app's own prompt and settings.

Short answer

In our test runs on September 27, 2026, GPT-6 Sol passed 4 of 5 of our writing prompts (both hard ones): tied 3rd of 15 models. Each reply counted as 1 message on Pro, where it has 125 a month, and cost $0.0023 to run on average.

Prompt by prompt

Each of the 5 writing prompts, and how GPT-6 Sol's reply did.

GPT-6 Sol on each writing prompt
PromptResultCostTime
Announce a second bakery shop on LinkedInEverydayPassed: Graded 4.8 of 5 on average (lowest 4).$0.00273.5 s
Rewrite corporate jargon in plain wordsEverydayFailed: Graded 3.7 of 5 on average (lowest 3).$0.00335.0 s
Decline a meeting and offer two timesEverydayPassed: Graded 5.0 of 5 on average (lowest 5).$0.00141.3 s
A product announcement with five rulesHardPassed: Graded 4.7 of 5 on average (lowest 4).$0.00141.6 s
Argue both sides of free busesHardPassed: Graded 4.3 of 5 on average (lowest 4).$0.00254.2 s

Every writing reply from GPT-6 Sol

  1. Announce a second bakery shop on LinkedIn

    An everyday prompt: “Announce a second bakery shop on LinkedIn”, in full.

    Passed: Graded 4.8 of 5 on average (lowest 4).

    We’re happy to share that Crumb & Co. is opening a second shop in Leeds on 14 November.
    
    When we opened our first bakery, we hoped to make good bread and bakes for our neighbours. The support we’ve had since then has meant more than we can say. Opening another shop is a big step for our small team, and we’re grateful to everyone who has helped us get here.
    
    If you’re in Leeds on opening day, please stop by and say hello. Coffee is on us all day, and we’d love to welcome you into the new space.
    
    Thank you for making room for Crumb & Co. in your routine. We can’t wait to meet more of our neighbours.

    389 tokens in, 189 out (39 of them reasoning) · 3.5 s · $0.0027 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  2. Rewrite corporate jargon in plain words

    An everyday prompt: “Rewrite corporate jargon in plain words”, in full.

    Failed: Graded 3.7 of 5 on average (lowest 3).

    In Q3, teams across the company worked together to put a more customer-focused approach into practice. In the enterprise customer segment, retention measures improved by 12%, while the rate at which customers left fell by 30 basis points (0.3 percentage points). We also adjusted how we connect with customers across channels to make better use of our capacity.

    417 tokens in, 246 out (169 of them reasoning) · 5.0 s · $0.0033 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  3. Decline a meeting and offer two times

    An everyday prompt: “Decline a meeting and offer two times”, in full.

    Passed: Graded 5.0 of 5 on average (lowest 5).

    Subject: Vendor Demo
    
    Hi Tom,
    
    Thanks for the invitation. I’m presenting to the board on Wednesday at 11:00, so I can’t make the demo then. Would Thursday at 10:00 or Friday at 14:00 work instead?
    
    Best,
    [Your Name]

    395 tokens in, 63 out · 1.3 s · $0.0014 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  4. A product announcement with five rules

    A hard prompt: “A product announcement with five rules”, in full.

    Passed: Graded 4.7 of 5 on average (lowest 4).

    Meet the Nordvik 2 desk lamp, designed to put light where you need it. Choose from five dimming steps, turn the arm through 270 degrees, and charge it via USB-C. Priced at €79, it ships on 1 December. Where will you put yours?

    406 tokens in, 63 out · 1.6 s · $0.0014 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  5. Argue both sides of free buses

    A hard prompt: “Argue both sides of free buses”, in full.

    Passed: Graded 4.3 of 5 on average (lowest 4).

    A city should make its buses free because a fare can keep people from jobs, classes, and medical appointments. Removing it would make every trip easier, especially for residents with the least money. Boarding would be faster without payment, and some drivers might leave their cars at home, easing traffic and pollution. Funding buses through taxes treats mobility as a public service: everyone benefits when more people can reach what they need.
    
    But free fares are not free service. Fare revenue helps pay for frequent, reliable buses; lose it without secure replacement funding, and riders may face longer waits and crowded routes. Even with new taxes, a city could spend that money on more drivers and better coverage instead. A zero-dollar fare does little for someone whose bus rarely comes. Targeted discounts for low-income riders would protect access while preserving resources to make the service worth taking.

    384 tokens in, 173 out · 4.2 s · $0.0025 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

How it compares on writing

GPT-6 Sol beside the job's picks among every model. 2 models share the hard-problems pick, and GPT-6 Sol isn't one of them.

GPT-6 Sol beside other models on writing
ModelPassedOn ProCost per reply
GLM 5.3Our pick: best value4 of 5250 a month on Pro$0.0006
GPT-6 LunaOur pick: hard problems (shared), everyday5 of 560 a day on Pro$0.0001
GPT-6 Sol4 of 5125 a month on Pro$0.0023
On Pro: Pro's count on each model. Cost: what OpenRouter charged us per reply, on average.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

How these runs were done

Rubric (graded). The grader model scores the reply from 1 to 5 on each published criterion. It passes with an average of 4 or more and no criterion under 3, and only if it also meets the prompt's automatic rules (length, words it must or mustn't use).

The grader is Claude Opus 5.5 at low reasoning effort; its own replies are graded by GPT-6 Astra, so no model grades itself. Its prompt and every rubric are on the methods page.

How the runs were done, and every writing prompt.

GPT-6 Sol, and writing, elsewhere

Questions

Is GPT-6 Sol good for writing?

In our test runs it passed 4 of 5 writing prompts, tied 3rd of the 15 models in llmwise. Every reply is on this page, so you can judge them yourself.

How many of my messages does a writing reply from GPT-6 Sol use?

1 message each on Pro, where it has 125 a month on Pro. The price of a message is fixed and shown before you send it, however long the reply.

How were these runs done?

The same way for every model: each prompt sent through llmwise's own pipeline, each reply checked the same way. The methods page has every prompt and how each is scored.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.