Skip to content

GPT-6.1 Sol · Writing

GPT-6.1 Sol for writing: our test runs

GPT-6.1 Sol ran our 5 writing prompts in llmwise, through the same pipeline your messages take. Here's every reply as it came, whether it passed, what it cost and what it counts as on Pro, next to the other models.

Based on 5 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

In our test runs on September 29, 2026, GPT-6.1 Sol passed 4 of 5 of our writing prompts (both hard ones): tied 3rd of 17 models. Each reply counted as 1 message on Pro, where it has 125 a month, and cost $0.0011 to run on average.

Prompt by prompt

Each of the 5 writing prompts, and how GPT-6.1 Sol's reply did.

GPT-6.1 Sol on each writing prompt
PromptResultCostTime
Announce a second bakery shop on LinkedInEverydayPassed: Graded 4.5 of 5 on average (lowest 4).$0.00134.2 s
Rewrite corporate jargon in plain wordsEverydayFailed: Graded 3.7 of 5 on average (lowest 3).$0.00133.3 s
Decline a meeting and offer two timesEverydayPassed: Graded 5.0 of 5 on average (lowest 5).$0.00071.8 s
A product announcement with five rulesHardPassed: Graded 4.3 of 5 on average (lowest 4).$0.00091.9 s
Argue both sides of free busesHardPassed: Graded 4.3 of 5 on average (lowest 4).$0.00144.6 s

Every writing reply from GPT-6.1 Sol

  1. Announce a second bakery shop on LinkedIn

    An everyday prompt: “Announce a second bakery shop on LinkedIn”, in full.

    Passed: Graded 4.5 of 5 on average (lowest 4).

    We’ve got some lovely news to share: Crumb & Co. is opening a second shop in Leeds on 14 November.
    
    As a small bakery, this is a big step for us, and we’re grateful to everyone who has helped us get here. Every visit, kind word and recommendation has meant a lot.
    
    We’re looking forward to welcoming familiar faces and meeting new neighbours, with the same care and love for baking that we bring to our first shop.
    
    To celebrate, we’ll be offering free coffee on opening day. Come in, say hello and help us mark the start of this next chapter.
    
    We can’t wait to welcome you.
    
    See you on 14 November!

    391 tokens in, 188 out (45 of them reasoning) · 4.2 s · $0.0013 · 1 message on Pro · answered by openai/gpt-6.1-sol via OpenAI ·

  2. Rewrite corporate jargon in plain words

    An everyday prompt: “Rewrite corporate jargon in plain words”, in full.

    Failed: Graded 3.7 of 5 on average (lowest 3).

    In the third quarter, teams across departments worked together to put customers first. The initiative increased customer retention by 12% and reduced the rate at which customers leave by 0.3 percentage points (30 basis points) in the enterprise segment. We also adjusted customer contact points across all channels to make better use of available capacity.

    419 tokens in, 185 out (113 of them reasoning) · 3.3 s · $0.0013 · 1 message on Pro · answered by openai/gpt-6.1-sol via OpenAI ·

  3. Decline a meeting and offer two times

    An everyday prompt: “Decline a meeting and offer two times”, in full.

    Passed: Graded 5.0 of 5 on average (lowest 5).

    Subject: Vendor Demo – Alternative Times
    
    Hi Tom,
    
    Thanks for the invitation. Unfortunately, I can’t attend the vendor demo on Wednesday at 11:00, as I’ll be presenting to the board then.
    
    Would Thursday at 10:00 or Friday at 14:00 work instead?
    
    Best,
    [Your Name]

    397 tokens in, 70 out · 1.8 s · $0.0007 · 1 message on Pro · answered by openai/gpt-6.1-sol via OpenAI ·

  4. A product announcement with five rules

    A hard prompt: “A product announcement with five rules”, in full.

    Passed: Graded 4.3 of 5 on average (lowest 4).

    Meet the Nordvik 2 desk lamp, designed to bring flexible lighting to your workspace. It dims in five steps, letting you choose the right brightness for focused work or a quieter evening. Its arm turns 270 degrees, so you can direct light where you need it, and it charges from USB-C for convenient powering.
    
    The Nordvik 2 costs €79 and ships on 1 December.
    
    Ready to see your desk in a new light?

    408 tokens in, 94 out · 1.9 s · $0.0009 · 1 message on Pro · answered by openai/gpt-6.1-sol via OpenAI ·

  5. Argue both sides of free buses

    A hard prompt: “Argue both sides of free buses”, in full.

    Passed: Graded 4.3 of 5 on average (lowest 4).

    A city should make its buses free because mobility is essential infrastructure, not a privilege reserved for those who can afford each trip. Removing fares gives low-income residents immediate access to jobs, schools, healthcare, and community life, while eliminating fare disputes and speeding boarding. If paired with reliable service, free buses can draw people out of cars, easing congestion and pollution. Funding transit broadly recognizes a simple truth: everyone benefits when more people can move freely, whether or not they ride.
    
    But a city should prioritize buses people can depend on over buses that cost nothing. Fare revenue can support frequent service, safe stops, and wider coverage; eliminating it without secure replacement funding risks longer waits and overcrowding. Free access means little when the bus never arrives. Universal subsidies also benefit riders who can comfortably pay, while targeted discounts protect those who cannot. With limited public money, reliable, affordable transit may deliver more opportunity—and attract more drivers—than free but deteriorating service.

    386 tokens in, 195 out · 4.6 s · $0.0014 · 1 message on Pro · answered by openai/gpt-6.1-sol via OpenAI ·

How it compares on writing

GPT-6.1 Sol beside the job's picks among every model. 2 models share the hard-problems pick, and GPT-6.1 Sol isn't one of them.

GPT-6.1 Sol beside other models on writing
ModelPassedOn ProCost per reply
GLM 5.3Our pick: best value4 of 5250 a month on Pro$0.0006
GPT-6 LunaOur pick: hard problems (shared), everyday5 of 560 a day on Pro$0.0001
GPT-6.1 Sol4 of 5125 a month on Pro$0.0011
On Pro: Pro's count on each model. Cost: what OpenRouter charged us per reply, on average.

Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.

GPT-6.1 Sol next to each model on writing

GPT-6.1 Sol against each model on the same 5 writing prompts. One run each, through OpenRouter: a wait depends on the provider and the load that day, so a lead under 10% counts as close.

  • Against Claude Fable 5.1: GPT-6.1 Sol answered 1.8× sooner (3.3 s to 6.0 s), cost 15.2× less, and passed 4 of 5 to its 3. Only GPT-6.1 Sol passed A product announcement with five rules; Argue both sides of free buses. Only Claude Fable 5.1 passed Rewrite corporate jargon in plain words.

  • Against Claude Opus 5.5: GPT-6.1 Sol answered 2.5× sooner (3.3 s to 8.2 s), cost 16.3× less, and passed 4 of 5 to its 4. Only GPT-6.1 Sol passed Argue both sides of free buses. Only Claude Opus 5.5 passed Rewrite corporate jargon in plain words.

  • Against Claude Sonnet 5.5: GPT-6.1 Sol answered 1.4× later (3.3 s to 2.3 s), cost 3.0× less, and passed 4 of 5 to its 4. Only GPT-6.1 Sol passed Argue both sides of free buses. Only Claude Sonnet 5.5 passed Rewrite corporate jargon in plain words.

  • Against Claude Sonnet 5: GPT-6.1 Sol answered 1.4× sooner (3.3 s to 4.5 s), cost 3.0× less, and passed 4 of 5 to its 4. Only GPT-6.1 Sol passed Announce a second bakery shop on LinkedIn. Only Claude Sonnet 5 passed Rewrite corporate jargon in plain words.

  • Against Claude Haiku 4.5: GPT-6.1 Sol answered 1.5× later (3.3 s to 2.2 s), cost within 10% of it, and passed 4 of 5 to its 4. Only GPT-6.1 Sol passed A product announcement with five rules. Only Claude Haiku 4.5 passed Rewrite corporate jargon in plain words.

  • Against GPT-6 Astra: GPT-6.1 Sol answered 1.5× sooner (3.3 s to 4.8 s), cost 9.8× less, and passed 4 of 5 to its 5. Only GPT-6 Astra passed Rewrite corporate jargon in plain words.

  • Against GPT-6 Sol: GPT-6.1 Sol answered within 10% of its time (3.3 s to 3.5 s), cost 2.0× less, and passed 4 of 5 to its 4.

  • Against GPT-6 Luna: GPT-6.1 Sol answered 1.4× later (3.3 s to 2.4 s), cost 11.4× more, and passed 4 of 5 to its 5. Only GPT-6 Luna passed Rewrite corporate jargon in plain words.

  • Against Gemini 3.1 Pro (preview): GPT-6.1 Sol answered 2.8× sooner (3.3 s to 9.1 s), cost 8.1× less, and passed 4 of 5 to its 4. Only GPT-6.1 Sol passed Argue both sides of free buses. Only Gemini 3.1 Pro (preview) passed Rewrite corporate jargon in plain words.

  • Against Gemini 3.8 Flash: GPT-6.1 Sol answered 1.2× sooner (3.3 s to 3.9 s), cost 1.2× less, and passed 4 of 5 to its 3. Only GPT-6.1 Sol passed A product announcement with five rules.

  • Against DeepSeek V4.1 Flash: GPT-6.1 Sol answered within 10% of its time (3.3 s to 3.1 s), cost 4.0× more, and passed 4 of 5 to its 4.

  • Against DeepSeek V4 Pro: GPT-6.1 Sol answered 1.5× later (3.3 s to 2.2 s), cost 2.5× more, and passed 4 of 5 to its 3. Only GPT-6.1 Sol passed A product announcement with five rules; Argue both sides of free buses. Only DeepSeek V4 Pro passed Rewrite corporate jargon in plain words.

  • Against Grok 4.7: GPT-6.1 Sol answered 1.2× sooner (3.3 s to 4.0 s), cost 3.7× less, and passed 4 of 5 to its 3. Only GPT-6.1 Sol passed Announce a second bakery shop on LinkedIn.

  • Against Kimi K3: GPT-6.1 Sol answered 1.1× later (3.3 s to 3.0 s), cost 3.4× less, and passed 4 of 5 to its 4. Only GPT-6.1 Sol passed A product announcement with five rules. Only Kimi K3 passed Rewrite corporate jargon in plain words.

  • Against GLM 5.3: GPT-6.1 Sol answered 1.9× later (3.3 s to 1.8 s), cost 2.0× more, and passed 4 of 5 to its 4. Only GPT-6.1 Sol passed Argue both sides of free buses. Only GLM 5.3 passed Rewrite corporate jargon in plain words.

  • Against GLM 5.3 Flash: GPT-6.1 Sol answered 2.6× sooner (3.3 s to 8.5 s), cost 3.5× more, and passed 4 of 5 to its 2. Only GPT-6.1 Sol passed A product announcement with five rules; Argue both sides of free buses.

How these runs were done

Rubric (graded). The grader model scores the reply from 1 to 5 on each published criterion. It passes with an average of 4 or more and no criterion under 3, and only if it also meets the prompt's automatic rules (length, words it must or mustn't use).

The grader is Claude Opus 5.5 at low reasoning effort; its own replies are graded by GPT-6 Astra, so no model grades itself. Its prompt and every rubric are on the methods page.

How the runs were done, and every writing prompt.

GPT-6.1 Sol, and writing, elsewhere

Questions

Is GPT-6.1 Sol good for writing?

In our test runs it passed 4 of 5 writing prompts, tied 3rd of the 17 models in llmwise. Every reply is on this page, so you can judge them yourself.

How many of my messages does a writing reply from GPT-6.1 Sol use?

1 message each on Pro, where it has 125 a month on Pro. The price of a message is fixed and shown before you send it, however long the reply.

How were these runs done?

The same way for every model: each prompt sent through llmwise's own pipeline, each reply checked the same way. The methods page has every prompt and how each is scored.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.