DeepSeek V4.1 Flash · Writing
DeepSeek V4.1 Flash for writing: our test runs
DeepSeek V4.1 Flash ran our 5 writing prompts in llmwise, through the same pipeline your messages take. Here's every reply as it came, whether it passed, what it cost and what it counts as on Pro, next to the other models.
Based on 5 of our test runs on , through OpenRouter with the app's own prompt and settings.
Short answer
In our test runs on September 27, 2026, DeepSeek V4.1 Flash passed 4 of 5 of our writing prompts (both hard ones): tied 3rd of 15 models. Each reply counted as 1 message on Pro, where it has 60 a day, and cost $0.0003 to run on average.
Prompt by prompt
Each of the 5 writing prompts, and how DeepSeek V4.1 Flash's reply did.
| Prompt | Result | Cost | Time |
|---|---|---|---|
| Announce a second bakery shop on LinkedInEveryday | Passed: Graded 5.0 of 5 on average (lowest 5). | $0.0002 | 2.9 s |
| Rewrite corporate jargon in plain wordsEveryday | Failed: Graded 3.7 of 5 on average (lowest 3). | $0.0003 | 1.2 s |
| Decline a meeting and offer two timesEveryday | Passed: Graded 5.0 of 5 on average (lowest 5). | $0.0002 | 3.1 s |
| A product announcement with five rulesHard | Passed: Graded 4.3 of 5 on average (lowest 4). | $0.0004 | 6.2 s |
| Argue both sides of free busesHard | Passed: Graded 4.3 of 5 on average (lowest 4). | $0.0003 | 3.9 s |
Every writing reply from DeepSeek V4.1 Flash
Announce a second bakery shop on LinkedIn
An everyday prompt: “Announce a second bakery shop on LinkedIn”, in full.
Passed: Graded 5.0 of 5 on average (lowest 5).
Some news we've been sitting on for a while: Crumb & Co. is opening a second shop, in Leeds, on 14 November. We started six years ago with one oven, one counter and a lot of very early mornings. The idea was simple — good bread, made slowly, sold to people who'd walk a bit out of their way for it. That we get to do it in a second place still feels a little unreal. The new shop will have the same loaves, the same pastries, the same slightly chaotic Saturday mornings. There's more seating, and a window that actually gets sun. To say thank you, coffee is free all day on opening day. No catch, no app, no code — just come in and we'll make you one. We'd love to see you there. Thank you for getting us this far. — The Crumb & Co. team
428 tokens in, 203 out (19 of them reasoning) · 2.9 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via GMICloud ·
Rewrite corporate jargon in plain words
An everyday prompt: “Rewrite corporate jargon in plain words”, in full.
Failed: Graded 3.7 of 5 on average (lowest 3).
In the third quarter, teams from across the company worked together on a project that put customers first. For our large business customers, retention rose 12%, and churn fell by 30 basis points. We also adjusted our channels — online, in-store, and other ways customers reach us — to use our resources more efficiently.
459 tokens in, 129 out (61 of them reasoning) · 1.2 s · $0.0003 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Parasail ·
Decline a meeting and offer two times
An everyday prompt: “Decline a meeting and offer two times”, in full.
Passed: Graded 5.0 of 5 on average (lowest 5).
**Subject:** Re: Vendor demo — Wednesday 11:00 Hi Tom, Thanks for the invitation to the vendor demo. Unfortunately, I can't make Wednesday at 11:00 — I'm presenting to the board at that time. I'd still like to join, though. Would either of these work instead? - Thursday at 10:00 - Friday at 14:00 Let me know which suits you, and I'll confirm. If neither works, feel free to send notes or a recording and I'll follow up. Thanks again for thinking of me. Best regards, [Your name]
434 tokens in, 157 out (30 of them reasoning) · 3.1 s · $0.0002 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Parasail ·
A product announcement with five rules
A hard prompt: “A product announcement with five rules”, in full.
Passed: Graded 4.3 of 5 on average (lowest 4).
Introducing the Nordvik 2, a desk lamp built for the way you actually work. Five dimming steps take you from bright task light to a soft evening glow. The arm turns a full 270 degrees, so you can aim light exactly where it is needed. Recharging is simple, via USB-C. The Nordvik 2 costs €79 and ships on 1 December. Would you like to see how it looks on your desk?
445 tokens in, 495 out (404 of them reasoning) · 6.2 s · $0.0004 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via CoreWeave ·
Argue both sides of free buses
A hard prompt: “Argue both sides of free buses”, in full.
Passed: Graded 4.3 of 5 on average (lowest 4).
Free buses work. They cut car trips, congestion, and carbon, and they help the riders who need them most. Boarding speeds up, so routes run faster and cheaper to operate. The cost is modest—often less than a parking garage or a lane widening—and downtown shops gain customers. A city that charges nothing at the farebox treats mobility as a right, not a fee. But somebody still pays. Fares cover a real slice of operating costs, and replacing them means higher taxes or cuts elsewhere—often taxes that hit the poor hardest. Most riders already pay, so the subsidy flows to everyone, including the comfortable. Crowding slows buses, driving away choice riders. Better to means-test discounts, keep service frequent, and spend scarce dollars where they move the most people.
423 tokens in, 282 out (122 of them reasoning) · 3.9 s · $0.0003 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via StreamLake ·
How it compares on writing
DeepSeek V4.1 Flash beside the job's picks among every model. 2 models share the hard-problems pick, and DeepSeek V4.1 Flash isn't one of them.
| Model | Passed | On Pro | Cost per reply |
|---|---|---|---|
| GLM 5.3Our pick: best value | 4 of 5 | 250 a month on Pro | $0.0006 |
| GPT-6 LunaOur pick: hard problems (shared), everyday | 5 of 5 | 60 a day on Pro | $0.0001 |
| DeepSeek V4.1 Flash | 4 of 5 | 60 a day on Pro | $0.0003 |
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
How these runs were done
Rubric (graded). The grader model scores the reply from 1 to 5 on each published criterion. It passes with an average of 4 or more and no criterion under 3, and only if it also meets the prompt's automatic rules (length, words it must or mustn't use).
The grader is Claude Opus 5.5 at low reasoning effort; its own replies are graded by GPT-6 Astra, so no model grades itself. Its prompt and every rubric are on the methods page.
DeepSeek V4.1 Flash, and writing, elsewhere
- DeepSeek V4.1 Flash: price, limits and messages on every plan
- The best AI for writing: our picks among every model
- DeepSeek for writing
- DeepSeek V4.1 Flash for coding: our test runs
- DeepSeek V4.1 Flash for math: our test runs
- DeepSeek V4.1 Flash for data analysis: our test runs
- Gemini 3.8 Flash vs DeepSeek V4.1 Flash
- GPT-6 Luna vs DeepSeek V4.1 Flash
- Rewrite for clarity: the prompt, and its price on each model
- Our test runs: 50 prompts on every model
Questions
Is DeepSeek V4.1 Flash good for writing?
In our test runs it passed 4 of 5 writing prompts, tied 3rd of the 15 models in llmwise. Every reply is on this page, so you can judge them yourself.
How many of my messages does a writing reply from DeepSeek V4.1 Flash use?
1 message each on Pro, where it has 60 a day on Pro. The price of a message is fixed and shown before you send it, however long the reply.
How were these runs done?
The same way for every model: each prompt sent through llmwise's own pipeline, each reply checked the same way. The methods page has every prompt and how each is scored.
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.